wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

STATMUL

Total questions: 50

Worksheet time: 25mins

Name
Class
Date
1.

Which statement best describes multivariate analysis?

a)

An analysis that considers only one variable at a time.

b)

An analysis of the relationship between only two variables without involving others.

c)

An analysis of multiple variables (more than two) simultaneously in one study.

d)

A data analysis that requires no statistical tests at all.

2.

Multivariate analysis should be used when...

a)

the study has only one dependent and one independent variable.

b)

all collected data are nominal scale.

c)

the aim is only simple descriptive statistics.

d)

the study involves many interrelated variables that need to be analyzed simultaneously.

3.

The measurement scale of variables must be considered when choosing multivariate techniques because...

a)

the scale does not affect the choice of statistical methods.

b)

the scale type (nominal, ordinal, interval, ratio) determines the appropriate multivariate method.

c)

all multivariate techniques require interval/ratio-scaled variables.

d)

nominal variables cannot be used in any multivariate analysis.

4.

How can multivariate outliers be detected?

a)

Create histograms for each variable and inspect distributions.

b)

Use the Durbin–Watson test on regression residuals.

c)

Use scatterplots for every variable pair.

d)

Compute Mahalanobis distance for each observation and flag those with the largest distances as outliers.

5.

In handling missing data, listwise deletion means...

a)

deleting variables with the highest missing percentage.

b)

deleting any observation that has at least one missing value.

c)

replacing missing values with the variable mean.

d)

ignoring missing data and using all available values.

6.

Which is a common assumption in many multivariate analyses?

a)

multivariate normality of the data.

b)

a nominal dependent variable.

c)

no more than three independent variables.

d)

all variables share identical ranges.

7.

If the assumptions underlying a multivariate technique are violated, a likely consequence is...

a)

the p-value automatically becomes 1.00.

b)

the regression model cannot be computed at all.

c)

the analysis results may be biased or invalid.

d)

the analysis becomes statistically stronger.

8.

The minimum sample size for multivariate analysis generally...

a)

is unimportant if the correct technique is chosen.

b)

increases with the number of variables and model complexity.

c)

is smaller than for univariate analysis because multivariate is more efficient.

d)

is always 30 observations by rule of thumb.

9.

To include a categorical (non-metric) variable in models requiring metric variables, you should...

a)

assign numeric codes (1,2,3,...) and treat them as interval.

b)

use dummy coding by creating 0/1 variables for the categories.

c)

ignore the categorical variable as it cannot be used.

d)

combine its categories into a single composite score.

10.

The term multivariate measurement refers to...

a)

measuring each construct with a single variable.

b)

analyzing many variables simultaneously.

c)

using interval or ratio scales for all variables.

d)

using two or more indicators to represent the same construct.

11.

A key feature of interdependence techniques is...

a)

no designated dependent variable; all variables are analyzed simultaneously.

b)

a single dependent variable influenced by independent variables.

c)

only two variables are analyzed.

d)

they require special statistical software.

12.

Which of the following is an interdependence technique?

a)

Multiple Regression Analysis

b)

Discriminant Analysis

c)

Exploratory Factor Analysis (EFA)

d)

Logistic Regression

13.

Which pair of methods are both dependence techniques?

a)

EFA and Discriminant Analysis

b)

EFA and Cluster Analysis

c)

Cluster Analysis and Logistic Regression

d)

Multiple Regression and Discriminant Analysis

14.

Classifying techniques as dependence vs interdependence is based on...

a)

the number of variables analyzed.

b)

the subject area of the study.

c)

whether a dependent variable is specified.

d)

the software or algorithm used.

15.

A researcher wants to predict a dependent variable (Y) from multiple independent variables (X1, X2, ...). The most suitable technique is...

a)

Factor Analysis

b)

Multiple Linear Regression

c)

Cluster Analysis

d)

Multidimensional Scaling (MDS)

16.

A researcher wants to segment 30 cities into homogeneous groups based on socio-economic indicators. The best technique is...

a)

Multiple Regression

b)

Discriminant Analysis

c)

ANOVA

d)

Cluster Analysis

17.

When choosing a multivariate technique, a researcher should consider...

a)

personal preference for a method.

b)

research objectives and variable characteristics (scale and dependent/independent roles).

c)

ease of interpreting software plots.

d)

the number of researchers involved.

18.

A key difference between cluster analysis and discriminant analysis is...

a)

cluster analysis is exploratory with no dependent variable, while discriminant analysis predicts a predefined categorical dependent variable.

b)

cluster uses categorical variables; discriminant uses metric variables.

c)

discriminant is interdependence; cluster is dependence.

d)

cluster is always more accurate than discriminant.

19.

Although both are interdependence techniques, EFA differs from clustering because...

a)

EFA requires a dependent variable; clustering does not.

b)

EFA groups variables by correlation patterns; clustering groups objects/cases by similarity.

c)

clustering uses variable correlation matrices, while EFA uses object distances.

d)

EFA is only for metric data; clustering only for categorical data.

20.

If the dependent variable is categorical (e.g., yes/no), the appropriate technique is...

a)

Multiple Linear Regression

b)

Factor Analysis

c)

Logistic Regression

d)

Cluster Analysis

21.

The main goal of Exploratory Factor Analysis (EFA) is to...

a)

segment respondents into homogeneous groups.

b)

predict a dependent variable from several independents.

c)

test mean differences across groups.

d)

identify latent factors that reduce many variables into fewer factors.

22.

The KMO index is used to...

a)

measure correlations among extracted factors.

b)

assess sampling adequacy for factor analysis.

c)

determine the optimal number of factors.

d)

test multivariate normality.

23.

Bartlett’s Test of Sphericity tests...

a)

whether the correlation matrix differs from identity, indicating sufficient correlations for EFA.

b)

whether data are normally distributed.

c)

whether the number of factors is correct.

d)

whether factor variances are homogeneous.

24.

By Kaiser’s criterion, retain factors with eigenvalues...

a)

greater than 0.5

b)

in the first component only

c)

less than 1.0

d)

greater than 1.0

25.

A scree plot is useful to...

a)

display correlations among variables.

b)

detect multivariate outliers.

c)

identify the number of factors via the elbow in eigenvalues.

d)

inspect residual distributions.

26.

Factor loadings represent...

a)

variable communalities.

b)

correlations between observed variables and factors.

c)

variable weights forming clusters.

d)

factor scores per observation.

27.

Communality indicates...

a)

the proportion of a variable’s variance explained by the factors.

b)

the magnitude of inter-variable correlations.

c)

the eigenvalue of the principal factor.

d)

the difference between initial and final variance.

28.

Factor rotation aims to...

a)

increase the first factor’s eigenvalue.

b)

reduce the number of factors.

c)

prevent factor correlations.

d)

clarify loading structure (high on one factor, low on others).

29.

The difference between orthogonal and oblique rotation is...

a)

orthogonal only for PCA, oblique for FA.

b)

orthogonal disallows factor correlations; oblique allows them.

c)

orthogonal yields more factors.

d)

oblique is always easier to interpret.

30.

A criterion to retain an item in factor models is...

a)

a high loading (≥0.50) on one factor with no high cross-loadings.

b)

very low communality (<0.30).

c)

moderate loadings on many factors.

d)

nominal/ordinal measurement.

31.

The primary goal of cluster analysis is to...

a)

reduce variables into factors.

b)

group objects into homogeneous clusters.

c)

predict a dependent variable.

d)

test causal relationships.

32.

In clustering, Euclidean distance is used to...

a)

determine the number of variables.

b)

compute p-values.

c)

measure dissimilarity between objects.

d)

measure inter-variable correlation.

33.

Standardizing variables before clustering is needed because...

a)

comparable scales prevent domination in distance computations.

b)

it ensures perfect normality.

c)

it auto-optimizes the number of clusters.

d)

it automatically removes outliers.

34.

A difference between hierarchical and k-means clustering is...

a)

hierarchical for categorical data; k-means for metric data.

b)

k-means is unsuitable for large data; hierarchical is efficient for big data.

c)

in k-means, cluster membership never changes.

d)

hierarchical needs no preset number of clusters; k-means does.

35.

In hierarchical clustering, once objects merge into a cluster they...

a)

cannot be separated or moved later.

b)

can be moved if closer to another cluster.

c)

become outliers.

d)

must be removed from analysis.

36.

Ward’s method is characterized by...

a)

complete linkage (farthest distance).

b)

susceptibility to chaining.

c)

minimizing the increase in within-cluster variance when merging.

d)

applicability only to non-metric variables.

37.

A dendrogram is...

a)

a plot of factor relationships in EFA.

b)

a tree diagram of cluster merging to choose the number of clusters.

c)

a regression residual plot.

d)

a multivariate correlation chart.

38.

To choose the optimal number of hierarchical clusters, you can...

a)

make clusters of equal size.

b)

use a cluster chi-square test.

c)

pick an arbitrary number.

d)

identify a large jump in fusion distances and cut before it.

39.

Extreme outliers can affect clustering because they...

a)

are automatically ignored by algorithms.

b)

increase the number of factors.

c)

may form their own cluster or distort distances.

d)

have no effect if n>100.

40.

A drawback of k-means is...

a)

it automatically fixes the number of clusters.

b)

results depend on initial seeds, risking local optima.

c)

it cannot handle more than 5 variables.

d)

it always yields equal-sized clusters.

41.

Multiple linear regression is appropriate when...

a)

all variables are categorical.

b)

there is only one independent variable.

c)

the dependent is metric and ≥2 independents are metric or dummied.

d)

relationships are non-linear.

42.

The coefficient of determination (R^2) indicates...

a)

the square root of the Y–X correlation.

b)

the average model error.

c)

the overall model p-value.

d)

the proportion of Y variance explained by the X’s.

43.

A slope coefficient of 2.5 for X means...

a)

a 1-unit increase in X raises Y by 2.5 on average (others constant).

b)

Y equals 2.5 when X=0.

c)

the X–Y correlation is 2.5.

d)

X causes Y to change 2.5-fold.

44.

A categorical variable can be included in regression by...

a)

coding as 1,2,3 and treating as interval.

b)

dropping a category to make it dichotomous.

c)

creating dummy variables (0/1) for categories.

d)

excluding it entirely.

45.

Homoscedasticity means...

a)

no multicollinearity.

b)

normally distributed residuals.

c)

no outliers.

d)

constant variance of residuals across levels of predicted values.

46.

A sign of high multicollinearity is...

a)

Durbin–Watson ≈ 2.

b)

residual Q–Q plot near the diagonal.

c)

large VIF (e.g., >10).

d)

very low R^2.

47.

High multicollinearity leads to...

a)

R^2 near 0.

b)

unstable coefficients with large SEs, hindering interpretation of each X’s effect.

c)

non-normal Y.

d)

model cannot be computed.

48.

The F-test in regression assesses...

a)

whether the overall model significantly explains Y.

b)

whether each individual coefficient differs from zero.

c)

residual normality.

d)

heteroscedasticity.

49.

The Durbin–Watson statistic detects...

a)

multicollinearity.

b)

homoscedasticity.

c)

residual autocorrelation.

d)

normality of Y.

50.

An influential observation can be detected by...

a)

a standardized residual near 0.

b)

a large Cook’s Distance or leverage.

c)

a high VIF on that variable.

d)

a very small t-test p-value.