WorksheetsUTS STATMUL
Total questions: 86
Worksheet time: 43mins
Multivariate statistical analysis permit the researcher to consider the effects of three or more variables at the same time.
The basic types of multivariate techniques are metric The basic types of multivariate techniques are metric
True
False
The type of measurement scales used will determine which multivariate statistical techniques are appropriate for the data
True
False
Several dummy variables can be included in a regression model
True
False
In multiple regression, the dependent variable must be continuous and interval-scaled
True
False
In multiple regression, dummy variables are those that have no effect on the dependent variable.
True
False
In a regression equation, the beta coefficients indicate the effect on the dependent variable of a 1-unit increase in any of the independent variables.
True
False
Partial correlations measure the variance inflation among independent variables
True
False
In multiple regression, the coefficient of multiple determination indicates the percentage of the variation in Y that can be explained by all independent variables
True
False
Multicollinearity in regression analysis refers to how strongly interrelated the independent variables in a model are.
True
False
MANOVA predicts multiple continuous dependent variables with multiple continuous independent variables
True
False
Discriminant analysis predicts a categorical dependent variable based on a linear combination of independent variables.
True
False
To determine whether the discriminant analysis can be used as a good predictor, information provided in the “confusion matrix” is used.
True
False
The purpose of factor analysis is to summarize the information contained in a large number of variables into as large a number of factors as possible.
True
False
A factor loading indicates how strongly a measured variable is correlated with a factor
True
False
In cluster analysis, each cluster should have low internal homogeneity and high external heterogeneity
True
False
Multidimensional scaling provides a means for placing objects in multidimensional space on the basis of respondents’ judgments of the similarity of objects.
True
False
Which of the following is a mathematical way in which a set of variables can be represented with one equation?
structuralism
variate
ANOVA
synergy
When a multivariate statistical technique is used to predict a dependent variable from several independent variables, the researcher is studying:
dependence
independence
interdependence
segments
When a researcher is attempting to predict sales volume by using building permits, amount of advertising, and the income levels of residents, the researcher is using:
univariate analysis
a chi-square analysis
multiple regression analysis
factor analysis
A variable that is coded as either zero or one and that has two distinct levels is called a(n):
regression variable
dummy variable
MANOVA variable
ANOVA variable
If the regression equation is: Y = 98.3 +.35X1 + 22.3X2, the predicted value for Y when X1 = 3 and X2 = 5 is:
118.45
210.85
67.23
98.3
the number of observations
the degrees of freedom of the denominator
the number of independent variables
the sample size
In the formula for the F-test in multiple regression, n k - 1 stands for:
the degrees of freedom of the numerator
the number of observations
the degrees of freedom of the denominator
the number of independent variables
Jeff is analyzing data and is concerned over how strongly interrelated the independent variables in his model are. Jeff is concerned about:
multicollinearity
MANOVA
degrees of freedom
convergence
Which of the following is computed by most regression programs and provide an indication of how much multicollinearity exists among a set of independent variables?
x2
beta
collinear coefficient
variance inflation factor (VIF)
Which of the following suggests problems with multicollinearity?
VIF > 5.0
Beta < 3.0
Power > 0.8
Alfa > 0.8
Which type of analysis attempts to predict a categorical dependent variable?
factor analysis
discriminant analysis
regression analysis
linear analysis
If a bank wants to differentiate between successful and unsuccessful credit risks for home mortgage loans, it should use:
factor analysis
multidimensional scaling
MANOVA
discriminant analysis
In discriminant analysis, a linear combination of independent variables that explains group memberships is known as a(n):
regression equation
discriminant function
discriminant factor
n-way ANOVA
Which multivariate analysis statistically identifies a reduced number of factors from a larger number of measured variables?
factor analysis
regression
discriminant analysis
logit analysis
Which of the following indicates how strongly a measured variable is correlated with a factor?
factor
discriminator
factor link
factor loading
All of the following are examples of dependence methods of analysis EXCEPT:
multiple regression analysis
multiple discriminant analysis
cluster analysis
multivariate analysis of variance
Which of the following is an example of an interdependence analysis method?
multidimensional scaling
multiple regression analysis
conjoint analysis
all of the above
All of the following are examples of interdependence methods of analysis EXCEPT
factor analysis
cluster analysis
multidimensional scaling
conjoint analysis
Ordinary least squares is used to estimate a linear relationship between a firm's quantity sold per month and its total promotional expenditures and the slope of the linear function is found to be positive and significantly different from zero. Assuming that all other variables, including product price, were constant during the period covered by the data set, this result implies that
the firm should spend more on promotional expenditures
the firm should spend less on promotional expenditures
promotional expenditures influence demand
promotional expenditures have no influence on demand
The coefficient of determination
is maximized by ordinary least squares
has a value between zero and one
will generally increase if additional independent variables are added to a regression analysis
All of the above are correct
The coefficient of correlation is
a measure of the strength and direction of thelinear relationship between two variables.
equal to the size of the change in the Y variable that is caused by a change in the X variable
is equal to the proportion of the variation in the Y variable that is due to variations in the X variable
All of the above are correct
Multiple regression analysis is used when
there is not enough data to carry out simple linear regression analysis.
the dependent variable depends on more than one independent variable.
one or more of the assumptions of simple linear regression are not correct.
the relationship between the dependent variable and the independent variables cannot be described by a linear function.
The adjusted value of the coefficient of determination
will always increase if additional independent variables are added to the regression model.
is equal to the proportion of the sum of the squared deviations of the dependent variable from its mean that is explained by the regression model.
is always greater than the proportion of the sum of the squared deviations of the dependent variable from its mean that is explained by the regression model.
is always less than the proportion of the sumof the squared deviations of the dependent variable from its mean that is explained by the regression model
If the F test statistic for a regression is greater than the critical value from the F distribution, it implies that
none of the independent variables in the regression model have a significant effect on the dependent variable.
all of the independent variables in the regression model have significant effects on the dependent variable.
one or more of the independent variables in the regression model have a significant effect on the dependent variable.
None of the above
The standard error of the regression measures the
variability of the independent variable(s) relative to its (their) mean.
variability of the dependent variable relative to its mean.
variability of the dependent variable relative to the regression line.
average error that will result if the regression line is used to predict
Multicollinearity refers to a situation in which
successive error terms derived from the application of regression analysis to time series data are correlated.
there is a high degree of correlation between the independent variables included in a multiple regression model.
. the dependent variable is highly correlated with the independent variable(s) in a regression analysis.
the application of a multiple regression model yields estimates that are nonlinear in form
Autocorrelation refers to a situation in which
successive error terms derived from the application of regression analysis to time series data are correlated.
there is a high degree of correlation between two or more of the independent variables included in a multiple regression model.
the dependent variable is highly correlated with the independent variable(s) in a regression analysis.
the application of a multiple regression model yields estimates that are nonlinear in form.
Heteroskedasticity refers to a situation in which the error terms from a regression analysis
do not have equal variance.
are not normally distributed.
do not have a mean of zero
All of the above are correct
The Durbin-Watson statistic is used to test for
multicollinearity
autocorrelation
heteroskedasticity
All of the above are correct
Autocorrelation may be the result of
the omission of an important explanatory variable.
the presence of a trend in the independent variable.
nonlinearities in the relationship between the dependent and independent variables.
All of the above are correct
One advantage of estimating a function in which all variables have been transformed into their natural logarithms is that
problems with multicollinearity will be eliminated
problems with heteroskedasticity will be eliminated
the estimated slope coefficients are all elasticities.
None of the above is correct.
What does a multiple linear regression analysis examine?
The relationship between more than one dependent and only one independent variable
The relationship between one or more than one dependent and only one independent variable
The relationship between one dependent and more than one independent variables
The relationship between more than one independent variables
What does the following expression (H0:β1=β2=0) mean?
One of the independent variables is useful in predicting the dependent variable
Both of the independent variables are useful in predicting the dependent variable
None of the independent variables is useful in predicting the dependent variable
There is a third independent variable predicting the dependent variable
Which of the following criteria is the most optimal for assessing the goodness of the fit of a multiple linear regression model?
Adjusted R2
R2
The intercept
The coefficient
In which cases are the standardised coefficients suggested to be used to identify the relative importance of the independent variables in a multiple regression model?
When all the independent variables are measured using the same metric
When not all the independent variables are measured using the same metric
When all the independent variables are measured using an ordinal scale ranging from 1 to 6
What is the post estimation command that you can use after the regress command in Stata to compute the predicted mean-Y values of interest?
pcorr
esttab
margins
marginsplot
What is the Null Hypothesis in regression?
The response is significantly affected by the predictors
The slope of the regression line is zero
The slope of the regression line is not zero
None of these
The Correlation Coefficient between the two variables was found to be -0.90, this means:
There is a weak correlation
There is a strong correlation
There is a strong positive correlation
There is a very weak negative correlation
A term used to describe the case when the predictors in a multiple regression model are correlated is called:
homoscedasticity
heteroscedasticity
multicollinearity
polynomial
In the below Versus Fits plot on the right side, the spread of residuals is increasing with the increase in the Fitted Value. Which of the regression assumptions is violated in this example?
homoscedasticity
independence
normality
multicollinearity
If the analysis predicts several continuous dependent variables with several categorical independent variables, the appropriate statistical technique is:
multiple regression
multiple discriminant analysis
conjoint analysis
MANOVA
A researcher has 57 variables in a large dataset and wishes to summarize the information from them into a reduced set of variables. Which multivariate technique should be used?
factor analysis
multidimensional scaling
logit analysis
regression analysis
In cluster analysis, the researcher wants clusters to have high ____ within-clusters and high between-cluster ____.
independence; dependence
significance; insignificance
heterogeneity; homogeneity
homogeneity; heterogeneity
A mathematical way of simplifying factor analysis results is
factor loading
factor reduction
factor rotation
factor analysis
General Mills would like to "see" a picture of how its brands are perceived by consumers compared to competitive brands. Which statistical technique can measure brands in multidimensional space on the basis of respondents' judgements of the similarity of the brands?
structural equations modeling
factor analysis
multidimensional scaling
partial positioning
Which technique allows a researcher to build and test a theory represented by a series of regression equations, each involving multiple item measures, that are solved simultaneously?
structural equations modeling (SEM)
synergistic regression
sequential regression modeling
sequential estimation modeling (SEM)
A multivariate tool that combines a factor analytic and regression approach to provide path estimates to a proposed model but falls short of providing an assessment of fit is called
partial least squares (PLS)
MANOVA
partial correlations
discriminant analysis
Pizza topping: olives - anchovies - pepperoni - banana
What type of question should be used?
Nominal
Ratio
Ordinal
Interval
TIme: 2,5 min - 5 min - 7,5 min
What type of question should be used?
Nominal
Ratio
Ordinal
Interval
Socioeconomics status: lower class - middle class - upper class
What type of question should be used?
Nominal
Ratio
Ordinal
Interval
Time: 1 o'clock - 2 o'clock - 3 o'clock
What type of question should be used?
Nominal
Ratio
Ordinal
Interval
none of them
0,75
0,25
4,00
Multicollinearity refers to the correlation among three or more independent variables. What is the impact of multicollinearity
Reduce any single independent variable’s unique predictive power by the extent to which it is associated with the other independent variables
Reduce any single independent variables which it is associated with dependent variable
The ability of an additional variable to improve the independent variable
Minimize the prediction from a given number of independent variables
Which is the incorrect statement about Sample Size considerations
The minimum ratio of observation is 5:1, but the preferred ratio is 20:1
Simple regression can be effective with a sample size of 20 but maintaining power at .80
The preferred ratio is 15:1, which should decrease when stepwise estimation is used
Requires a minimum sample of 50 and preferably 100 observations for most research situations
Based on the following figure, the interpretation from result of a Breusch-Pagan test
Reject HO because homoscedasticity is present
Fail to reject HO because homoscedasticity is present
Reject HO because heteroscedasticity is exists
Fail to reject HO because heteroscedasticity does not exists
Income & logHIV
logHB & logP
const, logHIV, income
logHB, logP, logT
The figure below is the coefficient value of the independent variable. Which statement is correct based on the picture below?
The highest beta weight is const
The lowest beta weight is logHIV
Income is the most independent variable on the dependent variable
All of the above are correct
Which is the correct statement based on the following figure
The value of R square means that the influence of independent variables on y is 85,4%
As per the above results, probability is close to zero. This implies that overall the regressions is not meaningful
84,9% variation in independent variables is explained by y
The value of Adj. R - squared decreases only when an additional variable adds to the explanatory power to the regression
Which statement is incorrect about interpreting the regression variate
Use the results of the regression model to interpret the unique impact of each independent variable relative to the other variables in the model
Use beta weights as a measure of comparing relative importance among dependent variable
Regression coefficients describe changes in the dependent variables
Multicollinearity may be considered "good" when it reveals a suppressor effect but
generally harmful
How can a multiple linear regression result be validated?
Re-assess the sample used for the study and compare the results
Get another sample from the population and compare the results
Infer from the R squared result
Any of the above
Which of the following data type can be used for the dependent variable of multiple linear regression? (1) Nominal (2) Ordinal (3) Interval (4) Rati0
1234
1,2
1,4
3,4
Which is the example of data that can be used as the dependant variable of a multiple linear regression study?
Weight
Hair type
Over 100K USD Income
Marital Status
How much missing data can be deleted?
15%
20%
30%
35%
How can outliers be detected
Use boxplot
Map two variables at a time
Mahalonobis distance
All of the above
Interval data has meaningful value of zero
True
False
HepatitisB has high correlation to other independent variables
Thinnes and HIV indicate multicollinearity occurs in a regression model
The result of HIV and HepatitisB indicate no correlation that means the multicollinearity occurs
All of the options are incorrect
The incorrect assumptions of multiple linear regression
Residuals come from a population that have constant variance
The error terms of the model are normally distributed
There is correlation between the residuals
There is a linear relationship between the predictors and the response variable
-99,5%
-50%
50%
99,5%
Presumably strong indication of multicollinearity.
95% confidence interval is left tailed
Indicate a slight negative autocorrelation.
R-Squared shows a "good" fit indication of the model.
