Font size
WorksheetsQMB 3200 EXAM 2
Total questions: 112
Worksheet time: 56mins
If a hypothesis is rejected at a 5% level of significance, it _____.
will always be accepted at the 1% level
will always be rejected at the 1% level
may be rejected or not rejected at the 1% level
will never be tested at the 1% level
If a hypothesis test leads to the rejection of the null hypothesis, a _____.
Type II error may have been committed
Type II error must have been committed
Type I error must have been committed
Type I error may have been committed
Two approaches to drawing a conclusion in a hypothesis test are _____.
one-tailed and two-tailed
null and alternative
p-value and critical value
Type I and Type II
Which of the following hypotheses is not a valid null hypothesis?
H0: μ ≥ 0
H0: μ < 0
H0: μ ≤ 0
H0: μ = 0
When the rejection region is in the lower tail of the sampling distribution, the p-value is the area under the curve _____.
less than or equal to the test statistic
greater than or equal to the critical value
greater than or equal to the test statistic
less than or equal to the critical value
Excel's __________ function can be used to calculate a p-value for a hypothesis test.
NORM.S.INV
NORM.S.DIST
COUNTIF
RAND
In the hypothesis testing procedure, α is _____.
the confidence level
the level of significance
1 − level of significance
the critical value
In a two-tailed hypothesis test, the null hypothesis should be rejected if the p-value is _____.
less than or equal to 2α
greater than or equal to α
greater than or equal to 2α
less than or equal to α
Which of the following is an improper form of the null and alternative hypotheses?
H0: μ = μ0 and Ha: μ ≠ μ0
H0: μ < μ0 and Ha: μ ≥ μ0
H0: μ ≥ μ0 and Ha: μ < μ0
H0: μ ≤ μ0 and Ha: μ > μ0
When each data value in one sample is matched with a corresponding data value in another sample, the samples are known as _____.
matched samples
independent samples
corresponding samples
dependent samples
The standard error of x̄1 - x̄2 is the _____.
standard deviation of the sampling distribution of x̄1 - x̄2
variance of the sampling distribution of x̄1 - x̄2
variance of x̄1 - x̄2
difference between the two means
A company wants to identify which of two production methods has the smaller completion time. One sample of workers is randomly selected and each worker first uses one method and then uses the other method. The sampling procedure being used to collect completion time data is based on _____ samples.
pooled
cross
independent
matched
Independent simple random samples are selected to test the difference between the means of two populations whose variances are not known. The sample sizes are n1 = 32 and n2 = 40. The correct distribution to use is the _____ distribution.
binomial
normal
t
uniform
In a simple regression analysis (where y is a dependent and x an independent variable), if the slope is positive, then it must be true that _____.
there is no correlation between x and y
there is a negative correlation between x and y
the y-intercept is 0
there is a positive correlation between x and y
Which of the following is correct?
SSR = SSE + SST
SSE = SSR + SST
SST = (SSR)2
SST = SSR + SSE
The proportion of the variation in the dependent variable y that is explained by the estimated regression equation is measured by the _____.
coefficient of determination
standard error of the estimate
correlation coefficient
confidence interval estimate
The least squares criterion is _____.
min Σ(yi - ŷi)2
min Σ(yi - ȳ)2
min Σ(xi - yi)2
min (Σyi - ŷi)2
A news reporter states that the average number of temperature in January has never dropped below 10 degrees Fahrenheit. You go online to research this claim. The appropriate hypotheses are:
H0: μ≤10 Ha: μ>10
H0: μ=10 Ha: μ=10
H0: μ>10 Ha: μ<10
H0: μ≥10 Ha: μ<10
The p-value
can be any value, negative or positive.
can be any positive value.
can be any negative value.
must be a number between 0 and 1.
A local bakery generates $2,500 of revenue per day, on average. The bakery's owner plans to increase expenses on the advertisement, hoping that this results in higher revenue. What is a Type I error for this situation?
A Type I error for this would be to conclude that the advertisement would result in lower revenue when the advertisement actually increases revenue.
A Type I error for this situation would be to fail to conclude that the advertisement would result in higher revenue when the advertisement actually increases revenue.
A Type I error for this situation would be to conclude that the advertisement would result in higher revenue when the advertisement actually does not increase revenue.
A Type I error for this situation would be to conclude that the advertisement has no effect on revenue when it has.
The normal probability distribution can be used to approximate the sampling distribution of as long as:
np≥5 and n(1−p)≥5
n≥30
n≥5
np≥30 and n(1−p)≥30
A student wants to determine if pennies are really fair, meaning equally likely to land heads up or tails up. He flips a random sample of 50 pennies and finds that 28 of them land heads up. What are the appropriate null and alternative hypotheses?
H0: p=.5 Ha: p=.5
H0: p≤.5 Ha: p>.5
H0: p≥.5, Ha: p<.5
H0: p≥28, Ha: p<28
Peter spends about 18 minutes repairing one smartphone, on average. After purchasing new equipment, he expects to perform some operations quicker, reducing the average repair time. What is a Type II error for this situation?
A Type II error for this situation would be to conclude that the new equipment would result in a higher repair time when the new equipment actually does not increase repair time.
A Type II error for this situation would be to fail to conclude that the new equipment would result in a lower repair time when the new equipment actually declines repair time.
A Type II error for this situation would be to conclude that the new equipment would result in a lower repair time when the new equipment actually does not decline repair time.
A Type II error for this situation would be to conclude that the new equipment does not change repair time when it does.
For a lower tail test, the p-value is the probability of obtaining a value for the test statistic:
at least as small as that provided by the population.
at least as large as that provided by the sample.
at least as small as that provided by the sample.
at least as large as that provided by the population.
The average hourly wage of computer programmers with 2 years of experience has been $21.80. Because of high demand for computer programmers, it is believed there has been a significant increase in the average wage of computer programmers. To test whether there has been an increase, the correct hypotheses to be tested are:
H0: μ≤21.80, Ha: μ>21.80
H0: μ=21.80, Ha: μ=21.80
H0: μ>21.80, Ha: μ<21.80
H0: μ≥21.80, Ha: μ<21.80
The average gasoline price of one of the major oil companies has been hovering around $2.20 per gallon. Because of cost reduction measures, it is announced that there will be a significant reduction in the average price over the next month. To test this, we wait one month, then randomly select a sample of 36 of the company's gas stations. We find that the average price for the stations in the sample was $2.15. The standard deviation of the prices for the selected gas stations is $.10. State the appropriate null and alternative hypotheses for testing the company's claim.
H0: μ≥2.20, Ha: μ<2.20
H0: xbar ≤2.15, Ha: xbar>2.15
H0: xbar ≥2.15, Ha: xbar<2.15
H0: μ≤2.20, Ha: μ>2.20
In hypothesis testing, if the null hypothesis has not been rejected when the null hypothesis has been true,
the correct decision has been made.
a Type II error has been committed.
a Type I error has been committed.
the level of significance is too low.
The average number of hours for a random sample of mail order pharmacists from company A was 50.1 hours last year. It is believed that changes to medical insurance have led to a reduction in the average work week. To test the validity of this belief, the hypotheses are:
H0: μ≤50.1, Ha: μ>50.1
H0: μ=50.1, Ha: μ=50.1
H0: μ>50.1, Ha: μ<50.1
H0: μ≥50.1, Ha: μ<50.1
For the case where σ is unknown, the test statistic has a t distribution. How many degrees of freedom does it have?
5
n-1
n
30
As the test statistic becomes larger, the p-value:
stays the same, since the sample size has not been changed.
becomes larger.
becomes smaller.
becomes negative.
Applications of hypothesis testing that only control for the Type I error are called:
assumption of calculations.
research tests.
level of significance.
significance tests.
The p-value is a probability that measures the evidence against the:
population
alternative hypothesis.
null hypothesis.
sample statistic.
Whenever the probability of making a Type II error has not been determined and controlled, only two conclusions are possible. We either reject H0 or:
do not reject H0.
reject Ha.
accept H0.
do not reject Ha.
Which of the following null hypotheses cannot be correct?
H0: μ=10
H0: μ≥10
H0: μ≤10
H0: μ=10
What are the most common choices for the level of significance?
0.95 and 0.99
0.01 and 0.05
0.10 and 0.20
0.50 and 0.75
For the case where σ is unknown, which statistic is used to estimate σ?
pbar
s
xbar
n
the hypothesis tentatively assumed true in the hypothesis testing procedure
Null hypothesis
type 1 error
alternative hypothesis
type 2 error
A hypothesis test in which rejection of the null hypothesis occurs for values of the test statistic in either tail of its sampling distribution
one-tailed test
critical value
two-tailed test
level of significance
The hypothesis concluded to be true if the null hypothesis is rejected. Known as the research hypothesis
null hypothesis
alternative hypothesis
p-value
level of significance
A probability that provides a measure of the evidence against the null hypothesis provided by the sample/probability used to test the null hypothesis. Smaller _______ indicate more evidence against H0.
p-value
type 1 error
type 2 error
null hypothesis
For a ______ test, the p-value is the probability of obtaining a value for the test statistic as small as or smaller than that provided by the sample
lower tail
upper tail
two-tailed
For an______ test, the p-value is the probability of obtaining a value for the test statistic as large as or larger than that provided by the sample
lower tail
upper tail
two-tailed
For a _______ test, the p-value is the probability of obtaining a value for the test statistic at least as unlikely as or more unlikely than that provided by the sample
lower tail
upper tail
two-tailed
The error/probability of rejecting H0 when it is true
type 1 error
level of significance
type 2 error
critical value
The probability of making a Type 1 error when the null hypothesis is true as an equality
two-tailed test
critical value
p-value
level of significance
The error/probability of accepting/not rejecting H0 when it is false/should have been rejected
one tail test
type 2 error
type 1 error
two tailed test
A value that is compared with the test statistic to determine whether H0 should be rejected
critical value
p-value
one tail test
two tail test
A hypothesis test in which rejection of the null hypothesis occurs for values of the test statistic in one tail of its sampling distribution
two tail test
type 1 error
one tail test
type 2 error
Samples selected from two populations in such a way that the elements making up one sample are chosen independently of the elements making up the other sample
Independent random samples/Independent sample design
Matched Samples/Matched sample design
Pooled Estimator of p
Prediction interval
Samples in which each data value of one sample is matched with a corresponding data value of the other sample
Independent random samples/Independent sample design
Matched Samples/Matched sample design
Pooled Estimator of p
Prediction interval
An estimator of a population proportion obtained by computing a weighted average of the point estimators obtained from two independent samples
Independent random samples/Independent sample design
Matched Samples/Matched sample design
Pooled Estimator of p
Prediction interval
The variable that is doing the predicting or explaining. It is denoted by x
Independent variable
Dependent variable
Outlier
MSR
The variable that is being predicted or explained. It is denoted by y
Independent variable
Dependent variable
outlier
MSR
The difference between the observed value of the dependent variable (yi) and the value predicted variable of the dependent variable (y-hat i) using the estimated regression equation; for the ith observation the ith residual is (yi -y-hat i). Represents the error in using (y-hat i) to estimate yi. The value of ___ is a measure of the error in using the estimated regression equation to predict the values of the dependent variable in the sample. Can think of a measure of how well the observations cluster about the (y-hat) line. Can be thought of as the unexpected portion of SST
SSE
SSR
MSE
MSR
The difference (yi- ybar) provides a measure of the error involved in using (ybar) to predict sales. The corresponding sum of squares, called the total sum of squares. Can think of ___ as a measure of how well the observations cluster about the ybar line
SSR
SST
SSE
MSE
To measure how much the (yhat) values on the estimated regression line deviate from (yhat), another sum of squares is computed. This sum of squares, called the sum of squares due to regression. Can be thought of as the explained portion of SST
SSR
SSE
SST
MSE
The unbiased estimate of the variance of the error term σ2 . It is called mean square error or s2
MSE
MSR
SSR
SSE
The sum of squares due to regression (SSR), divided by its degrees of freedom provides another independent estimate of σ2 . This estimate in called the mean square sue to regression, or simple mean square regression
MSE
MSR
SST
SSR
Regression analysis involving one independent variable and one dependent variable in which the relationship between the variable is approximated by a straight line
MSR
Regression model
Estimated regression equation
Simple linear regression
The equation that describes how y is related to x and an error term
Regression equation
Regression model
Coefficient of determination
Standard error of the estimate
The estimate of the regression equation developed from sample data by using the least squares method
Estimated regression equation
Scatter diagram
Regression model
Regression equation
The equation that describes how the mean or expected value of the dependent variable is related to the independent variable
Regression equation
Regression model
Estimated regression equation
Residual
A graph of bivariate data in which the independent variable is on the horizontal axis and the dependent variable is on the vertical axis
Scatter diagram
Residual plot
Standardized Residual
Influential Observations
A measure of the goodness of fit of the estimated regression equation. It can be interpreted as the proportion of the variability in the dependent variable y that is explained by the estimated regression equation
Standard error of the estimate
Coefficient of determination
Residual plot
Outlier
The square root of the mean square error, denoted by x. It is the estimate of 𝜎, the standard deviation of the error term
Regression equation
Standardized Residual
Standard error of the estimate
Influential Observations
The interval estimate of the mean value of y for a given value of x
Confidence interval
Prediction interval
The interval estimate of an individual value of y for a given value of x
Confidence interval
Prediction interval
Graphical representation of the residuals used to determine whether the assumptions made about the regression model appear to be valid
Residual plot
Residual
Outlier
Regression model
– the analysis of the residuals used to determine whether the assumptions made about the regression model appear to be valid. Residual analysis is also used to identify outliers and influential observations
Outlier
Residual
Standardized Residual
Standard error of the estimate
The value obtained by dividing a residual by its standard deviation
Standardized Residual
Coefficient of determination
Residual plot
Scatter diagram
A data point or observation that does not fit the trend shown by the remaining data
Outlier
Scatter diagram
Residual
Confidence interval
An observation that has a strong influence or effect on the regression results
Influential Observations
Coefficient of determination
Standard error of the estimate
Simple linear regression
Which of the following scenarios follows a matched sample design?
A dietitian had 50 clients follow a calorie-counting diet and another 50 clients follow a low-carb diet to see which is more effective for weight loss.
A company looks at the satisfaction of men and women to see which gender is more satisfied with the current work conditions.
A farmer tracks his sales of red and green apples to see which is preferred by his customers.
A teacher uses a pretest and then a posttest with her students to see how much they have improved.
When completing a two-tailed hypothesis test about the difference between two population means, the
samples must be of the same size.
p-value must be doubled.
test statistic must be doubled.
sample sizes must be added.
Regarding hypothesis tests about p1 - p2 , the pooled estimate of P is a:
simple average of and p-bar1 and p-bar2
the sum of p-bar1 and p-bar2
weighted average of p-bar1 and p-bar2
the difference of p-bar1 and p-bar2
To construct an interval estimate for the difference between the means of two populations with sample sizes of n1 and n2 when the two population standard deviations are known, what are the degrees of freedom to compute the z value needed for the interval estimate?
The degrees of freedom are n1 + n2.
The degrees of freedom are max(n1, n2).
The degrees of freedom are n1 + n2 - 2.
The z-distribution is independent of the degrees of freedom.
Regarding inferences about the difference between two population means, the alternative to the matched sample design is:
systematic samples.
independent samples.
dependent samples.
mutually exclusive samples.
A researcher recruits 25 people to participate in a study on alcohol consumption and its interactions with Tylenol. The 25 participants had to come to a check-in center every day at 7:00 a.m. for one week. They were given various amounts of alcohol. Each day, each participant would flip a coin to determine if they also took Tylenol with their alcohol. They found that their BAC was 25% higher on days when they were given Tylenol with their alcohol than when they drank alcohol alone. This is an example of a(n):
independent sample design.
double blind experimental design.
matched sample design.
dependent sample design.
If we are interested in testing whether the proportion of items in population 1 is larger than the proportion of items in population 2, then the:
null hypothesis should state p1-p2 >0
alternative hypothesis should state p1-p2 >0
null hypothesis should state p1-p2 <0
alternative hypothesis should state p1-p2 <0
The matched sample design often leads to a smaller sampling error than the independent sample design. The primary reason is that in a matched sample design:
variation between subjects is eliminated because the same subjects are used for both treatments.
variation in the response variable is eliminated because a control group is being utilized.
variation between the treatments is reduced because the sample size is essentially double.
variation in the sample design is reduced because the matched sample design creates pairs.
Suppose we have a t distribution based upon two sample means with unknown population standard deviations, which we are unwilling to assume are equal. When we calculate the appropriate degrees of freedom, we should:
round the calculated degrees of freedom down to the nearest integer.
not round the calculated degrees of freedom.
round the calculated degrees of freedom up to the nearest integer.
add the two sample sizes together and subtract 2.
A professor of statistics wants to identify whether exam scores at the end of the second semester compared to the first semester improve among her 40 students. The sampling procedure being used to collect data is based on
dependent samples.
independent samples.
convenience samples.
matched samples.
In regression analysis, the equation in the form y = 𝛽0 + 𝛽1x + ε is called the:
simple linear regression model.
estimated regression equation.
regression equation.
correlation equation.
The tests of significance in regression analysis are based on assumptions about the error term ɛ . One such assumption is that the variance of ɛ, denoted by 𝝈2, is:
unrelated to the value of x.
the same for all values of x.
greater as x increases.
less as x increases.
The tests of significance in regression analysis are based on assumptions about the error term ɛ. One such assumption is that the error term follows ɛ a(n) _____ distribution for all values of x.
uniform
normal
binomial
exponential
If a significant relationship exists between x and y and the coefficient of determination shows that the fit is good, the estimated regression equation should be useful for:
determining cause and effect.
estimation and prediction.
determining nonresponse error.
extrapolation
The value of the coefficient of correlation (r):
is always larger than the value of the coefficient of determination.
is always smaller than the value of the coefficient of determination.
can be equal to the value of the coefficient of determination (r2).
can never be equal to the value of the coefficient of determination (r2).
In regression analysis, the variable that is being predicted is the:
random variable.
confounding variable.
dependent variable.
independent variable.
Graphical representation of the residuals that can be used to determine whether the assumptions made about the regression model appear to be valid is called a:
normal probability plot.
regression plot.
scatter diagram.
residual plot.
Observations with extreme values for the independent variables are called:
high leverage points.
influential observations.
outliers
mistakes
When working with regression analysis, an outlier is:
any value that has a small residual.
any observation that is extreme in the x direction.
any observation that does not fit the trend shown by the remaining data.
any value that falls more than 1.5(IQR) above Q3 or below Q1
The model developed from sample data that has the form is known as the:
simple linear regression model.
estimated simple linear regression equation.
simple linear regression equation.
correlation equation.
If a residual plot of x versus the residuals, y - ŷ, shows a non-linear pattern, then we should conclude that:
the regression model is not an adequate representation of the relationship between the variables.
the regression model is useful for making predictions.
the regression model was not based upon a large enough sample size.
the regression model describes the relationship between x and y very well.
The mathematical equation relating the independent variable to the expected value of the dependent variable, , is known as the:
regression model.
estimated regression equation.
simple linear regression equation.
correlation equation.
The coefficient of determination:
can be negative or positive.
is the same as the coefficient of correlation.
cannot be negative.
is the square root of the coefficient of correlation.
Suppose a residual plot of x verses the residuals, y - ŷ, shows a nonconstant variance. In particular, as the values of x increase, suppose that the values of the residuals also increase. This means that:
as the values of x get larger, the standard deviation of the residuals becomes smaller.
as the values of x get larger, the values of y become larger.
as the values of x get larger, the ability to predict y becomes less accurate.
as the values of x get larger, the error term, , becomes smaller.
If you suspect that you have an influential observation, the first thing you should do is:
increase the value of the slope.
re-record the data to see if the observation shows up again.
remove the influential observation from the data set.
check to make sure no error has been made in collecting or recording data.
Larger values of r2 imply that the observations are more closely grouped about the:
origin
least squares line.
average value of the independent variables.
average value of the dependent variable.
The tests of significance in regression analysis are based on several assumptions about the error term ɛ. Additionally, we make an assumption about the form of the relationship between x and y. We assume that the relationship between x and y is:
quadratic
linear
constant
exponential
The tests of significance in regression analysis are based on assumptions about the error term ɛ. One such assumption is that the values of ɛ are:
limited
uniformly distributed.
independent
categorical
The difference between the observed value of the dependent variable and the value predicted using the estimated regression equation is called a(n):
point estimate.
prediction
residual
outlier.
The tests of significance in regression analysis are based on assumptions about the error term ɛ. One such assumption is that the error term ɛ is a random variable with a mean or expected value of:
x̄
ŷ
0
1
When studying the relationship between two quantitative variables, whenever we want to predict an individual value of y for a new observation corresponding to a given value of x, we should use a(n):
determination interval.
estimation interval.
confidence interval.
prediction interval.
When studying the relationship between two quantitative variables, an interval estimate of the mean value of y for a given value of x is called a(n):
determination interval.
estimation interval.
confidence interval.
prediction interval.
When constructing a confidence or a prediction interval to quantify the relationship between two quantitative variables, what distribution do confidence and prediction intervals follow?
Uniform distribution
t distribution
Chi-Square distribution
Normal distribution
A ________ is a graph of the standardized residuals plotted against values of the normal scores. This helps to determine whether the assumption that the error term has a normal probability distribution appears to be valid.
normal probability plot.
regression plot.
scatter diagram.
residual plot.
Which of the following statements is false?
Regression analysis can be interpreted as a procedure for establishing a cause-and-effect relationship between variables.
In practice, parameter values are not known and must be estimated using sample data.
In the estimated simple linear regression equation, b0 is the y-intercept and b1 is the slope.
ŷ is the point estimator of E(y) , the mean value of y for a given value of x.
An observation that has a strong influence or effect on the regression results is called a(n):
residual
influential observation.
outlier
mistake
If the coefficient of determination is a positive value, then the coefficient of correlation:
must also be positive.
must be zero.
can be either negative or positive.
must be larger than 1.
When constructing a confidence or a prediction interval to quantify the relationship between two quantitative variables, the appropriate degrees of freedom are:
k - 1
(r-1)(c-1)
n - 1
n - 2
An F test, based on the F probability distribution, can be used to test for:
significance in regression.
significance in the relationship between two categorical variables.
equality of the means of two populations.
equality of two population proportions.
In a simple linear regression model, the error term ε accounts for the variability in ______ that cannot be explained by the linear relationship between x and y.
y
the mean
x
the standard deviation
