Font size
WorksheetsAP Stats Exam Review Units 1 – 4
Total questions: 87
Worksheet time: 44mins
A(n) ________________ is the entire group of individuals we want information about, while a(n) ________________ is a subset of the population examined.
population, sample
sample, population
group, subset
individual, group
A(n) ______________ variable places an individual into one of several groups or categories.
categorical
continuous
dependent
quantitative
A(n) ______________ variable takes numerical values for which it makes sense to find an average.
quantitative
categorical
nominal
ordinal
Data from categorical variables are displayed with ______________ or ______________ charts.
bar, pie
line, scatter
histogram, box
area, bubble
Data from quantitative variables are displayed with ________________, ________ plots, or ________ plots.
histograms, dot, stem
bar, pie, scatter
box, line, area
frequency, pie, bar
When describing a quantitative distribution, you should always include ______________, ______________, ______________, ______________, and ______________.
shape, center, spread, odd features, context
mean, median, mode, range, variance
center, spread, frequency, probability, context
shape, median, mode, outliers, context
Choose the correct comparison of the mean and median for each shape given.
Skewed Left: Mean < Median;
Symmetric: Mean = Median;
Skewed right, Mean > Median
Skewed Left: Mean > Median; Symmetric: Mean < Median;
Skewed right: Mean = Median
Skewed Left: Mean = Median;
Symmetric: Mean > Median; Skewed right: Mean < Median
A(n) ____________________ is an individual value that falls outside the overall pattern of the distribution.
outlier
median
mode
mean
The ________________ is the average value, calculated by summing all values and dividing by the count. It is ________________ to outliers.
mean, resistant
median, resistant
mean, nonresistant
range, unaffected
The ____________________ is the midpoint of the distribution, with half the observations smaller and half larger. It is _______________ to outliers.
median; resistant
mean; nonresistant
median; nonresistant
range; nonresistant
The ____________________ describes the variation of the data set. It is the average distance from the mean. It is ________________ to outliers.
standard deviation, resistant
mean, nonresistant
median, resistant
standard deviation, nonresistant
The Interquartile Range (________) is the distance between the first quartile (________) and the third quartile (________). It is ____________________ to outliers.
IQR, Q1, Q3, resistant
IQR, Q2, Q4, sensitive
IQR, Q1, Q2, affected
IQR, Q2, Q3, vulnerable
The ____________________ summarizes a distribution using the Minimum, ________, Median, ________, and Maximum. These values create a boxplot.
five-number summary, Q1, Q3
mean summary, Q2, Q4
quartile summary, Q1, Q2
statistical summary, Q2, Q3
A rule for identifying outliers is any value falling outside the interval ________________ to ________________.
Q1 - 1.5(IQR), Q3 + 1.5(IQR)
Q1 - 2(IQR), Q3 + 2(IQR)
Q1 - 1(IQR), Q3 + 1(IQR)
Q1 - 0.5(IQR), Q3 + 0.5(IQR)
A ________________ measures how many standard deviations a value is from the mean.
z-score
mean
variance
median
Adding a constant, c, to every observation __________ the measures of center (mean, median, quartiles) by c but ________________ the measures of spread (range, IQR, standard deviation).
changes; does not change
decreases; increases
does not change; increases
increases; decreases
Multiplying every observation by a positive constant, b, ________________ both the measures of center and measures of spread by b.
multiplies
divides
subtracts
adds
The empirical rule states that _______% of data is within 1 standard deviation of the mean in a normal distribution, _______% is within 2 standard deviations, and _______% is within 3 standard deviations.
68; 95; 99.7
70; 90; 99
65; 92; 99.9
60; 85; 98
To find the percent of data lying within set boundaries in a normal curve, use ________________________ on the calculator.
normalcdf
normalpdf
invNorm
stdDev
To find the value that lies at a certain percentage in a normal curve, use ________________________ on the calculator.
invNorm
normalcdf
stdDev
mean
A ________________ displays the relationship between two ________________ variables.
scatterplot; quantitative
bar graph; categorical
pie chart; qualitative
histogram; discrete
When describing a scatterplot, you should discuss ____________, ________________, ________________, and ________________.
direction; shape; strength; outliers
mean; median; mode; range
slope; intercept; residuals; correlation
variance; standard deviation; skewness; kurtosis
The ________ variable is plotted on the horizontal (x) axis, and the ________ variable is plotted on the vertical (y) axis.
explanatory; response
response; explanatory
dependent; independent
independent; dependent
The ________________, denoted by r, measures the ____________ and ____________ of the linear relationship between two quantitative variables.
correlation coefficient; direction; strength
regression; magnitude; frequency
variance; spread; central tendency
mean; association; variability
The value of r must be between ________ and ________. A value close to 0 indicates a ________ linear relationship.
-1; 1; weak
0; 1; strong
-1; 0; strong
-1; 1; strong
____________________ does not imply ____________________.
Correlation; causation
Causation; correlation
Observation; inference
Prediction; certainty
The ________________ regression line is the line that minimizes the sum of the ________________ of the ________________ (vertical distances) from the data points to the line.
least-squares; squares; residuals
maximum-likelihood; cubes; errors
mean; absolute values; deviations
linear; products; differences
The general form of the least-squares regression line is ________________, where ŷ is the ________________ value of the response variable.
ŷ = a + bx; predicted
ŷ = bx + a; observed
ŷ = a - bx; calculated
ŷ = bx - a; measured
The ________ of the line represents the ____________ change in the ________ variable for every one-unit increase in the ____________ variable.
slope; predicted; response; explanatory
intercept; actual; explanatory; response
slope; actual; explanatory; response
intercept; predicted; response; explanatory
The ________ is the ________________ value of the ________ variable when the ________ variable is zero.
y-intercept; predicted; response; explanatory
slope; observed; explanatory; response
mean; actual; response; predictor
coefficient; estimated; predictor; response
In a computer output, the values needed for an LSRL are in the __________ column of numbers. The slope of the line is _______ and the y-intercept is _______.
first; b; a
first; a; b
Regression; m; c
Output; x; y
______________ is the use of a regression line to make ________ for x values that are far outside the range of the ____________ data. These predictions are often _____________.
Extrapolation; predictions; observed; unreliable
Interpolation; estimations; measured; reliable
Extrapolation; estimations; predicted; accurate
Interpolation; predictions; observed; reliable
A ________ is the difference between an observed value and the value predicted by the regression line.
residual
outlier
mean
variance
A(n) _______ plot is a scatterplot of the _______ against the _______ variable.
residual; residuals; explanatory
scatter; explanatory; residuals
line; residuals; response
trend; response; explanatory
A good fit for a linear model is indicated by a residual plot that shows ________ pattern.
no obvious
linear
curved
cyclical
The ________________, denoted by r² is the percent of the variation in the ________ variable that is accounted for by the ____________ with the ________ variable.
coefficient of determination; response; linear relationship; explanatory
correlation coefficient; explanatory; nonlinear relationship; response
regression coefficient; response; quadratic relationship; explanatory
standard deviation; explanatory; linear relationship; response
A(n) ________________ is an observation that lies outside the overall pattern of the other observations.
outlier
median
mode
range
A(n) _____________ point is an observation that, if removed, would significantly change the ____________ or ____________ of the regression line.
influential; slope; y-intercept
outlier; mean; median
critical; variance; standard deviation
leverage; correlation; residual
Points that are ____________ in the x direction relative to the rest of the data are ____________________ points.
outliers; high leverage
close; influential
far; outlier
close; outlier
The standard deviation of the residuals is ______. It measures the typical _______________________________.
s; residual
r; correlation strength
b; slope of the line
a; intercept value
A(n) ________________ attempts to collect data from every individual in the population.
census
survey
sample
estimate
A(n) ________________ is a value from a sample to used to estimate a population ___________________.
statistic; parameter
parameter; statistic
mean; median
sample; population
A(n) ________________ imposes a treatment on individuals to measure their responses.
experiment
survey
observation
census
A(n) ____________________________ observes individuals and measures variables without attempting to influence the responses.
observational study
experimental study
case-control study
survey
A(n) ________ is a method of choosing a sample where every individual and every possible group of size n has an equal chance of being selected.
simple random sample (SRS)
systematic sample
convenience sample
stratified sample
In _____________ sampling, the population is first divided into similar groups called ______________, and then a SRS is chosen from each group.
stratified; strata
cluster; clusters
systematic; systems
random; groups
In _____________ sampling, the population is first divided into representative groups called __________, and then all individuals from a randomly chosen subset of these groups are selected.
cluster; clusters
stratified; strata
systematic; systems
simple random; samples
____________ sampling selects individuals who are easiest to reach. This method often leads to ________________ bias, because not all of the population was eligible to be chosen.
Convenience; undercoverage
Random; response
Systematic; measurement
Stratified; nonresponse
______________ sampling allows individuals to choose to be in the sample by responding to a general invitation (e.g., an online poll). This often leads to ______________ bias.
Voluntary response; voluntary response
Random; selection
Stratified; measurement
Systematic; nonresponse
__________________ is the difference between the results of multiple samples taken from a population.
Sampling variability
Sampling bias
Measurement error
Nonresponse error
______________ is a systematic error in the design of the study that tends to favor certain outcomes.
Bias
Randomization
Sampling
Blinding
__________________ bias occurs when an individual chosen for the sample cannot be contacted or refuses to cooperate.
Nonresponse
Selection
Response
Measurement
__________________ bias occurs when respondents lie or give inaccurate answers, often due to the wording of the question or the identity of the interviewer.
Response
Selection
Sampling
Measurement
The four key principles of experimental design are ________________, ________________, ________________, and ________________.
Control, randomization, replication, comparison
Control, observation, prediction, analysis
Randomization, sampling, inference, estimation
Replication, correlation, causation, measurement
There are three steps needed to describe the random assignment of treatments in an experiment: ________________, ________________, ________________.
Label, random selection, assign
Observation, measurement, conclusion
Selection, grouping, analysis
Preparation, execution, evaluation
The individuals to whom the treatments are applied are called ______________ units. When they are human, they are called ______________.
Experimental; subjects
Control; patients
Sample; volunteers
Test; participants
A(n) ______________ is a condition applied to the experimental units. The treatments are formed by combining different levels of the ______________ variables (or factors).
Treatment; explanatory
Experiment; response
Variable; dependent
Factor; independent
In a(n) ______________ experiment, neither the subjects nor the people who interact with them and measure the response variable know which treatment a subject received.
Double-blind
Single-blind
Open-label
Placebo-controlled
In a(n) ______________ experiment, either the subjects or the people who interact with them know the treatment a subject received, but not both.
Single-blind
Double-blind
Randomized
Controlled
The ______________ effect occurs when subjects not receiving an active treatment show a response simply because they believe they are receiving a treatment.
Placebo
Nocebo
Hawthorne
Observer
In a(n) _____________________ experimental design, experimental units are randomly assigned to treatments.
Completely randomized
Matched pairs
Block
Factorial
In a(n) __________________ experimental design, experimental units are assigned to groups based on a common characteristic, then treatments are assigned within each group.
Randomized block
Completely randomized
Matched pairs
Factorial
In a(n) ____________ experimental design, experimental units are either matched up with one other unit or matched with themselves. Then treatments are randomly assigned.
Matched pairs
Completely randomized
Block
Factorial
Random assignment of treatments is important to show __________________.
Causation
Correlation
Generalization
Bias
Random sampling is important for _____________ results to a __________________.
Generalizing; population
Specifying; sample
Analyzing; variable
Comparing; group
The set of all possible outcomes of a chance process is called the _____________.
sample space
event
probability
experiment
A(n) ______________ is any collection of outcomes from some chance process.
event
sample
experiment
variable
The ______________ of any event must be a number between 0 and 1.
probability
mean
median
mode
The __________________________ states that if we observe more and more repetitions of a chance process, the proportion of times that a specific outcome occurs approaches a single value.
law of large numbers
central limit theorem
random variable principle
probability distribution rule
If two events have no outcomes in common, they are ____________________ (or disjoint). This means P(A ∩ B) = __________.
mutually exclusive, 0
independent, 1
complementary, A+B
dependent, AB
The probability that event B occurs given that event A has already occurred is called ________ ________ and is written as ___________.
conditional probability, P(B|A)
joint probability, P(A∩B)
marginal probability, P(B)
independent probability, P(A|B)
Two events A and B are ________________ if the occurrence of one event does not affect the probability that the other event occurs. This means P(A ∩ B) = ____________
independent, P(A) × P(B)
dependent, P(A) + P(B)
mutually exclusive, 0
complementary, 1
A ________________, or a ____________________ can be helpful tools for organizing and solving probability problems.
Venn diagram, two-way table
bar graph, pie chart,
scatter plot, box plot
dot plot, line graph
A ____________________ is a variable whose value is a numerical outcome of a chance process.
random variable
dependent variable
categorical variable
constant
A ____________ random variable X takes a fixed set of possible values with gaps between them.
discrete
continuous
uniform
normal
A ____________ random variable Y takes all values in an interval of numbers.
continuous
discrete
categorical
binary
The __________ value of a random variable is the mean of the outcomes, calculated as ______________
expected, Σx*p(x)
actual, simple addition
possible, mean
average, frequency count
When adding or subtracting two random variables, X and Y, the new ________ is always the sum or difference of the individual means.
mean
variance
standard deviation
median
When adding or subtracting two independent random variables, the new ________ is the square root of the sum of the______________.
standard deviation, variances
variance, standard deviation
variance, multiply
median, divide
The conditions for a ________________ distribution are ____________________, ____________________, ____________________, and ______________________________.
binomial, fixed number of trials, independent trials, two possible outcomes, constant probability of success
normal, continuous data, symmetric distribution, mean equals median, bell-shaped curve
poisson, rare events, independent events, constant average rate, discrete outcomes
uniform, equal probability, fixed range, independent outcomes, constant distribution
In a ____________ distribution, the variable of interest is the number of ______________ needed to get the ______________ success.
geometric, trials, first
binomial, successes, last
poisson, events, next
normal, samples, average
The mean (expected value) of a Binomial random variable is __________. The standard deviation of a Binomial random variable is ____________.
np, np(1−p)
np, np
np, n(1−p)
p, np(1−p)
The mean (expected value) of a Geometric random variable is ____________. The standard deviation of a Geometric random variable is ____________.
p1 , p1−p
p, p1−p
p1 , p1−p
p, p21−p
When calculating a binomial probability of an exact value, use binom _______. When calculating a binomial geometric probability of a ≤, use binom_______.
pdf, cdf
cdf, pdf
mean, variance
mode, median
What is the meaning of the term 'Categorical'?
A variable that can be divided into groups or categories that do not have a numerical value.
A variable that always represents a numerical value.
A variable that can only take continuous values.
A variable that is always used for mathematical calculations.
What does 'Standard Deviation' measure?
The typical distance from a mean.
It measures the average value in a set of data.
It measures the highest value in a set of data.
It measures the total sum of all values in a set of data.
Which term refers to the strength and direction of a linear relationship between two variables?
Correlation Coefficient
Mean
Standard Deviation
Coefficient of Determination
