WorksheetsBA 5
Total questions: 82
Worksheet time: 41mins
Name
Class
Date
1.
Which statement best captures the goal of running controlled experiments in a business setting?
a)
Guarantee the effect is positive for every individual user
b)
Pick the best-looking result among many metrics without correction
c)
Estimate causal impact of a change by comparing randomized treatment and control groups
d)
Eliminate uncertainty so confidence intervals are unnecessary
e)
Show that a metric moved after a launch, without accounting for confounding factors
2.
Which activity is explicitly presented as part of designing a trustworthy A/B test?
a)
Pre-define success criteria, metrics, and analysis plan before launching
b)
Assign treatment based on user choice to increase adoption
c)
Declare success if any tracked metric shows improvement
d)
Change the stopping rule whenever the p-value is close to 0.05
e)
Choose the primary metric after early results appear favorable
3.
Why do organizations run experiments?
a)
To eliminate the need for uncertainty estimates
b)
To guarantee every metric improves simultaneously
c)
To replace randomization with observational comparisons
d)
To measure causal impact rather than correlation
e)
To avoid defining success criteria in advance
4.
In the potential outcomes framework, what do Y(1) and Y(0) represent for a unit?
a)
Outcomes measured before and after data cleaning
b)
Outcomes for two different units in the same cluster
c)
The maximum and minimum possible outcomes for the unit
d)
Two different measurement instruments for the same outcome
e)
The outcome under treatment and the outcome under control
5.
What is a key benefit of random assignment in experiments?
a)
It removes the need to define hypotheses
b)
It makes p-values equal to the effect probability
c)
It guarantees no outliers occur in the outcome metric
d)
It ensures the treatment group is always larger than control
e)
It balances observed and unobserved factors on average
6.
Which activity is part of the experiment lifecycle described?
a)
Run sanity checks and detect sample ratio mismatch
b)
Train a predictive model to assign treatment deterministically
c)
Choose metrics only after seeing early results
d)
Remove the control group once significance is reached
e)
Skip analysis and ship whenever the mean increases
7.
What is the purpose of a guardrail metric in experiment metric design?
a)
To replace the primary metric when results are inconclusive
b)
To monitor unintended harm such as latency or error rate
c)
To guarantee the primary metric will improve
d)
To avoid collecting any diagnostic information
e)
To increase the number of metrics so p-values are smaller
8.
What is a primary use of an A/A test and related sanity checks?
a)
Detect instrumentation issues and sample ratio mismatch before shipping
b)
Increase conversion rate by exposing both groups to the new feature
c)
Identify the best user segment after the experiment ends
d)
Estimate the maximum possible treatment effect without a control
e)
Prove that randomization is unnecessary for causal claims
9.
What does Sample Ratio Mismatch (SRM) mean?
a)
The treatment effect changes sign across time periods
b)
The experiment uses multiple metrics instead of one
c)
The control metric has a higher mean than the treatment metric
d)
The outcome distribution is heavy-tailed
e)
The observed assignment ratio deviates from the planned design ratio
10.
Type I error is best described as:
a)
The difference between practical and statistical significance
b)
A false positive when the null hypothesis is actually true
c)
A failure to randomize users into groups
d)
The probability that the treatment effect is positive
e)
A false negative when the null hypothesis is actually false
11.
A p-value is best described as:
a)
The probability that the null hypothesis is true
b)
A measure of how extreme the data are under the null hypothesis
c)
A guarantee that the result will replicate
d)
The probability that the treatment has a positive effect
e)
The expected effect size in the population
12.
Which test is suggested for comparing conversion rates between two groups?
a)
Welch's t-test
b)
Mann-Whitney test on raw click counts only
c)
One-sample t-test on pooled users
d)
Paired t-test
e)
Two-proportion z-test
13.
MDE refers to:
a)
The maximum possible effect size that can be measured
b)
A method for correcting multiple comparisons
c)
A metric that must always decrease under treatment
d)
A practically significant minimum detectable effect used for design decisions
e)
A way to choose alpha after looking at the data
14.
Which input is listed as needed for power and sample size calculations for proportions?
a)
The names of the primary and guardrail metrics
b)
Baseline rate p0
c)
The calendar date when the experiment is launched
d)
The number of teams running experiments in parallel
e)
The random seed used for assignment
15.
A power curve for a fixed baseline, alpha, and MDE primarily shows how power changes with:
a)
Sample size per arm
b)
The randomization ratio chosen after observing outcomes
c)
The order in which users arrive during the test
d)
The experiment owner's confidence in the change
e)
The number of secondary metrics being tracked
16.
Naive peeking at results during an experiment tends to:
a)
Eliminate multiple testing concerns
b)
Reduce the need for randomization
c)
Guarantee higher power
d)
Inflate Type I error
e)
Make confidence intervals always narrower
17.
Which procedure is used to control the false discovery rate (FDR)?
a)
Bonferroni
b)
Stratified randomization
c)
CUPED adjustment
d)
Welch correction
e)
Benjamini-Hochberg
18.
When an outcome like revenue has heavy tails, a recommended approach is:
a)
Winsorizing or using robust statistics
b)
Switching to a one-sided test automatically
c)
Using only the median and never reporting uncertainty
d)
Dropping the control group to reduce variance
e)
Ignoring outliers because randomization fixes them
19.
What is a key benefit of variance reduction techniques?
a)
Ability to compute individual causal effects for every unit
b)
Automatic correction for multiple testing across many metrics
c)
Guaranteed elimination of selection bias without randomization
d)
Removal of the need to report confidence intervals
e)
Less noise, which can reduce required sample size or shorten test duration
20.
CUPED reduces variance by using:
a)
A smaller sample size and shorter duration by default
b)
A larger alpha level so more results are significant
c)
A pre-experiment covariate that is highly correlated with the outcome
d)
A post-treatment variable that changes only in the treatment group
e)
A different random seed chosen after seeing results
21.
In CUPED, theta is estimated as:
a)
A fixed constant chosen to match the target power
b)
Covariance of Y and X divided by variance of X
c)
Mean of X divided by mean of Y
d)
Variance of Y divided by covariance of Y and X
e)
Correlation of Y and X multiplied by sample size
22.
In CUPED, the adjusted metric can be written in words as:
a)
X minus theta times (Y minus the mean of Y)
b)
Y divided by X with a pooled denominator
c)
Y minus theta times (X minus the mean of X)
d)
The p-value transformed into an adjusted outcome
e)
Y plus theta times (X plus the mean of X)
23.
After forming the CUPED adjusted metric, the ATE is estimated by:
a)
The ratio of mean outcomes between treatment and control
b)
The correlation between treatment assignment and the covariate
c)
The difference in variances between treatment and control
d)
The pooled mean of both groups combined
e)
The difference in mean adjusted outcomes between treatment and control
24.
The approximate variance factor from CUPED is:
a)
One minus rho squared
b)
One plus rho squared
c)
Rho divided by sample size
d)
Alpha times beta
e)
The pooled proportion under the null hypothesis
25.
A key requirement for the CUPED covariate X is that it must:
a)
Be measured in a pre-period or otherwise not affected by treatment
b)
Have a perfectly nonlinear relationship with the outcome
c)
Depend on the p-value distribution under the alternative
d)
Be computed separately in each arm using post-assignment outcomes
e)
Be strongly affected by treatment to amplify the effect size
26.
Which step is included in the CUPED recipe described?
a)
Replace the primary metric with a proxy to make it easier to move
b)
Change the randomization ratio until the observed ratio matches the design
c)
Stop the experiment as soon as the first metric crosses 0.05
d)
Compute p-values by assuming the alternative hypothesis is true
e)
Run your usual two-sample test on the adjusted metric and report effect with uncertainty
27.
In the synthetic CUPED setup described, the post-period metric Y is modeled as depending on:
a)
The number of metrics being monitored
b)
Only random noise with no systematic component
c)
Only the treatment assignment with no other drivers
d)
A pre-period covariate X plus a true treatment lift
e)
The chosen multiple-testing correction method
28.
Compared with an unadjusted metric, a CUPED-adjusted metric typically has:
a)
Higher standard error (lower precision)
b)
No need for confidence intervals
c)
A guaranteed smaller point estimate
d)
Lower standard error (higher precision)
e)
A guaranteed larger point estimate
29.
What does a smaller standard error mainly imply for an effect estimate?
a)
A larger alpha level by default
b)
A change from two-sided to one-sided testing
c)
A higher probability that the null hypothesis is true
d)
A tighter confidence interval for the same confidence level
e)
A guaranteed elimination of all outliers
30.
Which OLS model is suggested as an alternative way to implement CUPED?
a)
Regress outcome on p-values; the p-value coefficient is the effect
b)
Regress the covariate on the outcome only; the slope is the ATE
c)
Regress outcome on experiment duration only; the duration coefficient is the effect
d)
Regress treatment indicator on the outcome; the outcome coefficient is the effect
e)
Regress outcome on treatment indicator and the covariate; the treatment coefficient is the effect
31.
Under the null hypothesis, p-values should be approximately:
a)
Always close to 1
b)
Always close to 0
c)
Equal to the effect size
d)
Exactly equal to alpha
e)
Uniform between 0 and 1
32.
Under an alternative where a real effect exists, p-values tend to be:
a)
Uniform between 0 and 1
b)
Skewed toward 1
c)
Fixed at the desired power level
d)
Skewed toward 0
e)
Centered near 0.5
33.
When is cluster randomization recommended?
a)
When you never plan to peek at interim results
b)
When the assignment ratio is 50/50
c)
When the outcome metric is always normally distributed
d)
When interference within clusters is likely
e)
When there is exactly one metric in the experiment
34.
Which diagnostic is explicitly recommended to include when reporting experiment results?
a)
Ignoring outliers because randomization fixes them
b)
Reporting only a point estimate with no uncertainty
c)
Declaring success if any metric becomes significant
d)
Selecting alpha only after seeing the final p-value
e)
Sample ratio mismatch (SRM) check
35.
In the email promotion mini-case, what is the primary metric?
a)
Email open rate only
b)
Click-through rate (CTR)
c)
Email sending throughput
d)
Unsubscribe rate
e)
Bounce rate
36.
In the mini-case analysis, which step is recommended?
a)
Decide to ship only if every tracked metric improves
b)
Stop the analysis once a single segment shows significance
c)
Check practical significance versus the MDE in addition to the p-value
d)
Assume causality without randomization because the goal is clear
e)
Ignore guardrails because the primary metric matters most
37.
Which hands-on task is listed ?
a)
Compute MDE using post-treatment outcomes only
b)
Run a simulation to validate Type I error at a chosen alpha level
c)
Choose alpha after peeking so the test stays significant
d)
Remove the control group to increase sensitivity
e)
Balance groups by outcome to guarantee comparability
38.
Which question is listed in the Quick Quiz section of the slides?
a)
What is the delta method used for in ratio metrics?
b)
Why do guardrail metrics matter in experimentation?
c)
When should you winsorize a heavy-tailed metric?
d)
Why is peeking problematic?
e)
How do you compute theta for CUPED in practice?
39.
Randomized tests primarily help convert:
a)
Variance into bias
b)
A guardrail metric into a primary metric
c)
Power into alpha
d)
Correlation into causation
e)
A p-value into a confidence interval
40.
Which resource is recommended for learning more about online controlled experiments and practical experimentation pitfalls?
a)
Causal Inference in Statistics: A Primer (Pearl et al.)
b)
Causal Inference: The Mixtape (Cunningham)
c)
Design and Analysis of Experiments (Montgomery)
d)
The Effect: An Introduction to Research Design and Causality (Huntington-Klein)
e)
Trustworthy Online Controlled Experiments (Kohavi et al.)
41.
For the two-proportion z-test statistic, the standard error under the null uses:
a)
A variance estimate taken from the confidence interval
b)
An unpooled proportion computed separately in each arm
c)
Only the treatment group's variance
d)
Only the control group's variance
e)
A pooled proportion (pooled variance under H0)
42.
In the appendix, power for proportions is described as depending on:
a)
Only the final observed p-value
b)
Only the number of guardrail metrics
c)
Only the experiment duration in days
d)
Baseline and alternative rates, alpha, and sample size
e)
Only the randomization ratio
43.
The provided two-proportion z-test function returns a:
a)
Confidence interval endpoints only
b)
Two-sided p-value
c)
Minimum detectable effect only
d)
Sample size per arm only
e)
One-sided p-value for increase only
44.
In the provided CUPED adjustment code, theta is computed using:
a)
Sample variance of y divided by sample covariance of y and x
b)
Sample covariance of y and x divided by sample variance of x
c)
Pooled proportion under the null hypothesis
d)
A fixed constant chosen before collecting any data
e)
Difference in means of y between treatment and control
45.
In the potential outcomes view, why is it impossible to compute an individual unit's causal effect directly?
a)
Only one of Y(1) or Y(0) is observed for each unit
b)
Randomization forces Y(1) to equal Y(0) for each unit
c)
Confidence intervals replace the need for causal effects
d)
The ATE is defined only for clusters, not individuals
e)
P-values are only defined for observational studies
46.
Random assignment primarily helps causal estimation by reducing which problem?
a)
The need to define success criteria
b)
The presence of any outliers in the data
c)
Confounding from differences between treatment and control groups
d)
Multiple testing across many metrics
e)
Heavy tails in the outcome distribution
47.
Which pairing best matches an experiment lifecycle stage to its described activity?
a)
Define - apply a multiple testing correction
b)
Run - perform sanity checks and detect SRM
c)
Design - compute p-values under the alternative
d)
Learn - ignore guardrails to focus on KPIs only
e)
Analyze - rebalance assignments to fix SRM
48.
Which metric setup best follows guidance on primary and guardrail metrics?
a)
Primary: p-value; Guardrail: confidence interval width
b)
Primary: experiment duration; Guardrail: number of metrics tracked
c)
Primary: conversion rate; Guardrail: latency or error rate
d)
Primary: number of users; Guardrail: conversion rate
e)
Primary: random seed; Guardrail: assignment ratio
49.
During an A/A test, you detect a sample ratio mismatch. What is the most appropriate next step?
a)
Ignore SRM if the primary metric looks stable
b)
Ship the feature because A/A tests cannot have issues
c)
Increase alpha so SRM is less likely to be significant
d)
Stop and investigate because results are not trustworthy
e)
Switch to observational analysis to avoid randomization
50.
You want to reduce false negatives (Type II error) while keeping alpha fixed. Which change most directly supports this goal?
a)
Remove guardrails to simplify the decision
b)
Increase sample size to increase power
c)
Increase the number of secondary metrics
d)
Decrease sample size to reduce variance
e)
Peek more frequently to stop early on success
51.
Which statement is an incorrect interpretation of a p-value?
a)
It is commonly reported with confidence intervals
b)
It is the probability that the null hypothesis is true
c)
It should be interpreted alongside an effect estimate
d)
It is not the probability that an effect exists
e)
It reflects how extreme the data are under the null hypothesis
52.
You are comparing mean revenue across groups and expect unequal variances. Which test choice is most appropriate?
a)
Chi-square test on pooled revenue
b)
Two-proportion z-test
c)
Paired t-test (same users in both groups)
d)
Welch's t-test
e)
One-sample t-test on the combined sample
53.
If the outcome distribution is unusual or highly non-normal, which two-sample option is mentioned?
a)
Welch's t-test
b)
Chi-square test for proportions
c)
Two-proportion z-test
d)
Paired t-test
e)
Mann-Whitney test
54.
An experiment shows p < alpha, but the estimated lift is smaller than the pre-defined MDE. What conclusion best matches recommended practice?
a)
It proves the treatment has a large business impact
b)
It is invalid because MDE and p-value cannot both be used
c)
It must be practically significant because p < alpha
d)
It implies the experiment should always be stopped early
e)
It may be statistically significant but not practically significant
55.
Holding baseline rate, alpha, and desired power fixed, what happens to required sample size when you target a smaller MDE?
a)
It stays the same by definition
b)
It becomes independent of the baseline rate
c)
It increases
d)
It decreases
e)
It becomes independent of noise
56.
You need interim looks during an experiment. Which approach avoids inflating Type I error?
a)
Use group-sequential or alpha-spending methods
b)
Increase the number of metrics and report the smallest p-value
c)
Stop the first time the effect estimate is positive
d)
Randomize only after observing early outcomes
e)
Run the standard test repeatedly without adjustment
57.
Which mapping of method to error-rate goal matches the slides?
a)
Bonferroni controls FDR; Benjamini-Hochberg controls FWER
b)
Neither method addresses false positives; they only reduce variance
c)
Bonferroni controls FWER; Benjamini-Hochberg controls FDR
d)
Both methods are designed to control FWER
e)
Both methods are designed to control FDR
58.
For heavy-tailed metrics like revenue or time-on-site, which mitigation is explicitly suggested?
a)
Always switch to one-sided p-values
b)
Ignore outliers because randomization removes them
c)
Replace confidence intervals with point estimates only
d)
Use only ratio metrics to stabilize variance
e)
Winsorize or use robust statistics
59.
For ratio metrics such as ARPU, which approaches are mentioned for inference?
a)
Post-stratification or propensity score matching
b)
Stratified randomization or cluster randomization
c)
CUPED or Bonferroni correction
d)
Delta method or Fieller's theorem
e)
Welch's t-test or Mann-Whitney test
60.
If you successfully reduce variance (noise) in your metric while keeping the same MDE and alpha, what is a likely design implication?
a)
You must increase sample size to keep results valid
b)
You must switch from randomization to observational analysis
c)
You must stop reporting confidence intervals
d)
You can achieve the same power with fewer samples
e)
You must increase the number of guardrail metrics
61.
In CUPED, which choice of theta minimizes the variance of the adjusted metric under the stated linear setup?
a)
theta = mean(Y) / mean(X)
b)
theta = Var(Y) / Cov(Y, X)
c)
theta = Corr(Y, X) * Var(X)
d)
theta = Cov(Y, X) / Var(X)
e)
theta = alpha / beta
62.
Why is it recommended to estimate CUPED theta on pre-assignment data (or a training split)?
a)
To make the p-values uniform under the alternative
b)
To ensure the assignment ratio is exactly 50/50
c)
To guarantee the treatment effect will increase
d)
To avoid multiple testing corrections
e)
To avoid leakage from treatment assignment into the adjustment
63.
If the pre-period covariate X is strongly correlated with outcome Y, what does the variance factor imply?
a)
Variance decreases roughly with (1 - rho^2), improving effective sample size
b)
Variance increases roughly with (1 + rho^2), reducing effective sample size
c)
Variance reduction requires changing the assignment ratio
d)
Variance is determined only by alpha and beta
e)
Variance becomes independent of rho, so no change occurs
64.
If the Y-X relationship is strongly nonlinear, which extension is suggested?
a)
A/A testing
b)
Two-proportion z-test
c)
CUPAC or ML-CUPED
d)
Sequential monitoring without adjustment
e)
Bonferroni correction
65.
When experiments are clustered (for example, assigned by group), which uncertainty practice is advised?
a)
Assume independence and use the same standard error formula
b)
Compute theta separately in each arm using post-period outcomes
c)
Avoid reporting diagnostics to keep analysis simple
d)
Replace confidence intervals with p-values only
e)
Use cluster-robust standard errors and confidence intervals
66.
Which reporting package best matches recommended practice?
a)
Only secondary metrics because they are more informative
b)
Only a single p-value with no effect size
c)
Effect with confidence interval plus diagnostics and next steps
d)
Only the effect size with no uncertainty
e)
Only a power curve summary with no text
67.
In the email promo mini-case, which items are most clearly guardrails rather than the primary KPI?
a)
The p-value distribution under H0
b)
Bounce rate and unsubscribe rate
c)
Click-through rate only
d)
The randomization ratio only
e)
The confidence level used for intervals
68.
You randomize users individually, but users within households influence each other's behavior. Which problem and design adjustment best match the slides?
a)
Heavy tails can violate SUTVA; consider winsorization at the household level
b)
P-values are invalid for clusters; consider reporting only point estimates
c)
Multiple testing causes interference; consider Bonferroni correction only
d)
SRM is unavoidable; consider switching to observational analysis
e)
Interference can violate SUTVA; consider cluster randomization at the household level
69.
An experiment team checks results daily and stops the first time p < 0.05 without any correction. What is the key statistical risk emphasized ?
a)
The p-values become uniform under the alternative
b)
The SRM check is no longer needed
c)
The false negative rate becomes zero
d)
The MDE automatically becomes smaller
e)
The false positive rate is inflated beyond the nominal alpha
70.
You detect an SRM and still observe a large lift on the primary metric. Which statement is most consistent with the slides?
a)
Do not trust the effect estimate until the SRM cause is understood and fixed
b)
The lift is automatically causal because it is large
c)
SRM only affects secondary metrics, not the primary metric
d)
SRM can be ignored if the confidence interval is narrow
e)
SRM is a sign of higher power, so shipping is safer
71.
You track many metrics and plan to declare success if any metric is significant. If your priority is to control the chance of at least one false positive, which approach best fits the slides?
a)
Control FWER, for example with a Bonferroni adjustment
b)
Use CUPED to eliminate the need for corrections
c)
Report only the smallest p-value across metrics without context
d)
Control only power by increasing alpha
e)
Use a nonparametric test for every metric regardless of type
72.
You care about limiting the expected proportion of false discoveries among the significant findings. Which approach best fits the slides?
a)
Remove guardrails to reduce multiplicity
b)
Avoid pre-registering endpoints to keep flexibility
c)
Peek frequently and stop when any metric crosses 0.05
d)
Control FDR, for example with Benjamini-Hochberg
e)
Control FWER, for example with Bonferroni
73.
A team proposes using CUPED with a covariate measured after users have been exposed to the treatment. What is the most important concern based on the slides?
a)
The p-values under H0 will no longer be uniform
b)
The MDE becomes undefined for adjusted metrics
c)
The adjustment can become biased because the covariate may be affected by treatment
d)
The adjustment will always increase power regardless of correlation
e)
The experiment will automatically develop SRM
74.
A team estimates CUPED theta using data after assignment and then applies different theta values in treatment and control. Which guidance is being violated?
a)
Winsorize the outcome before computing any covariance
b)
Estimate theta on pre-assignment data and apply the adjustment consistently across arms to avoid leakage
c)
Use only ratio metrics to stabilize theta
d)
Use Bonferroni to correct the theta estimate
e)
Replace randomization with stratification to eliminate theta
75.
In a clustered experiment, an analyst uses standard errors that assume independent observations. What is a likely consequence and the appropriate fix?
a)
MDE becomes smaller automatically; stop early
b)
Confidence intervals can be too narrow; use cluster-robust standard errors
c)
Confidence intervals are always too wide; use one-sided tests
d)
SRM is guaranteed; rebalance assignments mid-test
e)
Type I error becomes zero; use more metrics to compensate
76.
If X is weakly correlated with Y, which statement best follows from the variance reduction discussion?
a)
CUPED eliminates the need to choose an MDE
b)
CUPED becomes biased even under randomization
c)
CUPED provides little variance reduction because the variance factor is close to 1
d)
CUPED forces p-values under H1 to be uniform
e)
CUPED guarantees a large effective sample size gain
77.
Which situation most strongly suggests moving beyond basic linear CUPED to CUPAC or ML-CUPED?
a)
The p-value is larger than 0.05
b)
The experiment has only one metric
c)
The relationship between the covariate and outcome is strongly nonlinear
d)
The assignment ratio is exactly 50/50
e)
The baseline conversion rate is above 10 percent
78.
You run many A/A tests and find that p-values are frequently very close to zero. Which explanation best matches recommended diagnostics?
a)
It proves the alternative hypothesis is always true
b)
This is expected behavior under H0, so no action is needed
c)
It implies that confidence intervals should not be reported
d)
It means that Bonferroni always increases power
e)
There may be instrumentation issues or SRM, since under H0 p-values should be roughly uniform
79.
At a 5 percent significance level for a two-sided test, which relationship between the p-value and the confidence interval is most consistent with standard definitions?
a)
If the 95% confidence interval excludes zero effect, the two-sided p-value is above 0.50
b)
The p-value equals the confidence interval width divided by sample size
c)
If the 95% confidence interval includes zero effect, the two-sided p-value must be below 0.01
d)
If the 95% confidence interval excludes zero effect, the two-sided p-value is below 0.05
e)
The p-value and confidence interval are unrelated by definition
80.
With a very large sample size, a tiny lift can be statistically significant. Which practice most directly protects against overreacting to such results?
a)
Track more metrics so at least one is large
b)
Stop reporting confidence intervals to reduce confusion
c)
Define a practically significant MDE and success criteria in advance, not just a p-value threshold
d)
Avoid guardrails so the primary metric dominates the decision
e)
Increase alpha until the result stops being significant
81.
A primary KPI improves, but a guardrail metric shows clear harm. What is the most defensible action under the guardrail guidance?
a)
Declare success if any secondary metric improves
b)
Ignore the guardrail because it is not linked to business goals
c)
Ship immediately because the primary metric is always decisive
d)
Hold or iterate rather than ship until guardrail impact is acceptable
e)
Increase interim looks to confirm the KPI effect sooner
82.
When sample sizes are small for a proportions test, which approach is suggested as more appropriate than relying only on normal approximations?
a)
Always use Mann-Whitney because it is nonparametric
b)
Use exact methods for small samples
c)
Always use CUPED because it removes small-sample issues
d)
Always increase alpha because small samples reduce power
e)
Always compute p-values under H1 instead of under H0
100 %
