Amazon · Statistics & Data Analysis
Quantify improvement and compute required sample size
TrueInterview
October 7, 2026 · 1 min read
Suppose you assert that a new classifier lowers the spam rate from to , a relative decrease. For a two-arm randomized bucket experiment with significance level and power , using a two-sided z-test on proportions:
-
Derive the sample size per arm and present the formula you applied, making clear any pooled or unequal-variance assumptions.
-
If you see control at for 500k emails and treatment at for 500k, calculate the confidence interval for the absolute difference along with the corresponding p-value.
-
Discuss how class-imbalance drift and traffic seasonality can affect the test, and propose stratification or blocking to maintain validity.
-
If you monitor daily and stop as soon as , explain why the Type I error grows and outline a correction such as alpha spending or group-sequential boundaries.
-
If precision is critical for the business, explain how you would estimate a confidence interval for precision at a fixed recall using the delta method or bootstrap, and when each approach is appropriate.
Overview: This question measures a data scientist's skill in statistical experiment design and inference, spanning sample size calculation for proportions, two-sided hypothesis tests and confidence intervals, concerns from drift and seasonality, Type I error control under peeking, and precision estimation techniques like the delta method and bootstrap. It is often used to gauge knowledge of A/B test power and error control in the Statistics & Math domain, evaluating both conceptual comprehension and hands-on use of statistical methods.