Disney · Statistics & Data Analysis
Compute sample size and plan experiment
TrueInterview
October 7, 2026 · 2 min read
A product team plans to run an A/B test on a paywall copy change aimed at new signups, with the goal of raising the next-day subscription start rate.
Given:
- Baseline next-day subscription start rate for new signups: 18%.
- Minimum detectable effect: +7% relative, meaning the rate would rise to 19.26%.
- Two-sided significance level , power , equal allocation.
- Optionally apply CUPED with using pre-experiment engagement.
- If randomization is at the household level, the average household size is signups per household and the intraclass correlation is .
- You may plan up to four interim looks using O'Brien–Fleming alpha spending.
- Twelve secondary metrics are tracked; control the false discovery rate at 10% with the Benjamini–Hochberg procedure.
- 10% of users assigned to treatment will not actually see the new copy because of noncompliance, while 3% of control users may be exposed due to caching.
Questions:
- Compute the per-arm sample size without variance reduction or clustering, and show the formulas and approximations you use.
- Given CUPED with , what reduction in effective sample size does it provide? Recompute the required per-arm sample.
- Adjust for household clustering using the design effect . Recompute the per-arm sample size under clustering, both with and without CUPED.
- Explain how O'Brien–Fleming boundaries change the way Type I error is allocated, and what the practical implications are for timeline and power.
- Describe how you would keep the FDR at 10% across the 12 secondary metrics, and how you would interpret any discoveries.
- Compute the ITT and CACE effects using the given noncompliance rates, assuming monotonicity. How would you report both results responsibly to product stakeholders?
Overview: This question tests a candidate's skill in experimental design and applied statistics, covering sample size and power calculations, variance reduction through CUPED, clustering and design-effect adjustments, interim analysis with O'Brien–Fleming alpha spending, multiple-comparison control via Benjamini–Hochberg, and causal estimands including ITT versus CACE. It is frequently asked because interviewers want confidence that a candidate can turn business treatment goals into a rigorous experiment plan that balances Type I and Type II error, multiplicity, noncompliance, and operational constraints; the question sits in the Statistics & Math domain and stresses practical application alongside conceptual understanding.