Stripe · Statistics & Data Analysis
Design and power an A/B test
TrueInterview
October 7, 2026 · 1 min read
You are preparing to roll out an email targeting model: the treatment arm consists of users whose score is above a threshold and who receive the email; the control arm is held back.
(1) Pick one primary success metric and two guardrail metrics (such as unsubscribe rate or complaint rate), and explain your choices.
(2) Given a baseline 7-day purchase rate of 5% and an expected relative lift of 8%, calculate the minimum sample size per arm for a two-sided test at and 80% power; show your formulas and assumptions (a continuity-corrected normal approximation is acceptable).
(3) Propose a ramp plan with sequential monitoring that keeps the type-I error rate under control (for example, group-sequential or alpha-spending methods); specify the interim looks and stopping rules.
(4) Describe the pre-experiment checks you would run (randomization, covariate balance, holdout contamination), and explain how you would address interference and weekend-related seasonality.
(5) If legal or traffic constraints make a pure A/B test impossible, propose a credible quasi-experimental design (such as regression discontinuity at the score threshold or staggered difference-in-differences), and list the assumptions you would test and the plots you would include in your slides.
Overview: This question assesses a data scientist's skills in experiment design: choosing metrics and guardrails, computing statistical power and sample size, planning sequential monitoring and ramp strategies, running operational checks for randomization and contamination, and selecting credible quasi-experimental alternatives.