Meta · Statistics & Data Analysis
Design and analyze an A/B test
TrueInterview
October 7, 2026 · 2 min read
A marketplace intends to alter how it ranks search results so that merchants located close to the user are favored. You are asked to run a 14-day randomized A/B test split 50/50 at the user level, with roughly daily active users and a baseline conversion rate of . The anticipated relative improvement in conversion is ; the guardrails are that the cancellation rate may rise by no more than percentage points and that average delivery time may not grow by more than minutes.
(a) Determine the smallest sample size per arm needed for the conversion metric under a two-sided z-test with and power . Present the formulas and the numerical steps, treating the trials as independent Bernoulli draws. Then explain how clustering of users and repeated sessions or orders would cause variance to be larger than assumed, and how you would adjust for that (for instance, a variance inflation factor derived from an empirical design effect, or cluster-robust standard errors).
(b) Lay out the primary, secondary, and guardrail metrics; pre-register the hypotheses; and state the decision rule that combines the size of the effect with statistical significance (including the minimum detectable effect and non-inferiority thresholds for the guardrails).
(c) Describe a CUPED adjustment or a pre-period covariate adjustment that uses each user's 28-day pretest conversion propensity. Give the adjusted estimator and explain how you would confirm that the variance reduction is real (an A/A test and placebo checks).
(d) Sketch a ramp schedule (1% → 10% → 50% → 100%), how you would watch for novelty and learning effects, how weekday and seasonal patterns would be controlled, and how you would deal with attribution and tracking lag (events arriving 48 hours late).
(e) Suppose the test is split by geography (at the city level) rather than by user. Propose a difference-in-differences design with city fixed effects and calendar effects; list the assumptions and describe how you would check that pre-trends are balanced.
(f) Explain how you will watch for and correct peeking and the use of many metrics (alpha-spending, O'Brien–Fleming; FDR when there are many guardrails), and define a rollback plan.
Overview: This question assesses whether a candidate can design and analyze randomized experiments. It spans statistical power and sample-size calculation, cluster-robust variance adjustments, covariate adjustment (CUPED), specification of hypotheses and guardrails, ramping and monitoring strategies, and causal inference approaches such as geo-level difference-in-differences. It sits in the Analytics & Experimentation domain for data scientist roles. It is often used to gauge practical application of experimental statistics and operational decision-making under real-world constraints—weighing statistical rigor, control of multiple testing, and complications like repeat users, seasonality, and delayed attribution—and it requires both a conceptual grasp of causal inference and hands-on analytical execution.