Pinterest · Statistics & Data Analysis
Design rigorous A/B test and causal analysis
TrueInterview
October 7, 2026 · 2 min read
Answer every part with formulas, numeric results, and stated assumptions:
A) Sample size: With baseline conversion , a target minimum detectable effect of +7% relative (so ), two-sided , and power , compute the required sample size per variant for a standard two-proportion z-test. Show the z-scores used and the pooled variance assumption.
B) Duration: Given 1.2M daily visitors, a 60/40 A/B traffic split, and 80% eligibility, how many calendar days are needed to reach the sample size from part (A)? State any adjustments for repeat visitors and overlap with other experiments.
C) Variance reduction: If a pre-experiment covariate has with the outcome, quantify the effective MDE or sample-size reduction from using CUPED. Explain when CUPED increases bias (for example, covariate shift).
D) Sequential testing: You plan daily peeks for 21 days. Propose an alpha-spending or group-sequential design (such as Pocock or O’Brien-Fleming). Specify the spending function and the final critical z. Explain the pros and cons relative to always-valid sequential methods (SPRT/e-values).
E) Interference and clustering: When randomizing by user causes cross-unit spillovers, propose a cluster design (for example, geo or traffic-bucket). Compute the design effect for ICC with average cluster size and . How does this change the sample size?
F) SRM check: On day 3 you observe 110,000 users in A and 90,000 in B (expected 60/40 from 200,000 eligible). Perform a chi-square goodness-of-fit test and report the p-value. What actions do you take if SRM is significant?
G) Causal inference: The team ran an observational study with a strong pre-period trend. Sketch a DAG, choose an identification strategy (DID, IV, or RDD), list required assumptions (for example, exclusion restriction for IV; continuity for RDD), and propose concrete robustness checks (placebo tests, pre-trend tests, sensitivity to unobserved confounding).
Overview: This question evaluates a data scientist's competency in experimental design, sample-size and power calculations, variance-reduction methods (e.g., CUPED), sequential testing and alpha spending, clustering and interference effects, SRM checks, and causal identification strategies such as DID, IV, and RDD within the Statistics & Math domain.