Coinbase · Statistics & Data Analysis
Diagnose uplift drop in email A/B tests
TrueInterview
October 7, 2026 · 2 min read
An e-commerce company is running a test of personalized product emails to lift 7-day purchase conversion. Design the experiment, then debug the conflicting results from a rerun.
Part A — Design and sizing
-
Define one precise primary metric and 2–3 guardrails. Assume user-level randomization with intent-to-treat. Specify exposure and eligibility rules, and explain how to handle users who receive multiple emails.
-
Sample size: baseline 7-day purchase conversion is 3.5%. The test must detect a 10% relative lift with two-sided α=0.05 and power=0.80 under 1:1 allocation. Given 500,000 eligible users per day and 85% deliverability, how many calendar days are needed, including a full 7-day attribution window? Show the formula and the numeric answer.
-
Now suppose the business wants to power for 7-day revenue per randomized user, with mean $0.90 and SD $12.00. It wants to detect a +$0.10 absolute lift at the same α and power. What per-arm sample size and run length does that require?
Part B — Conflicting results and diagnostics
The first RCT ran from 2025-06-01 to 2025-06-14 with n=1,200,000 per arm. Control conversion was 3.50%, treatment was 4.20% (+20.0% relative, +0.70 pp). A rerun from 2025-08-15 to 2025-08-28 with n=900,000 per arm observed control 3.50% and treatment 3.57% (+2.0% relative, +0.07 pp).
-
For each test, calculate the two-proportion z-test p-value and a 95% CI for the absolute lift; then calculate a fixed-effects meta-analytic pooled lift across the two tests. Should you launch? Why?
-
List no fewer than six plausible explanations for the conflicting results (e.g., seasonality, targeting drift, novelty/creative fatigue, regression to the mean/winner’s curse, instrumentation/attribution changes, concurrency with promos, contamination, different triggered eligibility, population mix-shift). For each one, give 1–2 concrete checks (SQL or plots) to run and the exact data required.
-
Propose a re-analysis plan: pre-registration, CUPED or pre-period covariate adjustment, heterogeneity of treatment effects by region/device/recency, sequential monitoring corrections, and a holdout strategy for ramp. Describe the decisions you would make if the pooled lift is between +0% and +5%.
Overview:
This question assesses a data scientist's ability in experimental design, metric definition and guardrail selection, power and sample-size calculations, statistical inference (including two-proportion testing and fixed-effects meta-analysis), and debugging inconsistent A/B test reruns through instrumentation, population-shift, and heterogeneity checks. It is often asked because interviewers need to evaluate whether candidates can operationalize randomized email experiments, set run lengths and attribution windows, and diagnose conflicting results with applied analytics; the problem belongs to the Analytics & Experimentation domain and tests practical application built on conceptual statistical understanding.