Capital One · Statistics & Data Analysis
Design and analyze ad A/B test
TrueInterview
October 7, 2026 · 1 min read
You are comparing a new ad-ranking algorithm (B) with the current production system (A) on an online video platform. The primary metric is mean watch_time per impression, measured in seconds. Guardrail metrics are: (1) error rate at most 1%, (2) ad load, defined as ads per session, must not rise, and (3) click-through rate (CTR) must not fall by more than 2% relative to the control. Traffic is split 50/50 between A and B, with user-level randomization, and the test runs for 14 days. There is weekly seasonality and a known weekday-versus-weekend effect. Daily data are available for impressions, total_watch_time_sec, clicks, sessions, and errors, broken down by variant and platform (Web, Mobile). Design the experiment analysis plan: (a) State and and justify whether a one-tailed or two-tailed test is appropriate for the primary metric; (b) Specify the exact test for the primary metric (for example, a two-sample t-test on per-user means, CUPED with a covariate, or a cluster-robust approach) and justify the assumptions and clustering; (c) Define the variance reduction strategy you would use (for example, CUPED using pre-experiment watch_time) and how you would compute it; (d) Show how you will check guardrails with multiplicity control (for example, Holm-Bonferroni), and what decision rule you will use if a guardrail is violated; (e) Describe the stratification/segmentation you will pre-register (for example, by platform and weekday/weekend) and how you will combine strata (fixed vs random effects meta-analysis); (f) Provide a power/MDE calculation sketch assuming baseline mean = 70s, sd = 25s at user-day level, average 4 impressions/user-day, intra-user correlation 0.35, 200k users per arm over 14 days; (g) Explain how you would diagnose and mitigate traffic imbalance or novelty effects.
Overview: This question assesses a candidate's ability in online experimentation and statistical analysis, including hypothesis formulation, variance reduction, clustering and stratification, multiplicity control, power/MDE calculation, and operational metric guardrails.