Netflix · Statistics & Data Analysis
Plan and analyze a ranking A/B test
TrueInterview
October 7, 2026 · 1 min read
A search team puts forward a new ranking feature. You must design, run, and interpret the experiment:
(1) Randomization unit: pick between user, session, or query level, and defend the choice in light of carryover across sessions and potential network or interference effects.
(2) Metrics: specify the primary success measure (for instance query-level success rate, or paid conversion within 24 hours) along with guardrail metrics (latency, crash rate, ads revenue, bounce).
(3) Power and sample size: assume a baseline click-through rate of ; you want a relative uplift of (reaching ), with two-sided and power . Present the formula and calculate the per-variant sample size needed for a standard two-proportion z-test; then explain how clustering or CUPED would alter that number.
(4) Execution: describe SRM checks, triggered versus intent-to-treat analyses, consistent bucketing across services, burn-in for novelty effects, and sequential monitoring that does not inflate Type I error.
(5) Heterogeneity: suggest pre-registered segments (such as head versus tail queries, country, device) and how you would test for interaction while controlling false discovery.
(6) Interference and long-term effects: if ranking changes influence supply/demand dynamics, propose cluster randomization or switchback testing, and how to interpret the outcomes.
(7) Rollout: set stop/go criteria, a ramp plan, and how to update the ML training data so training is not entangled with experiment exposure.
Overview: This question assesses experimental-design and causal-inference skills for online A/B testing, spanning metric definition, randomization strategy under cross-session carryover and interference, power and sample-size computation, sequential monitoring, heterogeneity analysis, and safe rollout plus ML retraining concerns.