Netflix · Statistics & Data Analysis
Design and power a frequency-cap experiment
TrueInterview
October 7, 2026 · 2 min read
A product team is considering raising the per-user rolling 7-day frequency cap for a large video advertising campaign from 3 impressions to 4. Design an experiment, and provide power calculations that account for interference and clustering.
Context and requirements:
- Population: US users who are eligible for the campaign; roughly 4,000,000 eligible users per day are expected during the test.
- Randomization candidates: user_id, household_id, or geo cell; among eligible users the average household size is ; the household-level ICC for the primary metric is .
- Primary metric: 7-day conversion rate per unique exposed user, defined as any purchase within 7 days of first exposure; baseline .
- Guardrails: daily unique reach, average session watch time, and complaint rate per 1,000 impressions.
- Traffic allocation: 50% Treatment (cap = 4), 50% Control (cap = 3), planned duration 28 days, with 4 equally spaced interim looks including the final look.
- CUPED: a pre-period 7-day metric is available with to reduce variance.
- Interference risks: auctions shared across campaigns, overlapping advertisers, cross-device households, and pacing controls.
Tasks: (1) Select the randomization unit and defend it with a causal diagram: indicate where interference could arise and how your choice reduces it; if needed, propose cross-campaign holdouts or ghost bids. (2) Write precise metric definitions (numerators and denominators, exposure semantics, attribution window, cross-device de-duplication) and the data you would log so they can be computed unambiguously. (3) Compute the minimum per-arm sample size in unique users required to detect an absolute lift from 2.00% to 2.10% ( percentage points) with two-sided and using a two-proportion z-test. Adjust for clustering with , then adjust for CUPED by multiplying the variance by . Show the final effective sample size and discuss whether 28 days of traffic is sufficient. (4) Specify sequential monitoring with O’Brien–Fleming boundaries for 4 looks: give the approximate nominal at each look and describe the decision rules. (5) List at least three diagnostic checks (for example, covariate balance on pre-period exposures, saturation by user quantile, auction pressure) and the exact plots you would produce. Explain how you would interpret each one to decide whether to ship the higher cap.
Overview: This question tests experimental design and causal inference skills: power analysis, clustering and interference mitigation, metric engineering, sequential monitoring, and diagnostic interpretation for large-scale advertising experiments.