Netflix · Statistics & Data Analysis
Estimate ATE of personalization on streaming
TrueInterview
October 7, 2026 · 1 min read
You receive a user-level dataset from an online experiment where personalization was randomly assigned as the treatment and no personalization as the control. Suppose each user appears as a single row with these columns:
user_id(string or integer)treat(0 or 1): the randomized indicator for personalization assignmentminutes_streamed(float): total minutes streamed in the 7-day window after assignment- Optional pre-treatment covariates (some may be irrelevant or noisy): for example,
country,device_type,tenure_days,prior_7d_minutes,is_premium, and so on.
Tasks:
- Estimate the Average Treatment Effect (ATE) of personalization on
minutes_streamed. - Provide a 95% confidence interval and explain at least one valid method for computing it.
- Briefly explain whether and how you would use the given covariates, including why including irrelevant covariates may or may not be acceptable.
Assumptions:
- Randomization occurs at the user level; no interference (SUTVA).
- Use a two-sided 95% confidence interval.
- If you use regression, consider
treatas the only post-treatment variable; all other covariates are pre-treatment.
Overview: This question assesses causal inference and experimental analysis skills by asking for an estimate of the average treatment effect (ATE) of personalization on minutes streamed, a 95% confidence interval, and reasoning about pre-treatment covariates; it falls under the Analytics & Experimentation domain for a Data Scientist position. It is frequently used to evaluate the application of randomized experiment analysis and statistical inference while also probing conceptual understanding of causal assumptions and covariate adjustment, testing both conceptual and practical application.