Amazon · Statistics & Data Analysis
Estimate live sports impact on subscriptions
TrueInterview
October 7, 2026 · 2 min read
Amazon is weighing whether to stream live coverage of certain sporting events on Prime Video. Estimate, from observational data alone, the causal effect of that coverage on Prime memberships and on how much people engage. Be exact about three things: (1) the user-level treatment and comparison groups (for instance, those who see or actually watch the live sport versus those who do not), (2) the main outcomes (for instance, new subscriptions, retention, plan upgrades, watch time), and (3) the covariates that matter for selection on observables.
Begin with matching: pick one approach (k-NN, caliper, Mahalanobis, or propensity-score matching), specify the distance measure or the class of propensity model, describe how you prevent overfitting (regularization, cross-fitting), and give the balance diagnostics you would insist on (SMD cutoffs, variance ratios, overlap plots).
Lay out the identifying assumptions (unconfoundedness, overlap, SUTVA) and say how you would test or defend each.
Next, propose an instrument that handles unobservables by exploiting randomized differences in how prominently the live stream is promoted (for example, a hero banner versus a standard tile).
Write out the 2SLS in full: the first stage regresses LiveWatch on with controls and fixed effects; the second regresses the outcome on with the same controls.
Cover IV validity (relevance, exclusion, independence, monotonicity), weak-instrument diagnostics (the first-stage ), and over-identification tests when you have more than one instrument.
Last, sketch a DID/event-study alternative built on a staggered rollout, spelling out the regression, the fixed effects, and the heterogeneity, and enumerate the main threats (interference, spillovers, time-varying confounding) along with how you would mitigate them.
Overview: This prompt measures a data scientist's grasp of causal inference and observational study design in analytics and experimentation, spanning skills like defining treatment and control, choosing outcome metrics and covariates, and selecting identification strategies such as matching, IV, and difference-in-differences. It is commonly used to judge whether someone can deliver credible causal estimates when a randomized experiment is not possible. It probes both the conceptual side of identification assumptions (unconfoundedness, overlap, SUTVA, IV validity) and the applied side — matching algorithms, propensity models, regression specifications, diagnostics, and staggered-rollout/event-study designs for evaluating real programs.