Uber · Statistics & Data Analysis
Analyze results and large p-values correctly
TrueInterview
October 7, 2026 · 1 min read
Once the test has concluded, lay out precisely how the results will be examined: estimate the intent-to-treat effect at the level of user assignment using cluster-robust standard errors; apply CUPED or pre-period covariates to cut down variance; and deal properly with ratio metrics (via the delta method or Fieller) as well as skewed outcomes. Walk through the reasons session-level analysis is unsuitable in this setting (each user contributes repeated measurements, session counts are tied to treatment, and observations are not independent) and describe the remedies (roll up to the user, use mixed models, apply cluster-robust SE). Address non-compliance or partial exposure (people who never opened) and obtain the TOT through 2SLS, treating assignment as the instrument. When the p-value comes out large, determine whether to: not reject versus assert there is no effect, perform a post-hoc power/MDE check, and/or carry out equivalence or non-inferiority tests (TOST); as an option, contrast with a Bayesian posterior using a ROPE. Sketch out heterogeneity analysis and control for multiple testing.
Overview: The purpose of this question is to gauge how well a candidate has mastered experimental analysis and applied causal inference, covering intent-to-treat estimation, cluster-robust inference, variance reduction, treatment of ratio metrics and skewed outcomes, non-compliance and instrumental-variable approaches, frameworks for deciding when p-values are large, and heterogeneity alongside multiple-testing control. Frequently posed within the Statistics & Math area to test practical application paired with conceptual understanding, it assesses whether someone can reason about the right level of analysis, the soundness of inference, and how to interpret ambiguous findings, not merely their computational ability.