Uber · Statistics & Data Analysis
Formulate hypotheses and compute AB test significance
TrueInterview
October 7, 2026 · 2 min read
Use the A/B test snapshot below for the pickup ETA card experiment to answer every part.
Data (7-day snapshot):
- Primary metric (trip completion rate per request): • Control A: requests, completions • Treatment B: requests, completions
- Guardrail 1 (rider cancel rate per request): • Control A: • Treatment B:
- Guardrail 2 (wait time minutes, per request): • A: , , • B: , ,
- Five interim looks occurred at equally spaced information times, with no pre-registered alpha spending.
Tasks:
- Write the exact and for the primary metric; say whether the test is one-sided or two-sided and explain why.
- Pick the right test for the primary metric (difference in proportions), then compute the test statistic, p-value, and a 95% confidence interval for the lift. Show the formulas and the numeric results.
- For Guardrail 2 (mean wait time), choose the correct test (for example, Welch’s t-test) and compute the 95% confidence interval for the mean difference. State any distributional assumptions and explain why Welch rather than pooled.
- Apply a multiple-testing correction across the three outcomes (Primary, Guardrail 1, Guardrail 2) with Holm–Bonferroni at a familywise . State which effects remain significant.
- In plain language, explain what the p-value from part (2) does and does not mean.
- Because there were 5 unplanned interim looks, reassess significance with a simple Pocock or O’Brien–Fleming alpha-spending method; outline the approach and give an approximate adjusted conclusion. Exact boundaries are not required, but justify your choice.
- If the pre-period completion rate per rider has correlation with the in-experiment outcome, estimate the approximate variance reduction from CUPED and discuss how that would affect required sample size or interpretation.
- Conclude: ship, iterate, or stop? Defend your decision while taking the guardrails into account.
Overview: This question tests a data scientist’s ability in experimental design and statistical inference for A/B testing. It covers hypothesis formulation, difference-in-proportions testing and confidence intervals, guardrail analysis, multiple-testing correction, interim alpha-spending approaches, and variance-reduction techniques such as CUPED.
This item is drawn from a data scientist interview experience.
Loading comments…