DoorDash · Statistics & Data Analysis
Compute sample sizes and error control
TrueInterview
October 7, 2026 · 2 min read
Given the Biker experiment setting, calculate the necessary sample sizes and explain how error is controlled when practical limits apply. Provide formulas and numerical results wherever you can.
Assumptions:
- The test has two arms: control versus Biker exposure.
- Two-sided alpha is the default unless a specific item says otherwise.
-
Primary mean metric: The baseline mean delivery time is 42 minutes and the standard deviation is 15 minutes. The target relative improvement is . Use (two-sided) and power . Compute the per-arm sample size for a two-sample t-test on means.
-
Guardrail proportion metric: The baseline cancellation rate is 6%. You need non-inferiority with a margin of +0.5 percentage points, meaning the Biker cancellation rate must be at most 6.5%. Use a one-sided and power . Compute the per-arm sample size for a non-inferiority test on proportions.
-
Multiple metrics: There is 1 primary metric, 2 guardrails (cancellations and ETA accuracy), and 1 secondary metric (orders per courier-hour). Propose and justify an error-control method, such as Bonferroni, Holm/Hochberg, gatekeeping, or HMP. State the effective alpha for each family and explain how adjusted confidence intervals would be reported.
-
Cluster randomization: Randomization is at the zone-day level, with an average of orders per cluster and an ICC of 0.03 for the primary metric. Compute the design effect and the adjusted per-arm sample size. How many zone-days per arm are required to achieve that sample size?
-
Sequential monitoring: Four equally spaced looks are planned using O’Brien–Fleming spending. Explain qualitatively how the early critical values compare with the final look and how this affects runtime and MDE. Give the final-look alpha spending approximation and describe how boundary checks would be implemented in practice.
Overview: This question assesses skill in experimental design and applied inferential statistics, specifically sample size calculation for means and proportions, non-inferiority testing, multiple-comparison error control, cluster-randomized design effects, and sequential monitoring boundaries, within the Statistics & Math domain for a Data Scientist role. It is typically used to test the ability to apply statistical formulas and error-control principles under real-world constraints, balancing power and type I error while accounting for clustering and interim looks; the assessment focuses on practical application built on conceptual understanding of inferential methods.