DoorDash · Statistics & Data Analysis
Evaluate a new ranking model
TrueInterview
October 7, 2026 · 9 min read
A food-delivery business ranks store recommendations on its homepage using model V1.1. Its successor, V2.0, introduces several additional features and may need a different feature-set setup for users in the treatment arm. Design an experimentation and rollout plan for this upgrade. Because the marketplace has two sides, altering what the homepage shows can move consumer demand, merchant exposure, courier utilization, and delivery ETAs — so the plan has to bring together product metric design, causal inference (interference / SUTVA), and operational safety, not just a click-based A/B test. The question is split into seven parts. Handle them as one coherent plan: the metric, randomization, infrastructure, logging, validity threats, statistics, and launch criteria all need to fit together.
Constraints & Assumptions
- Two-sided marketplace: recommendations move merchant demand, courier load, and delivery times, so treating one user can change another user's experience (interference).
- The candidate pool is limited: a store has to be within delivery range and open right now to appear — and that eligible set shifts with time and place.
- V2.0 may rely on extra features, possibly including real-time ones, that V1.1 never used. Treatment has to be able to pull a different feature bundle from control.
- Homepage serving is latency-sensitive (per-retrieval-path budgets in the low single or double digits of ms), so every extra feature computation costs latency.
- Assume traffic is meaningful but finite — variance reduction and power planning matter, since you can't run indefinitely.
Clarifying Questions to Ask
- What is the company's true north — near-term orders/GMV, contribution margin, or long-run retention? That choice sets the primary metric.
- How big a lift must V2.0 produce to pay for the extra infra complexity and any latency cost (i.e., what counts as a practically significant effect)?
- How much interference should we expect — does V2.0 mostly reshuffle the same eligible stores, or does it redirect demand across stores enough to shift ETAs and supply?
- Which feature SLAs and freshness guarantees are in place today, and what are the current missingness and timeout rates for features at serve time?
- What is the baseline homepage-session→order conversion rate, and how much homepage traffic arrives per day (needed for power/MDE)?
- Do experimentation primitives already exist — a bucketing service, a config/feature-flag system, switchback tooling — that we have to build on top of or work around?
Part 1 — Primary success metric and guardrail metrics
Specify the primary success metric and the key guardrail metrics for a homepage recommendation model in a two-sided delivery marketplace. Argue for the primary metric against naive alternatives, and explain why guardrails are mandatory here.
Hint — Where to start: Begin with business value rather than engagement. Ask: which homepage action genuinely creates marketplace value? Then ask what that optimization could quietly break on the supply/operations side. Hint — Pitfall to name: Explain why CTR by itself makes a weak primary metric (noisy, easy to game, a model can push clicks up while real orders fall), and choose something closer to value (e.g. orders or GMV per session). Guardrails need to span both the consumer-latency path and the marketplace/operations side (ETA, cancellations, merchant fairness).
What This Part Should Cover
- One business-aligned primary metric (e.g. orders or GMV/contribution margin per session) with a clear argument for why it beats CTR.
- Guardrail metrics covering serving health (p95/p99 latency, timeout/error rate) AND marketplace health (delivery ETA, cancellation/refund rate, merchant-exposure concentration/fairness).
- Acknowledgement that pushing immediate orders to the max can hurt ETAs, courier load balancing, merchant fairness, and long-run supply diversity.
- A thin layer of secondary/diagnostic metrics (CTR, add-to-cart, reorder, basket/AOV, new-store discovery, retention) used for interpretation, not for the decision.
Part 2 — Unit of randomization
Pick the unit of randomization — user-level, session-level, geo-level, or switchback/time-based — given that recommendations can move merchant demand, delivery times, and marketplace balance. State your default and the condition that would make you change it.
Hint — Key tension: At bottom this is a bias-variance / SUTVA tradeoff. Finer units (user/session) buy power but can break the assumption that one unit's treatment leaves another unit's outcome untouched; coarser units (geo, switchback) contain interference but pay for it in power. Hint — Technique to surface: Call out switchback / geo-time clustered designs as the interference-robust option used in delivery and ride-sharing, and tie the choice to how much V2.0 actually shifts marketplace allocation instead of committing to one dogmatically.
What This Part Should Cover
- Explicit reasoning about SUTVA / interference: why a user-level A/B can be biased once treatment shifts which stores receive demand.
- A comparison of the options with candid pros and cons (power vs. interference containment; session-level contamination across variants).
- A decision rule: user-level when spillovers are small (a modest re-rank of eligible stores); geo-time switchback or zone-clustered when the ranker materially moves marketplace dynamics.
- Awareness of the power penalty from clustering (fewer independent units, geo heterogeneity).
Part 3 — Serving infrastructure for experiment-specific versions and feature configs
Explain how serving infrastructure should support experiment-specific model versions and feature-set configuration, so control and treatment can safely pull different feature lists. Show how you keep that safe and reproducible.
Hint — Decompose the system: Split the problem into three concerns: (1) deterministic assignment, (2) a config/registry that maps an arm to {model version, feature bundle}, and (3) safe behavior when a treatment-only feature is absent at serve time. Hint — Safety mechanism: Before anything goes live, how do you take latency and feature failures off the table without touching users? Consider computing V2.0 outputs without serving them.
What This Part Should Cover
- Deterministic bucketing (a stable hash of
user_idor the geo-time bucket) that records experiment ID and arm — no flapping between requests. - An experiment config service that maps arm → model version + feature bundle, so features aren't hard-coded in the app (control: v1.1/bundle A; treatment: v2.0/bundle B).
- A versioned feature registry (schema, types, freshness SLA, defaults, owners, optional/deprecated flags) plus backward-compatible serving (safe defaults and a missingness indicator, never a hard serving failure).
- Shadow mode ahead of live traffic: compute V2.0 scores in parallel and compare latency, score distribution, missingness, and calibration.
Part 4 — Logging requirements
Specify which events and metadata have to be logged so the experiment can be analyzed correctly and reproducibly.
Hint — What to anchor on: The test to apply: can you reconstruct exactly what each user saw and why? Log enough to attribute outcomes to an arm, a model version, and a specific candidate list — plus the diagnostics that later explain validity threats.
What This Part Should Cover
- Assignment-level fields: experiment ID, arm, unit ID (user/session/geo-time bucket), timestamp plus timezone, model version, feature-config version.
- Request/ranking-level fields: the candidate set before ranking, the ranked list actually shown, per-candidate scores where practical, feature-missingness/freshness indicators, per-component latency.
- Marketplace/ranking diagnostics: size of the eligible/serviceable pool, whether a fallback fired, stale-feature usage, store-level exposure.
- Downstream outcomes with an attribution window: click, add-to-cart, order, basket size, cancellation.
Part 5 — Practical validity threats
Explain how to deal with sample ratio mismatch, delayed conversions, feature missingness, novelty effects, selection bias, and spillover/interference. For each, give a concrete diagnostic or mitigation.
Hint — Triage order: Check experiment integrity before you read any lift — assignment and logging health first. Then deal with timing (delayed conversions), then the biases that blur infra quality or eligibility shifts into model quality, then interference. Hint — The subtle ones: With feature missingness, ask whether you're measuring the model or the infrastructure (if treatment has more missing real-time features). For selection bias, remember the eligible store set moves with time and place. For interference, tie back to the randomization choice you made in Part 2.
What This Part Should Cover
- SRM: define it (planned versus observed split), what it points to (assignment/logging bugs, crashes caused by treatment, geo routing), and that lift isn't trusted until it's resolved.
- Delayed conversions: an attribution window and an analysis window; why reading too early skews toward click-heavy variants.
- Feature missingness: missingness indicators, freshness logging, and slicing missing against non-missing so infra quality isn't mistaken for model quality.
- Novelty effects, selection bias, spillover/interference: watch for novelty decay over time; log eligible-pool composition so eligibility shifts aren't mistaken for lift; contain interference with switchback/zone clustering and marketplace-level outcome monitoring.
Part 6 — Power / MDE and variance reduction
Explain how to estimate power / MDE, and when stratification or CUPED helps. Show the quantitative reasoning, including how clustering changes the math.
Hint — Formula to reach for: For a binary metric, tie sample size to the baseline rate and the absolute detectable lift through the standard two-proportion sample-size approximation, then correct for clustered designs with a design effect driven by cluster size and intra-cluster correlation. Hint — Variance reduction: CUPED leans on a pre-period covariate correlated with the outcome (e.g. prior order count) to subtract predictable variance. Think about which pre-experiment covariates are both available and predictive in this setting.
What This Part Should Cover
- A power/MDE estimate for a binary metric using something like per arm, worked through on a concrete baseline (e.g. , ).
- The design effect for clustered/switchback designs: , and why geo-level tests demand more traffic or more time.
- CUPED: the adjustment , what represents, useful covariates (prior 7-day orders, prior sessions, pre-period spend), and the payoff (lower variance → smaller MDE → shorter test).
- Stratification: which slices matter (new versus returning, dense versus sparse markets, supply conditions, platform), and the Simpson's-paradox risk when the traffic mix differs across arms.
Part 7 — Ramping, rollback, and launch decision
Set the criteria for ramping, rollback, and final launch. Give the ramp sequence, explicit rollback triggers, and a multi-factor launch decision (not merely "is lift positive?").
Hint — Sequence: Stage exposure so operational failures surface before statistical ones: offline validation → shadow → canary ramp → full experiment. Pair each stage with what it is checking.
What This Part Should Cover
- A ramp sequence: offline replay/backtest with point-in-time features (NDCG, log loss, calibration) → shadow mode → canary ramp (e.g. 1%→5%→25%→50%) → full experiment spanning complete weekly cycles.
- Rollback triggers: latency regression, worsening ETA/cancellation, elevated missingness or timeouts, SRM/logging corruption, severe merchant-exposure skew.
- A launch decision framework: primary metric significant AND guardrails intact AND robust across key slices AND gain large enough to cover infra/latency cost AND not dependent on fragile real-time features AND not merely novelty.
What a Strong Answer Covers
Across all seven parts, a strong answer should come across as one coherent plan instead of seven disconnected checklists. Cross-cutting signals an interviewer looks for:
- Internal consistency — the Part 2 randomization choice, the Part 4 logging, the Part 5 interference handling, and the Part 6 power math reinforce one another (e.g. choosing switchback implies a design-effect penalty and marketplace-level logging).
- Marketplace/causal sophistication — interference, SUTVA, and two-sided effects are treated as first-class rather than an afterthought.
- Operational safety as a co-equal goal — guardrails, shadow mode, and staged rollout are integral rather than bolt-ons.
- Quantitative grounding — a concrete metric, baseline, MDE, and a defensible launch bar.
- Judgment over dogma — defaults paired with explicit switch conditions, and a launch decision that weighs lift against complexity, latency, and fragility.
Follow-up Questions
- Say user-level SRM is clean but the geo-time switchback reveals a strong day-part interaction (treatment wins at lunch, loses at dinner). How do you decide whether to launch, and to whom?
- Most of V2.0's lift comes from one real-time feature with a 2% serving timeout rate. How do you quantify how much of the measured lift is the model versus infra quality, and what would you demand before launch?
- Online conversion rose +0.4% while offline NDCG stayed flat. How do you reconcile the two, and which do you trust?
- After full launch, lift decays toward zero over three weeks. How do you tell a novelty effect apart from a real regression introduced during ramping, and what experiment would you run to find out?
Overview: This question assesses expertise in experimentation design and causal inference inside two-sided marketplace environments. It tests the ability to handle interference, SUTVA …[truncated]…