DoorDash · Statistics & Data Analysis
Diagnose Cold-Food Deliveries and Make a Launch Decision
TrueInterview
October 7, 2026 · 3 min read
You work as a data scientist at a food-delivery marketplace. Some customers report that their orders show up cold. The product team has suggested an intervention designed to keep food warm in transit and wants to know if it should be rolled out broadly. Walk through how you would diagnose the issue, set success criteria, assess the intervention, and make the launch decision. Name the data you would examine and the comparisons you would use. If randomized A/B testing is not feasible, offer a credible alternative.
Constraints & Assumptions
- Consider this a marketplace involving customers, merchants, and couriers. An outcome that helps one side while seriously hurting another is not acceptable as success.
- The term “cold food” is not yet a measurable operational metric. You can suggest a definition, but acknowledge its limits.
- The intervention is intentionally left vague. Specify which parts of your design would change depending on whether it acts at the order, courier, merchant, or market level.
- Available data includes order events, promised and actual delivery timestamps, support contacts, refunds, ratings, merchant and courier identifiers, geography, item attributes, and experiment-exposure logs when experiments are run.
- Do not treat a complaint or refund as a perfect signal of food temperature.
Clarifying Questions to Ask
- Which intervention is proposed, and at what unit can it be randomly assigned?
- Which customer segments, merchants, cuisines, order distances, or markets seem most affected?
- Is there a direct temperature measure, or only proxies like complaint reasons, refunds, ratings, and delivery time?
- Is the goal to lower cold-food incidents, boost retention, reduce support costs, or balance several objectives?
- What decision deadline, rollout risk, minimum meaningful effect, and observation window are in play?
Part 1: Diagnose and Measure the Problem
Explain how you would convert the vague “cold food” report into a measurable outcome. Outline the analyses you would perform to estimate how common the issue is, locate where it occurs, and separate likely operational causes from reporting artifacts. Hint: Begin with measurement and the delivery funnel before selecting a statistical model.
What This Part Should Cover
Part 2: Choose Primary, Secondary, and Guardrail Metrics
Suggest a hierarchy of metrics for assessing the intervention. Set one primary metric, helpful secondary metrics, and guardrails covering customers, merchants, couriers, and marketplace economics. For each important metric, indicate the direction of improvement and the unit of analysis. Hint: A metric tree can link the customer issue to operational drivers without treating every correlated metric as a success criterion.
What This Part Should Cover
Part 3: Estimate the Intervention's Causal Effect
Design a randomized evaluation if it is feasible. State the randomization unit, control condition, exposure logging, analysis population, and comparison. Cover interference, statistical power, novelty effects, and heterogeneous treatment effects. Then describe how you would estimate the impact if randomization cannot be used. Hint: Align the assignment unit with how the intervention is delivered, and make any observational design’s identifying assumptions testable when possible.
What This Part Should Cover
Part 4: Decide Whether to Launch
Assume the evaluation is complete. Outline a decision framework for launching, iterating, collecting more data, or stopping. Explain how you would treat a statistically significant but very small improvement, a promising average with harm in one segment, and an inconclusive result. Hint: Distinguish evidence quality, effect size, risk, and reversibility instead of reducing the call to a p-value threshold.
What This Part Should Cover
What a Strong Answer Covers
Follow-up Questions
- Complaints drop after launch, but refunds and delivery times stay the same. How would you determine if the product worked or simply changed how people report issues?
- The intervention is assigned to couriers, but those couriers handle both treatment and control orders. What bias could this introduce, and how would you redesign the experiment?
- The overall effect is positive, but long-distance orders get worse. What further analysis and launch policy would you suggest?
- Only one market can get the intervention at first. Which quasi-experimental approach would you consider, and what pre-period evidence would you require?
- How would you compute the minimum detectable effect and choose the experiment duration when cold-food reports are rare?
Overview: Work through a product analytics case focused on diagnosing a delivery-quality issue and deciding if an intervention should launch. The prompt assesses metric hierarchy, causal reasoning, study design, power considerations, and decision criteria.