Uber · Project Deep Dive
Demonstrate business impact from a project
TrueInterview
October 7, 2026 · 6 min read
Pick a past project and go deep on its business impact. State the problem, the decision your work made possible, and the main metric or metrics. Put numbers on the counterfactual and dollar impact, including baseline, lift, confidence or uncertainty, and sample sizes. Walk through the measurement approach—experiment or quasi-experiment—along with the key assumptions and how you checked them. Explain how you dealt with stakeholder disagreement or resistance, got adoption, and handled risks and guardrails. Cover the timeline, the trade-offs you accepted, what went wrong, and what you would change if you did the project again.
Overview: This question tests how well a data scientist can measure and communicate business impact, covering problem framing, metric definition, causal inference through experiments or quasi-experiments, uncertainty quantification, and stakeholder leadership.
Solution
How to Answer, Plus a Worked Example
This structure tends to score well in a technical screen:
- One or two sentences for business context and the decision.
- Two or three sentences to define the OEC (overall evaluation criterion) and guardrails.
- Measurement plan and assumptions, plus how you validated them.
- Numbers: baseline, lift, sample sizes, confidence interval, and dollar impact.
- Stakeholders, risks or guardrails, and adoption.
- Timeline, trade-offs, what failed, and the retrospective.
Here is a worked example built for a two-sided marketplace. The numbers are illustrative but realistic, and they show the expected depth.
Example Project: Cutting Rider Cancellations with Better Dispatch Scoring
- Problem and Decision
- Problem: A high rider cancellation rate during peak demand and in congested corridors caused lost trips and a poor experience. Internal analysis found cancellations were correlated with long, volatile pickup ETAs and suboptimal driver assignment.
- Decision enabled: Whether to launch a new dispatch scoring function that penalizes high-variance pickup ETAs and slightly widens the candidate driver pool in congested areas.
- Metrics (Definitions)
- Primary OEC: Trip Completion Rate (TCR) = .
- Guardrails:
- Pickup ETA, mean and 95th percentile.
- Driver cancellation rate and deadhead distance.
- Driver earnings per online hour.
- Surge minutes share, for marketplace stability.
- Counterfactual and Dollar Impact
- Measurement window: a 14-day geofenced online experiment across three large cities.
- Sample sizes: requests; requests.
- Baseline (control): cancellation rate , so .
- Treatment: cancellation rate , so .
- Lift: percentage points (pp) absolute.
Confidence/uncertainty (difference in proportions):
- .
- Standard error: .
- 95% CI: , or .
Dollar impact (illustrative unit economics):
- Assume a contribution margin per completed trip after variable costs of .
- Monthly requests in the three test cities are about .
- Incremental completed trips per month are about .
- Dollar impact per month is about .
- Using the bounds, the 95% CI is to per month in the test cities.
Scaling (with prudence):
- Suppose the broader network sees about requests per month, and we apply a 0.7 shrinkage factor for heterogeneity and operational differences.
- Scaled impact is about .
- With the CI on , a plausible range is roughly to per month, while acknowledging extra uncertainty from the shrinkage factor.
- Measurement Strategy
- Design: a cluster-randomized online A/B test to limit interference, since drivers and riders interact. We randomized by hex-grid geofences so each rider is always served under a consistent policy within a cell. Drivers were assigned to the policy of the pickup cell at assignment time.
- Why an experiment: it gives a direct counterfactual with high traffic and avoids bias from time trends and confounding supply shifts.
- Variance reduction: CUPED used rider-level and cell-level pre-period cancellation rates as covariates, with stratified randomization by hour of day and cell density.
- Assumptions and validations:
- Interference minimized: we chose sufficiently large cells; border analysis showed no significant spillovers; sensitivity checks excluding border cells were unchanged.
- Randomization balance: A/A checks and covariate balance on pickup ETA distribution, request mix, and weather or events passed.
- Stable logging: the A/A period showed no metric drift, and counter logs for dispatch decisions matched server truth at least 99.9%.
- Seasonality captured: the 14 days included two full weekends, and pre/post comparisons by day of week were consistent.
Quasi-experimental fallback (if A/B impossible):
- Matched difference-in-differences using synthetic control at the cell level with pre-period trends, instrumented by policy availability windows. Validated with placebo tests and parallel-trends checks.
- Execution, Stakeholders, and Adoption
- Stakeholders: Product for throughput, Operations for driver experience, Engineering for reliability, Finance for unit economics, and Policy/Support for edge cases.
- Misalignment/pushback:
- Ops worried about longer deadhead and driver satisfaction; Finance needed margin proof, not just TCR.
- Response: we agreed on an OEC with guardrails and a hard cap: mean pickup ETA could not worsen by more than 0.1 minutes, driver deadhead could increase by less than 0.5%, and earnings per hour had to stay within .
- Risks and guardrails:
- Real-time monitors for pickup ETA, driver cancellation rate, deadhead distance, and incident flags, with an automatic kill switch if any guardrail was breached for 15 consecutive minutes in a city.
- Phased rollout: 10% to 25% to 50% to 100%, with holdouts for continued monitoring.
- Adoption: we shared weekly readouts, city-level playbooks, and a rollback plan. PM and Ops co-owned the rollout gates; Finance validated the CM assumptions and signed off on the dollar impact.
- Timeline, Trade-offs, and Retro
- Timeline (approx.):
- Weeks 1–2: root-cause analysis and metric design, plus offline simulation from historical data.
- Weeks 3–5: feature engineering and model or scoring changes, plus load testing and logging hardening.
- Weeks 6–7: A/A testing and a small-city pilot to validate instrumentation and guardrails.
- Weeks 8–9: a 14-day cluster A/B test across three cities.
- Week 10: analysis, decision review, and rollout planning.
- Trade-offs:
- We accepted a small deadhead increase of +0.4% to gain +0.38 pp TCR, and tightened the variance penalty to keep pickup ETA change within +0.06 minutes.
- Surge minutes fell slightly by 0.2 pp from smoother matching, a positive side effect for user experience but one that required pricing team alignment.
- What failed/surprised us:
- Offline replay overstated gains, predicting about 0.8 pp versus 0.38 pp realized, because of unmodeled driver rejection behavior and road incidents; we fixed this by incorporating pickup ETA uncertainty and adding a driver-acceptance model in simulations.
- Early logging missed a rare reassignment path; A/A testing caught it, and we patched it before the main test.
- What I’d do differently:
- Start cluster A/A testing and border-spillover diagnostics earlier to quantify the design effect and power needs.
- Pre-register the analysis and CUPED covariates to reduce researcher degrees of freedom.
- Build an automated generalization or shrinkage pipeline for rollout forecasts, using learned effect modifiers like density, congestion, and driver supply volatility.
Teaching Notes and Templates You Can Reuse
- Counterfactual/dollars template:
- .
- Use CI bounds on to report a range, and apply a shrinkage factor when generalizing beyond the test scope.
- CI for difference in proportions:
- 95% CI: .
- Power planning (back-of-the-envelope):
- For an absolute effect with baseline and equal arms, per arm is about . Adjust by the design effect for clustering: .
- Guardrail playbook:
- Define thresholds in advance and wire kill switches; monitor in real time; keep a holdout group until post-launch stability is proven.
That level of specificity—a clear decision, precise metrics, a validated counterfactual, quantified dollars with uncertainty, and evidence of leadership—meets the bar for a Data Scientist technical screen.