Uber · Behavioral Stories
Navigate urgency, priorities, and conflict
TrueInterview
October 7, 2026 · 7 min read
Choose a single actual project in which you dealt with substantial ambiguity and reliance on other teams. Address the following:
(a) What draws you to this company and team, and why make this career shift at this point? Connect it to the mission, the metrics you would be accountable for, and particular domain problems.
(b) Outline the project's objective, the way you were assigned to it, the stakeholders involved (inside and outside the company), and the part you played. Which aspects were difficult, and for what reasons?
(c) Halfway through, a Sev-1 incident creates a deliverable that must be finished in 48 hours. Explain your framework for ranking priorities across four projects running at the same time, making the quality-versus-time tradeoff explicit. What did you drop, postpone, or run in parallel? How did you convey risk and obtain buy-in?
(d) Describe a time you disagreed with a central stakeholder about method or scope. How did you bring assumptions into the open, collect supporting evidence, and sway others without formal authority? What ended the disagreement, and which metrics shifted?
(e) Look back on the results and the lessons. What would you change on a future attempt, and how would you make those changes permanent (playbooks, dashboards, SLAs)?
Overview:
This question gauges how well a candidate handles ambiguity, ranks competing projects and deadlines, works through cross-team dependencies, persuades stakeholders without positional authority, and puts numbers on outcomes and process gains.
Solution
Context
Suppose the role is Data Scientist at a large two-sided mobility marketplace (riders and drivers). The work touches marketplace health, pricing, and reliability, with partners across Product, Engineering, Operations, Finance, Legal/Policy, and Customer Support.
(a) Why this team and why now
- Mission tie: I care about making transportation at scale dependable and fair. The marketplace's objective — cutting wait time while keeping driver earnings sustainable — lines up with what I do best: causal inference, experimentation, and analytics in production.
- Metrics I'd own: median and 95th percentile wait time (ETA), trip completion rate, cancellation rate, dispatch success rate, and fairness of driver earnings (variance and Gini).
- Domain challenges that excite me:
- Demand that spikes and shifts, varying by geography.
- Decisions made in real time under latency limits (scoring under 100 ms, data freshness SLAs).
- Incentives on both sides of the market, with fairness and regulatory limits.
- Observability and incident handling for services the marketplace depends on.
(b) Project: Dynamic Incentives to Stabilize Peak Demand
- Goal: Cut peak-hour wait time by 10% and cancellations by 3% across two metros by rolling out real-time, geo-targeted driver incentives driven by marketplace health signals.
- How I got staffed: I had already delivered a causal attribution framework for changes in driver supply, so Product asked me to head analytics and experimentation for this effort.
- Stakeholders:
- Internal: Product (Marketplace, Pricing), Engineering (Incentives Service, Data Infra), Operations (city teams), Finance (budget), Legal/Policy (fairness/compliance), Customer Support (rider/driver sentiment), Data Platform.
- External: Driver councils (feedback), a few B2B partners (scheduled rides).
- My role:
- Set success metrics and guardrails; design the experimental rollout; estimate budget and ROI; build monitoring dashboards; work with Eng to instrument events; own the analysis and readout.
- What was hard and why:
- Causal paths were unclear: incentives change driver supply, which moves surge and ETA, and weather or events confound it.
- Data latency and reliability: health signals had to be fresh within a minute, while some sources refreshed only every 5–15 minutes.
- Governance: keeping geographic fairness and preventing unwanted swings in earnings.
(c) Sev-1 incident and 48-hour deliverable
Partway through the pilot, a regression in the surge model priced certain zones too low during a major event. The symptoms were a 15% rise in median wait time and an 8% rise in rider cancellations in two metros, along with a surge in support tickets. Leadership asked for three things within 48 hours: (1) a quantified impact, (2) a mitigation plan, (3) a provisional recalibration.
- My prioritization framework across four concurrent projects:
- The live dynamic incentives A/B test (high impact, time-sensitive).
- Rider churn model refresh (high impact, less urgent).
- Fraud anomaly PoC (medium impact, medium urgency).
- Weekly city demand forecast (operational, routine). I ranked the work by Severity × Blast Radius × Reversibility, together with time-to-mitigation and fit with company KPIs.
- Tradeoffs (quality vs. time):
- To size the incident impact quickly, I ran a Difference-in-Differences (DiD) against a nearby control city untouched by the regression, instead of a full causal forest with heterogeneity.
- DiD estimator:
- Example: With wait time (min) going 5.6→6.5 in the treated area and 5.4→5.5 in the control, DiD = (6.5−5.6) − (5.5−5.4) = 0.9 − 0.1 = +0.8 min impact.
- Cut: heterogeneity by zone and driver cohort, long-run churn estimates, and a full backfill with late-arriving data.
- Defer: the churn model refresh by one sprint; the fraud PoC experiments by one week.
- Parallelize: One analyst gathered event-level logs; I handled the DiD and cost impact; Eng took rollback and the feature flag; Ops assembled qualitative support signals; Finance verified variable spend.
- To size the incident impact quickly, I ran a Difference-in-Differences (DiD) against a nearby control city untouched by the regression, instead of a full causal forest with heterogeneity.
- Communication and alignment:
- Stood up a 24-hour war room with a one-pager covering the incident summary, hypotheses, decision log, mitigation plan, and a risk register (likelihood × impact, owner, next check-in).
- Executive updates at T+12h and T+36h: the current impact estimate, confidence level, mitigation status, and next steps.
- Mitigation guardrails: rider wait time no more than target +5%, driver earnings variance change under 2 p.p., and no geography's incentives above the budget cap.
- Outcome at T+48h:
- Rolled back the faulty model and put a temporary surcharge floor in the affected zones.
- Estimated incremental impact: +0.8 min to median wait, −2.9 p.p. completion; projected revenue loss $420k (95% CI: $350k–$500k). Confidence: medium (the control city match was validated with pre-trend checks, p>0.1 for the pre-trend difference).
(d) Disagreement on methodology and scope
- Disagreement: Product wanted to call the pilot a success from pre-post gains in the pilot city and scale it globally. I held out for city-level randomized rollouts or, at minimum, DiD with pre-trend validation and a power analysis.
- How I surfaced assumptions:
- I wrote the assumptions down plainly: seasonality, event confounders, driver supply spillovers, regression to the mean.
- I gave a counterexample in which a pre-post reading "improved" because of a weather shift with no treatment at all.
- Evidence gathered:
- Back-tested the incentive policy on historical weeks (off-policy evaluation via inverse propensity weighting using prior propensity to be incented by zone and time).
- Ran a two-city stepped-wedge rollout with 20% zone-level randomization.
- Power analysis for the primary metric (median wait time). For a target detectable effect δ = 0.4 min, baseline σ ≈ 2.0, cluster-robust ICC ≈ 0.05, we calculated the zone-hours needed for 80% power at α = 0.05; we reached it in 10 days.
- Influence without authority:
- Framed the choice as "faster launch" versus "credible, scalable proof," with the quantified risk of false positives (type I error around 30% under plausible confounding).
- Offered a compromise: a limited ramp with holdout zones and a 7-day readout gate.
- Resolution and metrics moved:
- We agreed to run the stepped-wedge with guardrails.
- Results: −7.2% median wait time (−0.42 min), −3.1 p.p. cancellations, +4.3% driver online hours; budget +1.6% variable spend. Key guardrails stayed within limits. p<0.01 for primary outcomes; no significant negative movement in earnings variance.
(e) Outcomes, lessons, and institutionalization
- Outcomes:
- Shipped dynamic incentives to two metros; expanded to four after 6 weeks.
- Built a near-real-time marketplace dashboard: wait time percentiles, dispatch rate, cancellation rate, surge accuracy, incentive take-rate; with alerting on z-score anomalies and data freshness SLIs.
- Documented the incident, including root cause (model regressor drift plus insufficient shadow testing) and a plan to reduce time-to-detect (TtD).
- What I'd do differently:
- Add shadow deployments for pricing and surge models with automatic canary analysis before full enablement.
- Pre-register experiment designs and analysis plans so stakeholders align earlier.
- Tighten data contracts and freshness SLAs for the health signals feeding incentives.
- How I'd institutionalize improvements:
- Playbooks: an incident response runbook (Sev-1/2) with roles, first-hour checks, standard queries, DiD templates, and a communication cadence.
- Dashboards & alerts: SLOs for surge accuracy and ETA calibration; freshness monitors for critical Kafka topics that page on breach; guardrail alerting on cancellations and earnings variance.
- SLAs & ownership: clear RACI for model changes (DS sign-off, Eng owner, Product approver); a pre-launch checklist (instrumentation parity, shadow metrics within tolerance, rollback plan tested).
- Experiment standards: city/zone-level randomization as the default; guardrail metrics hard-coded in the experiment config; minimum detectable effect calculators embedded in experiment tooling.
- Validation/guardrails summary:
- Primary: median wait time, completion rate.
- Guardrails: cancellation rate, driver earnings variance, surge error (|actual − predicted|), support ticket rate, data freshness SLI.
- Post-launch: weekly DiD readouts; heterogeneity checks (zone, hour, cohort); counterfactual simulations under demand spikes.
This end-to-end approach shows handling ambiguity, coordinating across teams, making principled speed-versus-quality tradeoffs under pressure, and turning lessons into lasting processes and tooling.