Uber · Behavioral Stories
Demonstrate Leadership in Ambiguous Analytics Projects
TrueInterview
October 7, 2026 · 7 min read
Describe an analytics project you owned from start to finish where the objectives were unclear and the schedule was compressed. Address the following:
- Situation/Task: The business setting, the deadline for the decision, and how "success" was defined in quantifiable terms.
- Action: How you framed the problem, balanced stakeholders whose goals conflicted (for example, Growth versus Marketplace Health), and worked through trade-offs. Include a specific moment when you challenged a senior stakeholder who asked for a metric you considered misleading; what evidence did you bring, and how did you communicate your case?
- Result: Measurable impact with figures (for instance, +X% completion, −Y% cancellations, TTR cut by Z days). How did you make the work reproducible and hand it off (documentation, dashboards, alerts)?
- Reflection: One choice you would reverse with hindsight and why; what early indicator would have led you to reverse it sooner.
- Follow-up: If you had to run the project again as a two-week sprint beginning today, what would your day-by-day milestone plan be, and which interim leading indicators would you track? Overview: This question tests leadership, stakeholder management, hypothesis-driven analytics, metric design, and end-to-end delivery of analytics work under ambiguity and deadline pressure, including setting measurable KPIs and ensuring reproducibility.
Solution
Sample Response and Instructional Walkthrough
Here is a structured, end-to-end example suited to a marketplace analytics position. It shows how to scope under uncertainty, manage stakeholders, apply rigorous methods, and move quickly.
1) Situation/Task
- Context: A large city was approaching a holiday weekend that was expected to produce a demand surge. Recent data showed rider cancellations increasing, driven by long ETAs and sharp surge pricing. Leadership had 10 days to make a go/no-go call on a set of levers—targeted rider promotions, temporary driver incentives, and match-policy adjustments—to sustain growth without exhausting supply.
- Ambiguity: Different teams advocated different success definitions: Growth cared about request volume; Marketplace cared about supply health and reliability; Support focused on complaints. The goals had not been aligned.
- Success criteria (agreed in a working session):
- Primary KPI: Incremental completed trips (ICT) relative to a counterfactual.
- Guardrails:
- Cancellation rate: at least −10%.
- p90 ETA: improve by ≥ 1 minute (to avoid long-tail wait times).
- Driver earnings per online hour (DEPH): at or above baseline (no supply damage).
- Unit economics: contribution margin must be non-negative.
- Decision deadline: 10 days to recommend a plan; 14 days to pilot and decide on broader rollout.
2) Action
a) Scoping under ambiguity
- Problem framing: I built a metric DAG connecting upstream levers to outcomes: supply online hours → acceptance rate → ETA distribution → rider conversion → cancellations → completed trips → margin. This made failure modes and measurable touchpoints explicit.
- Hypotheses:
- Long-tail ETAs (p90) cause a disproportionate share of cancellations.
- Broad demand promotions without supply pacing make p90 ETAs worse, raising cancellations even when requests increase.
- Targeted promotions in time/geography pockets with spare supply and light driver incentives produce a net increase in completed trips without hurting DEPH.
- Data checks: Baseline by hour-of-week and zone; regression and partial dependence plots showing cancellation probability climbing sharply once ETA passes 9 minutes; variation by zone and weather.
b) Prioritizing stakeholders and negotiating trade-offs
- Stakeholders and objectives:
- Growth: maximize near-term request volume.
- Marketplace Ops: protect driver earnings and reliability (p90 ETA, cancellations).
- Finance: preserve margin.
- Support: reduce wait-time complaints.
- Trade-off framework:
- Proposed primary KPI (ICT) with guardrails (p90 ETA, DEPH, margin). The framing was: "We will grow only where reliability and supply health remain intact."
- Targeting: Restrict promotions to zones/hours with acceptance rate above 88% and idle time above 10% (a proxy for spare supply). Add small driver incentives in thin zones.
- Pacing: cap promotion issuance per 15-minute interval to prevent demand spikes.
c) Pushing back on a misleading metric (senior stakeholder)
- Request: A senior Growth leader wanted success to be defined by average ETA (mean) and total requests, arguing these were simple and quick to move.
- Why misleading:
- ETA distributions were heavy-tailed; the mean hid reliability problems. A small improvement for many riders could coexist with a worse tail for some riders, raising cancellations and complaints (Simpson’s paradox across zones).
- Evidence presented:
- A histogram of ETAs showed a long tail; a pilot simulation indicated mean ETA −8%, but p90 ETA +1.5 minutes in high-traffic zones, with a corresponding +3.2 percentage-point cancellation rate in those zones.
- Correlation: p90 ETA had a 2.3× stronger correlation with cancellations than mean ETA; NPS declines tracked tail deterioration.
- Small numeric example: Two zones with equal volume.
- Zone A: mean ETA 6→5.5 min (better), p90 10→9.5.
- Zone B: mean ETA 8→7.2 (better), p90 12→14 (worse). The overall mean looks better, but cancellations increased because of Zone B’s tail.
- Communication approach:
- A pre-read with a one-page visual, emphasizing customer harm stories tied to tail risk.
- Framed as risk mitigation: "Optimizing the mean risks hidden churn; p90 protects reliability."
- Proposed compromise: Primary KPI = ICT; reliability guardrail = p90 ETA; Growth would still see requests as a leading indicator, not a success metric.
- Outcome: Agreement to track p90 ETA and cancellations as guardrails; requests were downgraded to a leading indicator.
d) Experiment and analysis design
- Design: Geo-time segmented rollout across matched city zones, using synthetic control plus difference-in-differences to estimate incremental impact while controlling for seasonality and weather.
- CUPED adjustment: Reduced variance with pre-period outcomes X and the transformation , where .
- Power: Targeted detection of a +5% lift in completed trips at 90% power; variance estimates from eight weeks of history guided zone sample selection.
- Leading indicators monitored hourly: search-to-request conversion, acceptance rate, p90 ETA, driver idle time, DEPH, and wait-time complaint rate.
3) Result
- Outcomes over the 10-day pilot (test vs. control, CUPED-adjusted):
- Incremental completed trips: +6.8% (95% CI: +4.9% to +8.6%).
- Cancellation rate: −12.1% (from 16.5% to 14.5%).
- p90 ETA: −1.3 minutes (from 11.2 to 9.9).
- DEPH: +3.0% (supply health maintained).
- Contribution margin: +1.1 percentage points (after promotion and incentive costs).
- Operational learning: On day 2, one zone breached the p90 ETA guardrail because of a local event; auto-pacing cut promotions by 35% in that zone within 30 minutes, preventing wider tail deterioration.
- Reproducibility and handoff:
- Code: Versioned repository (SQL + Python), modular queries, data contracts on key tables, unit tests for transformations.
- Pipelines: Scheduled daily jobs (orchestration) to refresh metrics; anomaly checks on core KPIs with thresholds and backfills.
- Dashboards: Executive view (ICT, cancellations, p90 ETA, DEPH, margin), Ops view (zone/hour cuts, alerts).
- Docs: Two-page experiment design and assumptions, metric definitions, lineage; runbook for incident response and tuning.
- Alerts: Slack/email when guardrails are breached (for example, p90 ETA +1 min vs. baseline for at least two consecutive intervals), with auto-pacing hooks.
4) Reflection
- What I’d change: I would pre-register heterogeneous treatment effects (HTE) by zone type and ETA band and use that to target from day 1, instead of discovering it mid-pilot. We found that certain dense downtown zones gained less from promotions unless paired with stronger supply incentives.
- Early signal that would have prompted the change sooner: The first 24-hour uplift distribution showed high variance and a right skew in cancellations for downtown zones. A Bayesian shrinkage view of zone-level lift would have flagged a low posterior probability of positive lift earlier, which would have justified immediate re-targeting.
5) If re-running as a 2-week sprint (Day-by-Day)
- Day 1: Kickoff, confirm the decision deadline, align on success metrics (ICT as primary; p90 ETA, cancellations, DEPH, margin as guardrails). Run data audit and access checks.
- Day 2: Build metric DAG, baseline by zone/hour, set eligibility (zones with acceptance rate above 88%, idle time above 10%). Pre-register hypotheses and HTE plan.
- Day 3: Design experiment (geo-time DiD + synthetic control), sample size/power, CUPED covariates, guardrail thresholds and auto-pacing rules.
- Day 4: Build data models, QA tests, backfill eight-week baselines; create a monitoring dashboard skeleton.
- Day 5: Stakeholder readout to confirm trade-offs; finalize promotion/incentive parameters; implement alerts.
- Day 6: Soft launch in two pilot zones; validate data plumbing, latency, alert firing, and guardrail logic with dry runs.
- Day 7–8: Expand to matched zones; hourly monitoring; adjust pacing in response to guardrail breaches.
- Day 9–10: Interim analysis with CUPED; HTE cuts; decision checkpoint to keep/stop/retarget zones.
- Day 11–12: Consolidate results, margin analysis, sensitivity checks (weather/events), falsification tests.
- Day 13: Executive readout with recommendation; pre-approve scaled rollout conditions.
- Day 14: Handoff: finalize dashboards, runbook, pipeline ownership; schedule post-rollout review.
Interim leading indicators
- Demand: search-to-request conversion, request growth (as a signal, not a success metric).
- Supply: online hours, acceptance rate, driver idle time, driver cancellations.
- Reliability: p50/p90 ETA, wait-time complaint rate.
- Health: DEPH, contribution margin per trip.
- Safety valves: Guardrail breaches trigger auto-pacing and/or a zone-level pause.
Methods Summary (for learning)
- Incremental lift via DiD: .
- CUPED variance reduction: , with .
- Tail-aware reliability: Optimize p90 ETA (and cancellations due to wait) instead of mean ETA to avoid masking tail risk.
- HTE-first targeting: Use early zone-level posteriors to reallocate exposure quickly.
Pitfalls and Guardrails
- Pitfall: Vanity metrics (requests, mean ETA) can hide harm; use causal KPIs and reliability guardrails.
- Pitfall: Spillovers between zones; use geographic buffers and sensitivity checks.
- Guardrails: Pre-commit to thresholds and automatic throttling to prevent runaway tail risk.
- Validation: Placebo tests in the pre-period, backtest on prior events, and monitor data quality alerts.
Loading comments…