Uber · Product & Business Case
Explain grocery-specific product strategy and scrappy XP
TrueInterview
October 7, 2026 · 8 min read
Imagine a delivery platform rolling out "Eats Grocery" in New York City: what genuinely separates grocery from restaurant delivery, and how should that difference show up in the product and in your experimentation? Name at least five distinct aspects (for example substitution behavior, perishability and cold chain, basket size and variety, inventory accuracy and stockouts, picking and packing latency, delivery windows, fees and margins, refund and replace policies); choose one growth bet and one efficiency bet, spell out crisp success metrics and leading indicators, and design a scrappy, low-risk experiment (treatment, control, assignment, duration) to validate them; list the top three risks with mitigations, and explain how experimentation feasibility differs from restaurants (e.g., store-level randomization, switchbacks within store hours, catalog volatility) along with the product knowledge you would need in advance.
Overview: This question assesses a data scientist's product-strategy and experimentation-design competencies, covering operational constraints, metric definition, hypothesis framing, risk identification, and cross-functional leadership through a grocery vertical launch.
Solution
1) How grocery diverges from restaurant delivery (and why that matters)
Nine grocery-specific factors follow, each with product and experimentation implications.
- Substitution behavior (item-level fulfillment)
- Why it differs: A single order spans many SKUs, and stockouts are common, so shoppers and pickers routinely swap items.
- Product: Substitution preferences set up front ("same brand" versus "lowest price"), pre-approved alternatives, price-parity rules, and picker-facing suggestion UX.
- Experimentation: Item-level outcomes such as fill rate and substitution acceptance rate carry weight; cluster at the store or order level to keep arms from interfering.
- Perishability and cold chain
- Why it differs: Frozen and chilled items need temperature control, and long waits degrade quality.
- Product: Cold-chain badges, enforcement of courier thermal-bag use, and routing that delivers chilled items last.
- Experimentation: Guardrails on late arrivals of temperature-sensitive goods; store-hour switchbacks may be needed to respect capacity.
- Basket size and variety
- Why it differs: Orders hold more items and more weight and volume, and they mix planned shops with quick top-ups.
- Product: Aisle navigation, bundles, free-delivery thresholds, and scheduled delivery.
- Experimentation: AOV has fatter tails, so power analyses need variance-aware designs and longer durations.
- Inventory accuracy and stockouts
- Why it differs: The catalog shifts constantly, APIs and feeds are imperfect, and what is on the shelf does not equal what the feed reports.
- Product: Real-time stock badges, out-of-stock prediction, hiding or de-ranking risky SKUs, and proactive substitutions.
- Experimentation: Stores vary widely; stratify by partner and integration type, and watch for non-compliance when feeds fail.
- Picking and packing latency
- Why it differs: In-store picking adds preparation time, and shopper efficiency varies.
- Product: Aisle-aware pick lists, batching, image search, and barcode-scan verification.
- Experimentation: Staff behavior shifts, so prefer switchbacks while holding picker training constant.
- Delivery windows (scheduled vs ASAP)
- Why it differs: Customers plan baskets, stores have hours and cutoffs, and staging enables batching.
- Product: One- to two-hour windows, dynamic pricing by window, and capacity gating.
- Experimentation: Window availability must respect capacity; randomize at the store-hour (switchback) level rather than the user level.
- Fees, margins, and contribution economics
- Why it differs: CPG margins are thinner, picking cost is significant, and ad and retail-media revenue are possible.
- Product: Smart fees, minimums, threshold-based promos, and retail-media placements.
- Experimentation: Optimize contribution margin per order, not conversion alone, and include refunds and shopper time.
- Refund and replace policies
- Why it differs: Return and refund volume is higher, and partial refunds are routine.
- Product: Self-serve refunds, price-adjusted substitutions, and post-order make-goods.
- Experimentation: Guardrails that catch moral hazard and fraud, with outcome metrics net of refunds.
- Compliance and restricted items (ID, EBT/SNAP, alcohol)
- Why it differs: Legal and ID checks apply, tender types vary, and age-restricted flows exist.
- Product: ID-at-door flows, tender selection, and restricted-item gating.
- Experimentation: Segment experiments to exclude flows under strict compliance, or monitor them separately.
2) Two focused bets
A) Growth bet — Scheduled delivery windows to unlock planned baskets
- Hypothesis: Offering one- to two-hour scheduled delivery windows (with smart capacity gating) raises conversion and AOV for planned baskets without pushing lateness past guardrails.
- Target segment: ZIP codes within two miles of participating stores; users browsing more than five minutes or viewing more than six products (planning signals); store-hours with spare picking capacity. Metrics
- Primary:
- Order conversion rate (sessions to orders).
- Average order value (AOV) and units per order (UPO).
- Secondary:
- New-to-grocery buyer rate (first grocery order).
- 28-day grocery repeat rate.
- Guardrails:
- On-time-in-window rate (OTW) at or above .
- Support contacts per 1,000 orders at or below .
- Courier and picker utilization within safe bounds.
- Leading indicators:
- Click-through on window selection and the spread of demand across windows.
- Add-to-cart after a window is selected. Scrappy, low-risk experiment
- Treatment: Show one- to two-hour scheduled windows (including off-peak incentives such as $0.99 delivery for low-demand windows) on eligible store-hour inventory.
- Control: ASAP only, the status quo.
- Randomization unit: Store-hour switchback — each participating store toggles treatment and control in two-hour blocks. This respects capacity and avoids cross-user interference.
- Assignment: Blocks split 50/50, stratified by day of week and peak versus off-peak. A capacity gate prevents offering windows when pickers or couriers are constrained.
- Duration: Two to three weeks to cover weekly cycles.
- Power note (example): If baseline conversion is 7% and we expect a gain of 0.7 pp (10% relative), with pooled standard deviation at the session level, roughly 60,000–80,000 sessions per arm are needed; that is feasible across multiple stores. Tiny numeric example
- Baseline: 7% conversion, AOV $48, OTW 92%.
- Target lift: conversion to 7.7%, AOV to $52, holding .
- Contribution margin: if margin per order is $6 at baseline, an extra $4 of AOV at 25% gross margin adds about $1, which easily covers a $0.50 window incentive used on under 30% of orders.
B) Efficiency bet — Pre-cart OOS prediction with proactive substitutes
- Hypothesis: Predicting likely out-of-stock items and recommending in-stock substitutes up front reduces pick and pack time, cancellations, and refunds while preserving conversion.
- Target segment: Items with high out-of-stock probability at participating stores; stores with historical OOS labels or API feeds. Metrics
- Primary:
- Item fill rate, computed as .
- Order-level cancellation rate.
- Secondary:
- Substitution acceptance rate (customer- or picker-approved).
- Picker minutes per item and total pick time per order.
- Refund dollars per order.
- Guardrails:
- Session-to-order conversion no worse than below baseline.
- AOV no worse than below baseline.
- Leading indicators:
- Click-through on suggested substitutes.
- Share of high-risk items avoided (grayed out or hidden) before add-to-cart. Scrappy, low-risk experiment
- Treatment: For items with predicted OOS at or above (e.g., a 0.8 precision threshold per store-SKU-hour), gray out or de-rank the item and show one or two in-stock substitutes with clear labels ("In stock; similar brand").
- Control: The current experience, with no pre-cart OOS treatment.
- Randomization unit: Store-level switchback by day (whole catalog treated versus not), to avoid item-level interference and simplify ops. Alternatively, item-level within a store for only the top 100 high-risk SKUs if traffic is limited.
- Assignment: 50/50 days per store; stratify by weekday and weekend.
- Duration: Two to four weeks.
- Safety valve: Start with a small share of traffic (e.g., 20%) and a high-precision threshold to minimize false positives; ramp after three days if guardrails hold. Tiny numeric example
- Baseline item fill rate: 92%, cancellations 3.0%.
- Target impact: fill rate up 2 pp to 94%, cancellations down 0.5 pp.
- Economics: If refunds average $1.20 per order and picker time is 18 minutes per order at $0.35 per minute, a one-minute saving per order plus $0.30 fewer refunds yields roughly $0.65 of margin improvement per order.
3) Top risks and mitigations
- Prediction errors harm conversion (false OOS) or trust
- Mitigations: Use high-precision thresholds; limit treatment to SKUs with strong signal; make "see alternatives" prominent and overriding easy; monitor continuously with fast rollback.
- Capacity and SLA breaches from scheduled windows
- Mitigations: Capacity gating by store-hour; conservative initial quotas; dynamic throttling; guardrail monitoring and auto-disable.
- Partner or store pushback (perceived cannibalization or catalog suppression)
- Mitigations: Share store-level results and net sales; run an opt-in pilot; exclude key SKUs from treatment initially; co-design substitutes with merchants.
4) How grocery experimentation differs from restaurants
- Randomization unit: Grocery often needs store-level or store-hour switchbacks (capacity, picker ops), versus user-level for restaurants.
- Catalog volatility: SKUs and availability change intra-day; define intent-to-treat and track non-compliance when items go out of stock.
- Multi-stage fulfillment: Shopper picking introduces extra latency and variability; measure pick-time KPIs and control for picker shifts.
- Interference risk: Couriers and pickers shared across arms can cause spillovers; cluster by store and time blocks.
- Longer decision cycles: Planned baskets span days; longer experiments are needed to capture repeat behavior and weekend effects.
- Heavier tails: AOV and order times have high variance; consider nonparametric estimators, CUPED, or hierarchical models.
- Compliance constraints: Age-restricted and EBT flows limit randomization; segment them or exclude them from tests.
5) Product and operational knowledge needed up front
- Inventory integration: Feed latency and coverage, API reliability, which stores have real-time versus batch updates, historical OOS and error taxonomies.
- Picking model: Who picks (in-store staff versus platform shoppers), training, typical pick times, barcode scanning coverage, aisle maps.
- Capacity model: Picker and courier availability by hour and zone; batching rules; SLA definitions (promised window, handoff times).
- Catalog and merchandising: Top SKUs by store, substitution maps, private-label constraints, restricted items.
- Financials: Fee structure, commission and markup, incentives, retail-media revenue, picking cost per minute, refund liability.
- Compliance and policy: ID checks, alcohol, EBT/SNAP eligibility, building-access norms in NYC (doormen, walk-ups), cold-chain requirements.
- Experimentation platform: Ability to randomize at store or store-hour level, support switchbacks, log item-level events (add-to-cart, OOS, subs), guardrail alerting and kill switches.
- Baselines: Current conversion, AOV, fill rate, cancellations, OTW, picker times, refund rates by store and time of day.
Appendix: Key formulas and definitions
- Fill rate: .
- Substitution rate: ; acceptance rate: .
- On-time-in-window (OTW): .
- Contribution margin per order: .
- Picker time per item: . Together, these bets and designs let you validate demand for planned grocery missions while improving fulfillment efficiency, with low operational risk and clear guardrails.