Instacart · Product & Business Case
Lead a zero-to-one initiative effectively
TrueInterview
October 7, 2026 · 8 min read
Walk through how you would move an ambiguous 'improve shopper retention' directive from concept to shipped product. Specify the problem statement, success metrics and guardrails, discovery plan, PRD outline, stakeholder map, milestones, and kill criteria. Describe how you would reduce risk with a prototype, secure resources, coordinate change across CX/Legal/Sales, and conduct a post-launch review. Add a 30/60/90-day plan and one concrete example of a difficult trade-off you would accept.
Overview This question tests whether a data-driven leader can turn an unclear directive into a measurable program, assessing problem framing, metric design, discovery and experiment planning, stakeholder alignment, prioritization, and trade-off judgment in a two-sided marketplace.
Solution
1) Problem Statement
- Goal: Raise shopper retention in a manner that durably strengthens marketplace health and unit economics without damaging customer experience, partner relationships, or compliance.
- Assumptions (stated explicitly):
- A 'shopper' is an independent contractor who completes orders.
- Retention is tracked at several horizons (D7, D30, D60, D90), with cohorts defined by shopper start date.
- Marketplace health is driven by fill rate, on-time delivery, shopper earnings per hour, and customer NPS.
- Proposed v1 scope: Concentrate on early-lifecycle retention (the first 30–60 days), where the churn hazard is usually greatest.
2) Success Metrics and Guardrails
- North star
- 60-day retention of new shoppers (): the share of a start-month cohort that completes at least one batch during days 31–60.
- Formal definition: is the survival probability to day . Target: percentage points within two quarters.
- Leading indicators
- D7 activation rate: the proportion completing at least one batch in the first 7 days.
- Time-to-first-batch (TTFB): median hours from onboarding completion to the first batch.
- Early hours worked: median online hours during week 1.
- Early experience quality: cancellations and support contacts within the first 10 batches.
- Economic metrics
- Incremental LTV per shopper (ILTV) = .
- Cost per retained shopper (CPRS) = (incentive + ops cost) / incremental retained shoppers.
- Marketplace health guardrails (must be no worse than these thresholds relative to control):
- Order fill rate: pp.
- On-time delivery: pp.
- Customer NPS/CSAT: pt.
- Shopper earnings per hour: baseline (no reduction in p50 or p25).
- Support contacts per order: .
- Compliance: pay transparency, classification, and disclosure requirements are satisfied.
- Budget guardrail: CPRS $150 (example) with ROI at 6 months.
Small numeric example
- Baseline: 10,000 new shoppers per month; , so 3,500 are retained at D60.
- Target: 39%, so 3,900 retained; that is +400 incremental retained shoppers.
- If each retained shopper completes 40 orders over the next 60 days at $1 margin per order, that yields +$16k gross margin over 60 days.
- If incentives and ops cost total $60k, short-term ROI is below 1, but over 6–12 months ILTV may justify it; set CPRS and ROI thresholds to keep the program sustainable.
3) Discovery Plan
- Quantitative (weeks 1–4)
- Build cohort retention curves (D7/D30/D60/D90) plus survival and hazard analyses.
- Segment by cohort month, geography, tenure, device, earnings decile, time of day, batch type, cancellations, and support contacts.
- Use feature correlation or importance (for example, SHAP over a survival model) to find drivers such as TTFB, idle time, pay variability, batch distance, cancellation exposure, and support friction.
- Map the funnel from onboarding complete to first login to first batch to 5th batch to 20th batch, and locate the largest drop-offs.
- Run power analyses for experiments; estimate variance and minimum detectable effects.
- Qualitative (weeks 1–4)
- Conduct 1:1 interviews with 20–30 shoppers (new, ramped, churned or reactivated), plus ride-alongs or shadowing.
- Run a survey to quantify pain points such as earnings clarity, navigation, batching fairness, tip expectations, and payment timing.
- Tag support tickets for early-lifecycle themes.
- Competitor and policy review
- Benchmark payout cadence, sign-up bonuses, earnings transparency, guarantees, and compliance constraints.
- Synthesize hypotheses
- H1: Long TTFB and idle time are major drivers of early churn.
- H2: Earnings variability and weak transparency lower perceived fairness.
- H3: Early negative experiences (cancellations, complex substitutions, parking problems) create outsized churn risk.
- H4: Payment cadence, especially a slow first payout, weakens the reinforcement loop.
- Prioritize using RICE or impact versus effort, and run a pre-mortem asking how this could fail.
4) PRD Outline (for the initial wedge)
- Title: Fast Start for New Shoppers (v1)
- Background and problem
- Goals and non-goals
- Hypotheses
- Target users and segments
- User stories, for example: 'As a new shopper, I want my first batch quickly with predictable earnings.'
- Solution overview
- Example components: queue prioritization for the first 3 batches, an earnings guarantee on the first day, real-time guidance, and an accelerated first payout.
- Requirements
- Feature flags, geo eligibility, communications, and payments configuration.
- Experiment and rollout plan
- Geo-cluster randomization, holdouts, duration, sample size, and MDE.
- Metrics and telemetry
- Primary, leading, guardrails, and dashboards.
- Risks and mitigations
- Demand displacement, fairness perceptions, and legal concerns.
- Compliance and legal checklist
- Support and operations plan
- Analytics plan (analysis, HTE, CUPED or diff-in-diff as needed)
- Launch criteria and kill criteria
5) Stakeholder Map (RACI example)
- Responsible: PM (Shopper Experience), Data Scientist, Engineering Lead, Designer, and UXR.
- Accountable: GM or Director for Supply/Marketplace.
- Consulted: Shopper Ops, CX/Support, Trust & Safety, Legal/Compliance, Finance, Marketing/CRM, Dispatch/Matching, Payments, Data Engineering, and Partner/Sales.
- Informed: Executive sponsor, Regional Ops, and Partner Success.
6) Milestones and Kill/Gate Criteria
- Milestones
- Weeks 0–2: Finalize metric definitions; complete discovery; stack-rank proposals; choose the v1 wedge.
- Weeks 3–4: PRD v1, experiment design, instrumentation plan; secure resource and budget approvals.
- Weeks 5–8: Build plus internal dogfooding; launch a geo pilot behind feature flags.
- Weeks 9–14: Run the pilot for 6 weeks with weekly guardrail monitoring.
- Weeks 15–16: Readout; go/no-go; iterate or scale.
- Kill/Gate criteria (examples)
- Pre-pilot: No green light without legal signoff and guardrail dashboards.
- During pilot: Pause immediately if any guardrail breach persists for more than 48 hours (for example, on-time delivery down 1 pp, earnings per hour down $0.25 at p50) or if incident rates spike.
- End of pilot:
- Proceed if: pp (lower bound of the 95% CI is above +0.5 pp), CPRS $150, fill rate pp, and NPS .
- Iterate if: the effect is positive but below the ROI threshold; redesign and retest.
- Kill if: the effect is null or negative, or cost or guardrail violations occur; document and sunset.
7) De-risking with Prototype/MVP and Experimentation
- MVP concept: 'Fast Start'
- Queue prioritization for the first 3 batches to shorten TTFB.
- A first-week earnings guarantee (for example, $X for the first Y hours), reconciled at payout.
- Early guidance such as checklists, in-app tips, and a hotline for the first 5 batches.
- An accelerated payout for the first week.
- Low-lift prototypes
- Wizard-of-Oz operations: manually assign priority in a few geographies for 2 weeks.
- Off-app communications: SMS nudges, a help line, and survey prompts to validate messaging.
- A mock earnings transparency card (no ML yet) that shows a conservative expected earnings range.
- Experiment design
- Geo-cluster A/B with matched markets, or time-based randomized windows if interference risk is high.
- Use CUPED or pre-period covariate adjustment to reduce variance.
- Sample-size sketch: to detect +3 pp on S30 from a 50% baseline (two-sided, , power = 0.8), you need roughly 3,200–4,000 shoppers per arm; adjust for clustering and seasonality.
- Heterogeneity of treatment effects: tenure, geography, time of day, and demand elasticity.
- Interference checks: monitor control-geography displacement in fill rate.
8) Obtaining Resources and Budget
- One-pager or PRFAQ: problem, hypotheses, expected ROI, risks, plan, and asks.
- Business case
- Forecast ILTV uplift and CPRS under base, best, and worst scenarios.
- Show sensitivity to demand, seasonality, and incentive size.
- Asks
- People: 1 PM, 1 DS, 2–3 engineers, 0.5 DE, 0.5 designer, UXR support, plus ops coverage.
- Budget: an incentive pool for guarantees and payouts, plus UXR incentives.
- Platform: experiment infrastructure, feature flags, and dashboards.
- Governance: biweekly steering with GM/Legal/Finance; use pre-reads to speed up approvals.
9) Change Management (CX/Legal/Sales/Partners)
- CX/Support
- Update macros and FAQs; train agents on Fast Start rules and edge cases.
- Establish a real-time escalation path during the pilot; monitor contact reasons.
- Legal/Compliance
- Review incentive structures, disclosures, pay transparency, and classification sensitivities.
- Ensure clear terms and conditions and opt-in where needed; perform regional policy checks.
- Sales/Partners
- Communicate pilot geographies and the expected impact on fill rate and on-time delivery; align on service levels.
- Set expectations that there are no partner pricing changes; share guardrails and monitors.
- Shopper communications
- Use in-app banners, lifecycle emails, and SMS; explain how guarantees work without over-promising.
10) Post-Launch Review and Learning Plan
- Timing: interim review at 2–3 weeks; full readout at 6 weeks; final at 12 weeks.
- Contents
- Primary effects (S30/S60), leading indicators, guardrails, CPRS, and ILTV.
- Heterogeneity analysis, external validity, and seasonality effects.
- Operational learnings such as support load, fraud or abuse, and gaming of guarantees.
- Decision: scale, iterate, or sunset; revise the PRD for v2 if applicable.
- Documentation: write-up, dashboard links, and a decisions log; add to the playbook.
11) 30/60/90-Day Plan
- Days 0–30
- Align on definitions; build cohort and survival dashboards; run discovery (quantitative and qualitative).
- Select the v1 wedge (Fast Start); write the PRD and experiment design; complete legal review.
- Create the instrumentation plan; run power analysis; obtain resource and budget approvals.
- Days 31–60
- Build the MVP; set up feature flags and monitoring; train CX; finalize communications.
- Launch the geo pilot; hold weekly reviews; enforce guardrails.
- Start concurrent low-lift tests such as nudges and the transparency card to increase learning velocity.
- Days 61–90
- Complete the pilot; deep-dive results and heterogeneity of treatment effects; perform ROI analysis.
- Make the go/no-go call; create a scale plan or pivot to the next hypothesis (for example, payment cadence or batching fairness).
- Publish learnings; build the v2 backlog; align the roadmap for the next quarter.
12) Example Tough Trade-off
- Choice: a broad hourly incentive for all shoppers versus a targeted 'Fast Start' for new shoppers.
- A broad incentive likely boosts short-term supply but is expensive and may not improve long-term retention; it also risks fairness expectations and demand displacement.
- A targeted Fast Start focuses on the highest-hazard period with lower CPRS and a clearer causal link to retention, but it may create perceived inequity among veteran shoppers.
- Decision: choose targeted Fast Start to maximize ILTV and ROI, with mitigations:
- Communicate the purpose clearly, and add lightweight recognition for veterans (for example, loyalty badges or occasional targeted boosts) without undermining economics.
Common Pitfalls and Guardrails
- Goodhart's Law: do not optimize D7 at the expense of D60 or earnings per hour.
- Selection bias: use randomized or strong quasi-experimental designs, and avoid survivorship bias.
- Interference: use geo-level randomization to prevent spillovers in a shared marketplace.
- Seasonality and competitor actions: use matched-market controls and extend pilots across multiple periods.
- Compliance: incentives, disclosures, and communications must be legally vetted.
Optional Alternatives if RCT Is Hard
- Matched-market difference-in-differences with synthetic controls.
- Interrupted time series with guardrails and falsification tests.
- Instrumental variables (for example, exogenous weather shocks) for diagnostics only.
This plan converts an ambiguous mandate into a measurable, de-risked, cross-functionally executable path from idea to launch, with explicit metrics, safeguards, and decision gates.