DoorDash · Statistics & Data Analysis
Define and compute retention and churn precisely
TrueInterview
October 7, 2026 · 1 min read
Define retention and churn for a transactional consumer app, and show how you would calculate each correctly:
- Pick precise definitions for cohorts (signup versus first purchase), activity (active when someone places at least one order in the period), retention types (N-day, week N, rolling, bracket), and churn (no activity for K consecutive periods). Justify the choices you make based on the decision use-case.
- Give formulas for cohort retention and churn rates that use the correct risk sets, and address right-censoring and delayed conversion. Explain pitfalls such as survivorship bias, Simpson’s paradox, and seasonality.
- Explain how to measure the long-term retention effect of a treatment, such as a 20% discount, through survival analysis: define time-to-churn, hazard, and cumulative incidence; state how you would compare the curves, for example with log-rank or stratified tests, and how you would adjust for covariates.
- Show how rolling retention can diverge from strict cohort retention and how you would reconcile the two for executives. Include an example with invented numbers to show the difference and compute both correctly.
- Explain how you would choose windows — washout, observation, and attribution — and how those choices affect experiment power and bias.
Overview: This question tests a data scientist’s skill in statistically measuring user retention and churn. It covers cohort definition, activity rules, risk-set-aware retention and churn formulas, handling censoring and delayed conversion, and survival-analysis concepts including time-to-churn, hazard, and cumulative incidence.
Loading comments…