Stripe · Statistics & Data Analysis
Evaluate a new product with experimentation
TrueInterview
October 7, 2026 · 1 min read
Introducing a recommendation module can create cross-user interference and seasonal traffic effects. Outline how you would evaluate it. (1) Specify an Overall Evaluation Criterion (OEC) for a commerce app plus three guardrail metrics (such as churn, p95 latency, and complaint rate), including exact formulas and units. (2) Pick a test design—user-level RCT, geo-cluster, or time-based switchback—and defend the choice with respect to interference, non-stationarity, and operational constraints. (3) Describe the ramp strategy and pre-registration plan: stopping rules, power target, variance reduction (CUPED or covariate adjustment), and small-area risk controls. (4) If randomization cannot be done, propose a quasi-experimental fallback such as synthetic control or difference-in-differences, and list the assumptions and falsification tests you would run. (5) Suppose partway through the test the OEC is flat while add-to-cart rises and conversion falls; provide a metric-debugging checklist and the exact cuts you would request to isolate the issue (for example, device, geography, new vs returning users, latency buckets). Be specific and include equations where relevant.
Overview: This question assesses experimental design and causal inference: defining an Overall Evaluation Criterion (OEC) and guardrail metrics, choosing a test design and ramp/power strategy, specifying quasi-experimental fallbacks, and running metric diagnostics for a recommendation module in a commerce app.