Shopify · Statistics & Data Analysis
Design robust experiment for ambiguous core change
TrueInterview
October 7, 2026 · 1 min read
Suppose you need to assess a central product change whose impact is likely subject to network effects—for instance, adjusting matchmaking in a large online game with 8 million daily active users. State the main success metric and guardrail metrics—such as D1/D7 retention, ARPDAU, and crash rate—select the unit of randomization (user, session, or cluster), and defend that choice given the risk of interference. Lay out a complete test plan covering pre-registration, ramp approach, stopping rules (sequential or alpha-spending), power and minimum detectable effect targets, and expected duration. Describe your variance reduction method (for example CUPED using pre-period engagement), treatment of outliers, tests for novelty decay, and diagnostics for spillover. Calculate the necessary sample size per variant when baseline D1 retention is 40% and the target is a +1.0 percentage point absolute increase, at and power = 0.80; include the formula and assumptions you used. Explain how you will detect heterogeneous treatment effects across cohorts such as geography, payer status, and device; how you will control multiple comparisons (e.g., false discovery rate); and your fallback plan if randomization cannot be used—for instance, difference-in-differences with parallel trends checks. Finally, specify clear ship and rollback thresholds, data quality service-level objectives, and how findings will be shared asynchronously with stakeholders in a remote-first setting.
Overview: This question tests a data scientist's skills in experimental design and causal inference, covering success metrics and guardrails, randomization unit selection under interference, sample-size and power calculations, variance reduction, detection of heterogeneous treatment effects, multiple-testing control, and operational data-quality and rollout criteria. It appears often in Analytics & Experimentation interviews because companies need to defend randomized evaluations for features with network effects and operational limitations; the domain is Analytics & Experimentation, and the difficulty is mainly applied work that rests on conceptual statistical knowledge.