Snowflake · Statistics & Data Analysis
Design an A/B test for ML model launch
TrueInterview
October 7, 2026 · 1 min read
You are swapping the existing ranker for a new model in a feed. The baseline CTR is 2.0%. You anticipate a +5% relative improvement in CTR and require 90% power at α=0.05. Eligible traffic is 1,000,000 users per day; assignment is user-level 50/50 and remains stable over time. A) Calculate the necessary sample size and shortest test duration with a two-proportion power analysis; list all assumptions (variance, independence, no interference). Show your formulas and reasoning; adjust for an expected 5% bot/invalid traffic rate and a possible 1% sample ratio mismatch. B) Define guardrail metrics (bounce rate, crashes, latency p95, revenue per user) and decision thresholds. Explain sequential monitoring with α-spending (for example, Pocock or O’Brien–Fleming) so the test can stop early without raising Type I error. C) Handle novelty and day-of-week effects: propose CUPED (covariate adjustment) based on pre-experiment user CTR; state the exact covariate and how you would confirm variance reduction without introducing bias. D) Reduce interference and contamination: keep user-level bucketing, avoid cross-arm content spillover, and prepare a geo or holdout design if network effects are suspected. E) After the test, describe checks for realized power, heterogeneity of treatment effects across cohorts, and how you would decide to ramp to 100%. Overview: This question tests a data scientist's ability in experimental design, statistical power analysis, sequential monitoring, covariate adjustment (such as CUPED), and practical operational issues including guardrails, interference, and ramping for ML model launches.