Snowflake · Statistics & Data Analysis
Decide and justify product metrics amid trade-offs
TrueInterview
October 7, 2026 · 2 min read
You are rolling out a new 'Smart Sort' ranking for a content feed; it should improve relevance, but it could lower short-term ad impressions. Select one primary success metric and several guardrail metrics, then lay out an experiment and a decision framework for situations where metrics move in opposite directions.
Sub-questions:
- Metric choice: From the candidate set {7-day retention, session minutes per DAU, revenue per DAU, CTR, creator supply health}, choose one primary metric and two to three guardrails. Justify each selection using statistical characteristics such as variance, sensitivity to bots or heavy tails, weekday stability, and exposure to Simpson’s paradox. Give exact formulas and aggregation levels (per-user versus per-session), and state whether winsorization or log transformations should be applied.
- Experiment design: Describe the A/B design, covering the randomization unit (user versus session), likely interference risks from ranking changes, mitigation approaches such as sticky bucketing or ghost ads, and how long the test should run. Calculate the sample size needed for an MDE of 1.0% relative change in the primary metric with 90% power and ; list the required inputs and explain how CUPED or stratification would lower variance. Assume the baseline window is 2025-08-18 through 2025-08-31, and the test must end by 2025-09-01.
- Decision framework: If the primary metric rises by +0.8% while a guardrail such as creator payout per impression falls by −2.5% with , describe a principled trade-off approach—for example, an LTV delta combining engagement and revenue, constrained optimization, or multi-metric decision rules. Also cover how you would handle novelty effects, ramp plans, sequential monitoring, and heterogeneous treatment effects by country.
- Post-launch: Specify the on-call dashboards and alert thresholds you would put in place for the first week after a 10% rollout.
Overview: This question tests skill in product metric selection, statistical experiment design, causal inference, and post-launch monitoring for feed-ranking trade-offs; it assesses competencies in metrics engineering, A/B testing, variance reduction, and decision-framework reasoning in the Analytics & Experimentation area for a Data Scientist role. It is often used to evaluate whether a candidate can justify metric choices under competing business goals, design defensible randomized experiments, and define monitoring and alerting strategies; it checks both practical application—hands-on experiment setup and monitoring—and conceptual understanding of statistical properties, bias, heterogeneity, and trade-off reasoning.