ByteDance · Statistics & Data Analysis
Design robust metrics for a feature launch
TrueInterview
October 7, 2026 · 1 min read
You're rolling out a new in-app Quick Reply feature for a messaging app and need to specify metrics and guardrails for a one-week A/B test scheduled from 2025-08-25 through 2025-09-01. Be exact about denominators, units of analysis, and attribution windows.
-
Define the primary success metric (north-star) that represents meaningful user value delivered by Quick Reply. Give the precise formula, including the unit of analysis (user or user-day), numerator, denominator, inclusion criteria such as the exposure definition, and a 24-hour attribution rule from click to reply send.
-
Propose at least two guardrail metrics that defend long-term health—for example, reply quality or abuse rate, churn, or app crashes. For each one, specify the measurement unit, formula, and acceptable movement thresholds.
-
Suppose the overall reply send rate goes up, while average conversation length falls and the complaint rate among Ads-acquired users rises. Explain how you would segment and interpret these metrics to avoid Simpson’s paradox, and how you would decide whether to ship, hold, or iterate.
-
Define a metric that would catch “empty engagement” — for example, accidental taps or replies that get deleted before sending. Describe how to implement it with existing telemetry and how it would affect the launch decision.
-
Lay out a pre-analysis plan: how you will pre-register primary and secondary metrics, set stopping rules, and prevent metric fishing. Include how you would confirm that “exposed” really means the user saw the Quick Reply entry point, not merely that they were eligible.
Overview: This question assesses a data scientist’s ability to define sound A/B test metrics and guardrails, specify exact units of analysis, denominators, exposure and attribution windows, detect empty or accidental engagement through telemetry, and pre-register a valid analysis plan.