Google · Statistics & Data Analysis
Design tests to measure latency impact
TrueInterview
October 7, 2026 · 2 min read
Question
Imagine you are a data scientist working on a large consumer product such as YouTube. The engineering team is rolling out a change that should cut client-side, or video-start, latency by about 100 ms for a subset of users, although the new code path could raise the error rate and alter buffering. Outline and evaluate the measurement plan for this change.
- Experiment design (latency to business impact). Propose an A/B test to measure the causal effect of lower latency on engagement. Specify:
- A hypothesis and a single primary decision metric, including why that metric is sensitive to latency and how you weigh sensitivity against business relevance.
- Diagnostic metrics that help pinpoint where any movement originates, such as funnel steps or latency percentiles.
- Guardrail metrics covering quality, reliability, and revenue so a regression is not shipped—especially playback error rate and rebuffering, because the new stack may be unstable.
- Randomization unit and interference. Select among user-, device-, and session/request-level randomization and defend your choice. Describe how you prevent contamination and interference, including sticky bucketing, CDN/cache routing, and cross-device spillover.
- Power, MDE, and duration. Explain how you would estimate the required sample size and experiment length: which inputs you need, how you choose the minimum detectable effect, and how you account for heavy-tailed watch time. Mention the main variance drivers, such as user heterogeneity, seasonality, and outliers.
- Variance reduction. Provide at least two methods for lowering variance or improving sensitivity, and describe when each is suitable—for example, CUPED, stratification, triggering, winsorization, or clustered standard errors.
- Analysis plan and conflicting movements. State the estimator you would use, how you manage multiple metrics and heterogeneous effects such as WiFi versus cellular, and how you resolve cases where results move in opposite directions—for instance, watch time rises while the error guardrail also rises.
- Diagnosing a change in a ratio metric. Suppose leadership follows a ratio metric such as or , and it moved by +0.3%. Lay out a structured way to diagnose why it changed, break the movement into numerator and denominator contributions, and protect against misleading readings such as Simpson's paradox.
- When randomization is not possible: propensity score matching. Now suppose you cannot randomize the latency change—for example, it was rolled out selectively because of infrastructure constraints—and you only observe that some users had lower latency than others. Describe how you would apply propensity score matching (PSM) to estimate the impact, state the assumptions PSM requires, and explain how you would validate or sensitivity-test those assumptions.
Assumptions
- Users are global; traffic changes by time of day and day of week, and latency effects may differ by network type, such as WiFi versus cellular.
- A pre-period can be defined for computing baselines or covariates.
- Logging exists for latency, exposure, errors, buffering, and the main engagement outcomes.
- Use one consistent reporting timezone, such as UTC, for daily metrics to avoid boundary artifacts.
Overview: A Google data scientist onsite case about measuring the causal effect of lower YouTube video-start latency on engagement. It covers A/B test design, including the primary metric, diagnostics, and error/buffering guardrails; randomization unit and interference; power and MDE for heavy-tailed watch time; variance reduction through methods such as CUPED and stratification; analysis of conflicting metric movements; ratio-metric decomposition with Simpson's paradox; and propensity score matching when randomization is not possible.