Google · Statistics & Data Analysis
Diagnose a metric drop in search time
TrueInterview
October 7, 2026 · 2 min read
Across the last three calendar months, the metric 'searching time per user per session' fell by 35%. A teammate suggests modeling two distributions: T1 = time before the first successful result and T2 = time before giving up. Critique this proposal and design a robust analysis to find the root cause.
- Define exactly what the metric and population mean; address multi-tab sessions, background inactivity, and timeouts. Specify inclusion and exclusion rules.
- Discuss the biases caused by splitting the problem into T1 and T2: selection bias (omitting sessions with no success), right-censoring, competing risks (success versus abandonment), and left-truncation. How would you detect and correct them?
- Choose methods (for example, survival/hazard models with censoring; mixture models) and show how to estimate hazards and compare them across cohorts such as device, locale, and query type.
- Exclude non-product explanations: instrumentation changes, seasonality, traffic mix shifts, bot filtering, and release flags. List concrete checks and guardrail metrics.
- Build a time-series decomposition and change-point analysis; state the covariates and counterfactual baselines.
- Propose a minimal experiment or holdout (such as rolling back a ranking feature) with success criteria and expected directional outcomes. Overview: This question assesses a data scientist's ability to define metrics and populations precisely; to use time-to-event and survival/hazard modeling; to perform time-series decomposition and change-point analysis; to investigate instrumentation and traffic; and to design experiments for root-cause identification. Community answers Answer by SS
- First: Critique the T1 / T2 idea Your teammate proposes: T1: time until first success T2: time until the user gives up Problem with this approach It sounds intuitive, but it is biased: Selection bias: T1 only includes sessions where users succeeded → it ignores failures Right-censoring: Some sessions are still in progress → we never observe the final outcome Competing risks: A user can either succeed OR abandon → the two paths are not independent Left-truncation: If tracking starts late (e.g., page reload), we miss early behavior Conclusion: Splitting into T1 and T2 loses information and produces misleading results.
- Define the metric properly (very important) Metric: "Searching time per user per session = elapsed time from first query until either success or exit" Define: Start: first search query End: success (click a meaningful result) OR abandonment (exit / inactivity) Handle edge cases: Multi-tab sessions: Treat each tab independently OR merge them when they fall within the same session window Background inactivity: Remove idle time (e.g., no activity > 30 sec) Timeouts: Cap sessions (e.g., maximum 30 mins) Inclusion rules: Include sessions with ≥1 search Exclude: bot traffic broken/incomplete logs extremely long idle sessions
- What could cause the 35% drop? Before modeling, rule out non-product causes: Data / logging checks Did the tracking definition change? Are events missing or delayed? Did the timeout logic change? Traffic mix changes More mobile users? More easy queries? Different geographies? Seasonality Holidays → simpler queries? Bots / filtering
Loading comments…