Google · ML System Design
Build and evaluate a full ML pipeline
TrueInterview
October 7, 2026 · 1 min read
You need to output two things: (1) the chance a user spends more than $0 within the next 7 days (a classification task), and (2) the amount they are expected to spend in that window (a regression task). The training set contains events and orders through 2025-08-31, while predictions begin on 2025-09-01. Lay out a complete pipeline covering feature creation (with time-windowed aggregates), leakage safeguards (such as dropping post-cutoff signals like refund_time), time-based cross-validation, class imbalance handling, and model selection for each of the two tasks. Define the metrics (for example PR-AUC, calibrated Brier score, pinball loss for quantiles), describe a calibration approach, and explain how you would choose a threshold when the cost matrix is asymmetric. Explain how you would spot and fix regressions that affect only certain segments, pick and defend an offline/online evaluation strategy (including rollout and holdbacks), and set up monitoring after deployment for drift, label delay, and model decay. Finish with two concrete feature examples that are predictive but leakage-prone, and show how you would redefine them safely.
Overview: This question tests whether a candidate can design and run machine learning pipelines from start to finish, touching feature engineering, leakage prevention, temporal cross-validation, evaluation and calibration, thresholding with asymmetric costs, deployment rollouts, and post-deployment monitoring.