Expedia · ML System Design
Validate and monitor ranking model end-to-end
TrueInterview
October 7, 2026 · 1 min read
Using the same Expedia hotel-ranking model: a) Design an offline evaluation approach that avoids leakage through time-based splits and user-level grouping, corrects for position bias with propensity-weighted/IPS metrics or counterfactual learning-to-rank, and reports calibration of conversion probability estimates, such as reliability curves or expected calibration error (ECE). b) State the ranking metrics you would report—for example, revenue-weighted NDCG@10 or ERR—along with their formulas and an explanation of why they capture client value. c) Describe diagnostic methods for finding the main drivers without exposing target information, such as SHAP with an appropriate background dataset, permutation checks, and stability across cross-validation folds. d) Specify the rollout and monitoring plan: shadow mode, canary releases, guardrail alarms, data and concept drift detection, handling of late-arriving data, and an automatic rollback policy with defined thresholds. e) Explain how you would check that optimizing a surrogate objective still lifts the client KPI, and what actions you would take if the offline–online relationship no longer holds.
Overview: This question tests a data scientist’s command of learning-to-rank concepts, offline evaluation and metric design, position-bias correction, diagnostic analysis, deployment safeguards, and alignment of surrogate objectives with business KPIs in hotel search ranking.