Meta · ML System Design
Choose and compute recommender evaluation metrics
TrueInterview
October 7, 2026 · 1 min read
You are building a restaurant recommendation system whose logistic regression model outputs , the probability that a user will engage with a suggested restaurant. For the time being, A/B tests and user surveys are not available.
-
Offline evaluation design: Suggest a sound offline procedure for comparing two models M0 and M1 when no A/B test is possible, for example counterfactual evaluation using inverse propensity scores derived from the historical logging policy. State the precise metric or metrics you would calculate, such as PR-AUC, Precision@K, or calibrated Brier score, explain how you would prevent leakage, and describe how you would select K. Identify one failure mode of IPS when propensity values are very small, along with a mitigation.
-
Thresholded metrics: On a holdout set of 1,000 recommendations evaluated at threshold 0.7, you observe TP=120, FP=30, TN=820, and FN=30. Calculate precision, recall, specificity, F1, and accuracy. Discuss why accuracy may be misleading in this setting, and identify the metric that best matches the objective that every displayed item should be relevant.
-
Calibration: Explain how you would assess and correct probability calibration, for instance with reliability diagrams or Platt scaling versus isotonic regression. Why is good calibration important when defining business rules such as only showing an item if its score is at least ?
-
Model choice: Defend the use of logistic regression instead of more complex models in this scenario, and name two failure modes—such as multicollinearity among features or class imbalance—along with concrete remedies.
-
Network effects: If features based on friend activity create interference, which offline split strategy best reduces leakage: time-based, user-disjoint, or graph-clustered splits? Describe the trade-offs.
Overview: This question tests skill in offline evaluation for recommender systems, propensity-weighted policy comparison, interpretation of classification metrics, probability calibration, model selection trade-offs, and management of data leakage and network interference.