Uber · ML System Design
Build and assess CTR prediction
TrueInterview
October 7, 2026 · 2 min read
Your task is to estimate the likelihood that an ad impression gets a click within 24 hours. The positive rate is roughly 0.7%. Available features are user_age, device_type, locale, time_of_day, ad_id (high-cardinality), campaign_id, past_7d_impressions, past_7d_clicks, and referrer. Labels are delayed, since some clicks only appear up to 24 hours afterward.
-
Modeling: Suggest two model families that can handle extreme class imbalance along with sparse or high-cardinality features. How would you encode ad_id and campaign_id without leaking information? Explain a time-based cross-validation design that accounts for label delay.
-
Imbalance: Contrast class weighting, focal loss, under-sampling, and calibrated thresholding. In what situations would you steer clear of synthetic oversampling? Support your answer by describing the likely impact on ranking versus calibration.
-
Evaluation: Model A reaches ROC-AUC=0.91 and PR-AUC=0.14; Model B reaches ROC-AUC=0.88 and PR-AUC=0.22. Explain why these metrics can disagree when prevalence is only 0.7%, state which one you would trust for email/ad CTR, and describe how to pick operating thresholds for different business goals with a cost matrix that weighs missed clicks against wasted impressions.
-
Calibration and thresholds: Explain how you would evaluate and improve calibration—for example, isotonic regression versus Platt scaling—and how you would choose thresholds to (a) maximize F1 and (b) maximize expected profit. How would you calculate precision@top1% and use it to compare models?
-
Online validation: Outline a bucket test for confirming lift from the model’s scores, such as top-k targeting. Which logs are needed to detect covariate drift and label delay once the model is live, and what steps would you take to protect against feedback loops?
Overview: This question tests predictive modeling and applied data science skills for CTR prediction, spanning extreme class imbalance, delayed feedback, sparse or high-cardinality feature encoding, time-aware validation, evaluation and calibration of probabilistic scores, and online A/B validation; it sits firmly in the Machine Learning domain and checks both conceptual understanding and hands-on application. It is frequently asked because it probes reasoning about real-world production concerns—ROC versus PR metric choice, thresholding under business costs, calibration approaches, drift detection, and avoiding feedback loops—without needing specific implementation details.
See the complete Uber Data Scientist interview account from which this question was taken.