Uber · ML System Design
Select the better $5 promo-targeting model
TrueInterview
October 7, 2026 · 1 min read
Suppose you have two models that score users for a $5 coupon: M0, the current model, and M1, a new one. Each outputs . You can send no more than B promotions per day. Define a decision rule that each day targets the top K users while keeping expected spend at or below B. a) Specify a profit-aligned success metric, for example , and explain why AUC or accuracy can mislead when evaluating targeting. b) Given a historical randomized dataset with columns {user_id, features, assigned_treatment ∈ {coupon, control}, outcome redeem ∈ {0,1}, gmv, timestamp}, derive off-policy estimators for comparing the policies induced by M0 and M1: inverse propensity scoring (IPS), self-normalized IPS (SNIPS), and doubly robust (DR). Write out the formulas, state the assumptions required for unbiasedness, and discuss variance trade-offs and cross-fitting. c) Describe how you would calibrate the predicted probabilities (for instance with isotonic regression or Platt scaling), choose a daily threshold that respects budget B as conditions drift, and directly optimize expected profit under guardrails such as opt-out rate and complaint rate. d) List three concrete leakage risks—for example, features that capture prior coupon exposure, post-treatment variables, or proxies for future engagement—and explain how you would detect or prevent each. e) Explain how to handle delayed redemption labels and per-user redemption caps in both training and evaluation so they do not introduce bias. f) Outline a monitoring plan for non-stationarity and cold-start users, including shadow deployment and canarying.
Overview: This question tests proficiency in policy evaluation, off-policy estimation, probability calibration, leakage detection, delayed-label handling, and production monitoring for budget-constrained coupon targeting; it falls within the Machine Learning and Data Science domain.