Cvs Health · ML System Design
Build an uplift model for targeting
TrueInterview
October 7, 2026 · 1 min read
You are given last season's campaign logs, which include randomized holdout groups. Design a treatment effect modeling approach for deciding which customers to contact by SMS or email in the upcoming flu-shot campaign.
Data available
- Features: demographic information, previous visits, vaccination history, engagement signals (opens/clicks), distance to a store, and appointment history.
- Labels: indicates vaccination within 30 days; treatments ; randomized assignment with known probabilities; exposure indicators (delivered/opened).
- Costs: SMS costs $0.02 and email costs $0.001; the budget permits reaching at most 40% of eligible customers.
Tasks
- Modeling
- Choose and justify one approach: separate response models with two-model uplift, direct uplift methods (for example, meta-learners such as T-learner, S-learner, or DR-learner), or multiclass treatment modeling. Address leakage from post-treatment features, class imbalance, and calibration.
- Evaluation
- Define offline evaluation using uplift/Qini curves and AUUC; compute incremental ROI after accounting for channel costs; apply policy evaluation with inverse propensity weighting (IPW) or doubly robust estimators.
- Policy
- With a budget that allows contacting up to 40% of eligible customers, describe how to rank customers by predicted incremental effect and select each customer's channel (for example, argmax over channel-specific uplift minus cost). Explain guardrails such as do-not-contact lists and fairness across age or state.
- Online validation
- Propose a gated rollout test that compares model-based targeting with uniform random targeting. Define success metrics and stopping rules.
- Diagnostics
- Show how you would identify segments with harmful persuasion, meaning negative uplift, and how you would handle them in targeting.
Overview: This question tests a data scientist's ability to carry out treatment-effect and uplift modeling for multi-armed marketing interventions, including causal inference with randomized holdouts and propensity scoring, cost-sensitive channel selection, calibration and class-imbalance handling, policy evaluation, and diagnostic and operational safeguards in the Machine Learning domain. It is often asked because it examines practical use of causal ML and decision-policy design—requiring model selection, offline and online evaluation under budget and fairness constraints—and is mainly a practical application task with significant conceptual causal-inference elements.