Stripe · ML System Design
Design a leak-free time-split model
TrueInterview
October 7, 2026 · 2 min read
You have one calendar week, with an expected workload of about 4–6 hours, to prepare slides and runnable code that estimate, for every active user, the chance that the user completes a purchase during a fixed 30-day future window.
Requirements:
- Use
snapshot_ts = 2025-10-01 00:00:00 UTC. A positive label means an order is placed in the half-open interval[snapshot_ts, snapshot_ts + 30 days). - Spell out the exact steps that prevent label leakage. Cover records that are late-arriving (
arrival_timeafterevent_time) and every feature that must not peek at data aftersnapshot_ts. - Present a lightweight but strong baseline and a primary model. Compare logistic regression with monotonic or regularized features against gradient-boosted trees, justify your final choice, and describe the first five features you would engineer.
- Select the main threshold-independent evaluation metric for the expected class imbalance—for example PR-AUC versus ROC-AUC—and a separate business metric for the slide deck; explain trade-offs.
- Design an honest temporal validation strategy, such as rolling-origin or blocked cross-validation, and explain how to tune hyperparameters quickly while limiting overfitting within the available time.
- Describe how you will detect and correct data leakage, target leakage, and train-test contamination; include at least two concrete checks you would implement in code.
- Outline a calibration approach, such as Platt scaling versus isotonic regression, choose a decision threshold for an email targeting campaign with a cost per message, and explain how the slides will communicate calibration quality and expected business impact.
- Provide a minimal ablation plan that can be executed within the week; state which components you would remove first if time is short, and give the exact headline text for 5–7 slides that form a coherent narrative for a hiring manager.
Example 1:
Input:
snapshot_ts = 2025-10-01 00:00:00 UTC
event_time = 2025-10-11 14:00:00 UTC
arrival_time = 2025-10-13 08:00:00 UTC
Output: label = 1
Explanation: The order was placed inside the 30-day window; the record arrived two days later, so the labeling logic must rely on event_time, not arrival_time.
Example 2:
Input:
snapshot_ts = 2025-10-01 00:00:00 UTC
event_time = 2025-10-29 23:00:00 UTC
arrival_time = 2025-11-02 10:00:00 UTC
Output: label = 1
Explanation: The purchase occurred before the window closed even though the record was received after the window; late-arriving data should still be counted.
Example 3:
Input:
snapshot_ts = 2025-10-01 00:00:00 UTC
event_time = 2025-10-31 00:00:00 UTC
Output: label = 0
Explanation: The interval is half-open, so an event exactly at snapshot_ts + 30 days is excluded.
Constraints:
- Expected total effort: about 4–6 hours across one week.
- Deliverables: slides plus runnable code.
- The prediction is a probability for each active user, not just a binary label.
- You must justify metric, validation, calibration, and ablation choices.
- Every feature must be computable using only information available at
snapshot_ts.