Capital One · ML System Design
Build and evaluate airline delay prediction model
TrueInterview
October 7, 2026 · 8 min read
You are provided several CSVs for the well-known airline delay challenge, with columns such as flight_date, carrier, flight_num, origin, dest, sched_dep, sched_arr, dep_delay_min, arr_delay_min, distance, aircraft_type, weather_features_*, and holiday_flag. a) Define a binary target and justify it—for instance, late_arrival = arr_delay_min > 15. b) Describe a leakage-aware feature set: include weather forecasts at origin/dest, route history aggregates up to t−7 days, time-of-day, day-of-week, month, distance, carrier- and airport-level rolling statistics; exclude or properly lag any feature that encodes future information (e.g., actual arrival times). c) Specify a time-based split (e.g., train through 2024-06, validate 2024-07–2024-09, test 2024-10–2025-03), class imbalance handling, and primary metrics (PR-AUC, calibrated Brier). d) Compare a strong baseline (regularized logistic regression with target encoding) against gradient boosting (e.g., XGBoost/LightGBM): hyperparameters to search, early stopping, monotonic constraints if used. e) Explain how you would perform rolling-origin cross-validation and backtest threshold policies (e.g., proactive swaps or buffers) with cost-sensitive evaluation that prices false negatives at 5× false positives. f) Productionization: 20 ms/flight latency budget, 50 MB model size, feature store vs on-the-fly aggregation, drift detection, and periodic retraining cadence. g) Deliverables: reproducible notebook, clean data pipeline, model cards with fairness slices across carriers/airports, and an exec summary with recommended operational policy and estimated ROI.
Overview: This question assesses a data scientist's ML skills: target definition, leakage-aware feature engineering, temporal splitting and backtesting, model comparison and hyperparameter tuning, cost-sensitive evaluation, and production constraints like latency, model size, monitoring, and retraining.
Read the full Capital One Data Scientist interview experience that includes this question.
Solution
Framing the problem
The business objective governs everything: predict, before departure, whether a flight will arrive meaningfully late, so operations can take action (rebook crews, pre-position aircraft, warn passengers, add buffer). That framing imposes two strict rules: the model may only use information known at prediction time (a fixed horizon before scheduled departure), and evaluation must simulate how the model would have been used historically. These are exactly where most airline-delay solutions quietly fail due to leakage, so I treat leakage as the primary risk, not an afterthought. I will assume predictions are made at a fixed cutoff—say, 2 hours before scheduled departure—and freeze what is knowable to that moment.
a) Target definition and justification
Primary target: late_arrival = (arr_delay_min > 15), a binary label.
Justification:
- Operationally meaningful, not arbitrary. The 15-minute cutoff is the long-standing industry standard for an on-time arrival (the DOT/BTS on-time definition), so the label corresponds to a metric the business already tracks and is accountable for. Predicting it yields something stakeholders can act on and compare against.
- Binary classification beats raw regression for this use case. We could regress
arr_delay_mindirectly, but the decision (intervene or not) is inherently a threshold decision, and the delay distribution is heavy-tailed and zero-inflated (most flights on time, with a long right tail of severe delays). A calibrated classifier for P(late) is easier to threshold against a cost policy than a noisy minute-level regression. - Watch the label's edge cases.
arr_delay_minmust be the actual realized arrival delay used only for labeling, never as a feature. Cancelled or diverted flights lackarr_delay_min; decide explicitly—typically treat a cancellation as a positive (did not arrive on time) if the downstream cost is similar, or model cancellation separately. Document the choice; silently dropping cancellations biases the label toward optimism. Secondary targets worth defining for richer policy work (optional, mention but do not over-build): - A multi-class or ordinal version: on-time / minor (15–60 min) / major (>60 min), because the cost of a delay is non-linear.
- A regression head on
arr_delay_min(e.g., quantile regression at the 0.5/0.9 quantiles) if operations wants an expected-buffer estimate, not just a flag. I would lead with the binary classifier and keep the ordinal/quantile variant as a stretch goal.
b) Leakage-aware feature set
The governing rule: every feature must be reconstructable from data available at the prediction cutoff (2 hours before departure). I group features by source and explicitly state the lag. Schedule / static (known at booking time — safe):
distance,aircraft_type,carrier,origin,dest, scheduled block time (sched_arrminussched_dep).- Time encodings: hour-of-day of
sched_dep(cyclical sin/cos),day_of_week,month,holiday_flag, and a flag for the day before/after a holiday. Cyclical encoding avoids the artificial discontinuity from 23 to 0. - Route (
origin,dest) and a directional flag. Weather — forecasts only, never actuals. Use the forecast for the departure/arrival window issued before the cutoff (this is whatweather_features_*should represent at serving time). Using realized weather at the actual arrival time is leakage. Specifically: forecasted precipitation, wind, ceiling/visibility, convective probability at origin and dest for the scheduled departure and arrival hours. Historical aggregates — strictly lagged to t−7 days or earlier: - Route-level: mean/median
arr_delay_min, P(late), variance over the route's flights in a trailing window (e.g., last 7 or 28 days), as of t−7. - Carrier-level and airport-level rolling delay rates (origin departure-delay rate, destination arrival-delay rate) over trailing windows.
- Carrier×airport and aircraft-type rolling statistics for fleet/station effects.
- Critical leakage trap: these aggregates must be computed using an expanding or rolling window that ends strictly before the row's own date, not over the entire training set. Computing a route's mean delay over all dates (including future) and joining it back is the most common leakage bug here and greatly inflates offline metrics. Same-day upstream propagation (the highest-signal feature, but the trickiest):
- The single biggest driver of arrival delay is whether the inbound aircraft and crew are already late. Schema caveat: the listed columns have no tail number or registration, and
aircraft_typeis an aircraft class (e.g., B738), not an aircraft identifier—so you cannot link the exact physical inbound airframe from these columns alone. The cleanest proxy the schema does support is the prior leg flown under the samecarrier+flight_numearlier the sameflight_date, or a same-day leg arriving into this flight'soriginchained by (carrier,dest=origin,sched_arrapproximately this flight'ssched_depminus turnaround time). Use that proxy, and note that a true tail-rotation linkage would require adding a registration column. - Given a linkable prior leg, its current departure delay—observed at the cutoff—is enormously predictive and legitimate, because at 2 hours before departure we genuinely know that earlier leg's status. Include
inbound_dep_delay_so_far,inbound_in_air_flag, andturnaround_buffer=sched_dep−inbound_sched_arr(all derived from the proxy linkage above). - If the cutoff is before the inbound leg has departed, then you only have its forecast, not its realized delay—encode that honestly (use the inbound's own predicted P(late), or mark as unknown). Explicitly excluded (future-encoding) features:
arr_delay_min(the label), actual arrival time, actual taxi/airborne times, realized en-route weather, anddep_delay_minof this same flight if the cutoff precedes departure (we do not yet know it). If the cutoff is after pushback, you could use realizeddep_delay_min, but then state that explicitly—it changes the product. Encoding plan: high-cardinality categoricals (origin,dest,carrier, route,aircraft_type) via target/mean encoding computed within the CV fold with smoothing and out-of-fold predictions to avoid target leakage; or native categorical handling for the GBM (LightGBM/XGBoost). Never fit the target encoder on the full data before splitting.
c) Time-based split, imbalance handling, metrics
Split — strictly temporal, no shuffling. Random k-fold is invalid here: it leaks future into the past and allows near-duplicate same-day flights to straddle folds. Use the proposed scheme:
- Train: through 2024-06
- Validation (model selection, early stopping, calibration, threshold): 2024-07 to 2024-09
- Test (reported once, untouched): 2024-10 to 2025-03 Add a small embargo or gap (e.g., drop the few days around each boundary) so trailing-window features computed near the boundary do not peek across it. Recompute all rolling aggregates within each split's own causal window. Class imbalance. "Late" is the minority class (roughly a quarter or less of flights, depending on threshold and season—state qualitatively, do not quote a fixed number). Approach in priority order:
- Do nothing to the data first—GBMs handle moderate imbalance fine. Set
scale_pos_weightapproximately equal to (#neg / #pos) (XGBoost) oris_unbalance/class_weight(LightGBM), or class weights in logistic regression. This adjusts the loss, not the label prior. - Avoid naive oversampling or SMOTE for this problem: it distorts calibration (which we care about—see Brier), and SMOTE-ing time-series rows breaks temporal structure. If used, only on the training fold, and recalibrate afterward.
- Optimize a threshold on the cost curve, not 0.5 (part e), and calibrate probabilities (isotonic or Platt on the validation set) so the cost-sensitive threshold is meaningful. Primary metrics:
- PR-AUC (average precision)—the right ranking metric under imbalance; ROC-AUC is overly flattering when negatives dominate. PR-AUC focuses on the positive (late) class we care about.
- Calibrated Brier score plus a reliability diagram—because the downstream policy uses probabilities, not just rankings; a well-ranked but mis-calibrated model makes the cost-optimal threshold wrong.
- Secondary: recall at a fixed operational precision (e.g., recall at precision 0.5), ROC-AUC for continuity with prior reporting, and the expected cost per the 5:1 cost ratio (part e) as the ultimate business metric.
d) Baseline (regularized logistic regression) vs gradient boosting
Baseline — regularized logistic regression with target encoding.
- Pipeline: out-of-fold smoothed target encoding for high-cardinality categoricals, one-hot for low-cardinality, standardize numeric features,
LogisticRegressionwith L2 (or elastic-net viasaga). - Hyperparameters to search: inverse-reg strength
C(log-grid, e.g., to ), penalty (L2 vs elastic-netl1_ratio), class weight. Small search; LR is cheap. - Value: fast, fully interpretable coefficients, a calibration-friendly probabilistic baseline, and a sanity floor. If the GBM cannot clearly beat a well-tuned LR, something is wrong (often leakage in the GBM, or no real signal). Gradient boosting — XGBoost / LightGBM.
- Why it should win: non-linear interactions (route × weather × time-of-day × inbound delay) are exactly what trees capture, and native categorical/missing handling suits this messy data.
- Hyperparameters to search (Bayesian/Optuna over a temporal validation split, not random CV):
num_leaves/max_depth(capacity),learning_rate(small, e.g., 0.03–0.1) paired with early stopping on validation PR-AUC/logloss to choosen_estimators—never fix the tree count by hand.min_child_weight/min_data_in_leaf(regularize against tiny noisy leaves),subsample(bagging_fraction),colsample_bytree(feature_fraction),reg_alpha/reg_lambda.scale_pos_weightfor imbalance.
- Monotonic constraints—use them deliberately. Domain priors that hold monotonically: higher forecasted precipitation/wind → higher P(late); larger inbound delay → higher P(late); shorter turnaround buffer → higher P(late). Encoding these as monotone constraints buys robustness, easier stakeholder trust, and guards against weird non-monotone fits in sparse regions. Do not constrain features where the relationship is genuinely non-monotone (e.g., hour-of-day).
- Calibration: GBMs are often mildly mis-calibrated; fit isotonic regression on the validation set after training, then report Brier on test. Decision rule: pick by validation expected cost (5:1) and Brier, with PR-AUC as a tiebreaker—favor the GBM only if its cost advantage survives calibration and holds on the untouched test window. Keep the LR as the interpretable fallback and as a monitored shadow.
e) Rolling-origin CV, backtesting threshold policies, cost-sensitive eval
Rolling-origin (walk-forward) cross-validation. Instead of a single train/val cut, slide the origin forward to test stability across regimes (summer thunderstorms vs winter ops). Concretely, several expanding-window folds:
| Fold | Train through | Validate |
|---|---|---|
| 1 | 2024-03 | 2024-04 |
| 2 | 2024-04 | 2024-05 |
| 3 | 2024-05 | 2024-06 |
| … | … | … |
| Always train before validate, recompute features causally per fold, and report mean ± variance of PR-AUC / cost across folds. High variance flags a model fragile to seasonality—important to surface before it hits production. | ||
| Cost-sensitive evaluation. Define the confusion costs from the brief: a false negative (we said on-time, flight was late → no proactive action, expensive downstream recovery) costs 5× a false positive (we said late, it wasn't → wasted buffer/swap) | ||
| […]truncated[…] |