Capital One · ML System Design
Evaluate and monitor a credit risk model
TrueInterview
October 7, 2026 · 1 min read
You are building a probability-of-default model (12-month horizon) for a consumer lender whose default rate is 1%.
Misclassification costs work out as follows: letting a defaulter through (a false negative) runs $1,200, while turning away a creditworthy applicant (a false positive) runs $60.
Supervisors care more about stability and interpretability than about raw accuracy.
Name the three evaluation priorities that matter most in this setting and explain why each one earns its place.
Then lay out a complete plan spanning: (1) offline evaluation (temporal cross-validation, how you handle class imbalance, and which metrics you select — AUC-PR, expected cost, calibration error, KS — with a reason for each); (2) choosing a threshold that minimizes expected cost while keeping the decline rate at or below 4%; (3) backtesting across out-of-time cohorts and stressed periods; (4) a live champion–challenger test with guardrails; (5) production monitoring (data and label drift, calibration tracking, PSI thresholds, segment-level stability, plus an alert and runbook).
Spell out the calculations, the acceptance criteria, and what you would do if calibration drifts while rank ordering stays intact.
Overview:
This item tests the ability to design, evaluate and monitor cost-sensitive consumer credit probability-of-default (PD) models, emphasizing calibration, rank ordering, threshold selection, backtesting, champion–challenger experimentation and production monitoring under regulatory expectations for stability and explainability.