Apple · Statistics & Data Analysis
Evaluate a model and choose metrics
TrueInterview
October 7, 2026 · 1 min read
You are responsible for a fraud-screening model used on e-commerce orders.
Fraud occurs at a base rate of 0.7%.
Costs attached to each action: sending an order to manual review costs $3; a legitimate order that gets flagged incurs $1 in friction; letting a fraudulent order through costs $120; passing a legitimate order correctly costs $0.
Two candidate models are scored on a validation set of 100,000 orders containing 700 positives, both at threshold 0.5, producing: Model A: TP=490, FP=4,900, FN=210, TN=94,400. Model B: TP=560, FP=8,400, FN=140, TN=90,900.
Tasks: (a) Calculate precision, recall, F1, a ROC-AUC proxy taken from the TPR/FPR points, and expected cost per order for A and B at 0.5. Which model is preferable given the costs above? (b) Derive the cost-optimal threshold in general form, expressed through calibrated and the cost parameters; then apply it to this setting, assuming perfect calibration and the stated base rate. (c) Discuss PR-AUC versus ROC-AUC when imbalance is extreme, calibration diagnostics (Brier, ECE), and decision curve analysis / net benefit. (d) Suggest an offline evaluation plan that holds up under prevalence shift, plus a safe online A/B with guardrails (manual review SLAs, false accusation rate, a holdout for drift), and describe how you would watch for concept drift and fairness across user segments after launch.
Overview: This prompt assesses a data scientist on cost-sensitive model evaluation, extreme class imbalance, calibration and threshold derivation, experiment design, and post-launch monitoring and fairness, all within the Analytics & Experimentation domain.