Meta · ML & AI Fundamentals
Evaluate fraud classifier with cost-sensitive metrics
TrueInterview
October 7, 2026 · 1 min read
Imagine you take over an existing binary classifier used for fraud detection. a) At one chosen threshold you get a holdout confusion matrix of TP=200, FP=800, FN=100, TN=99,900 — compute precision, recall, F1, and the false-positive rate. b) Suppose a missed fraud case costs $20 and a wrongly flagged case costs $0.20, with prevalence at 1% — explain how you would use predicted probabilities to select a threshold that maximizes expected utility, and name which metric (PR-AUC versus ROC-AUC, say) better captures improvements at low prevalence and why. c) Sketch a calibration check together with a fix (a reliability plot, isotonic or Platt scaling, for example), and describe how calibration and threshold selection interact. d) Propose an online evaluation plan with guardrails that stop legitimate users from being over-blocked while the catch rate improves (shadow evaluation run alongside, or a second review stage for low-confidence positives, for instance), and define success criteria and rollback triggers.
Overview: This question assesses competence in binary classifier evaluation, cost-sensitive decision-making, probability calibration, and safe online deployment in the Machine Learning domain for a Data Scientist role, requiring hands-on application grounded in conceptual understanding.