Amazon · ML & AI Fundamentals
Diagnose and fix underperforming ML model
TrueInterview
October 7, 2026 · 1 min read
You've taken over a binary fraud model with severe class imbalance, where positives make up roughly 2% of cases. On a time-split validation set, current results are AUC=0.61 and precision at 0.90 recall is only 0.05. You have one day to make a meaningful recall improvement while keeping review capacity fixed. 1) Explain how you would quickly determine whether the model is underfitting or overfitting using learning curves, calibration plots, PR versus ROC trade-offs, and leakage checks. 2) Suggest three targeted changes that can be shipped within a day—for example, class-weighted loss, monotonic gradient boosting with categorical encoders, or threshold moving driven by cost-sensitive utility—and explain why each should improve performance. 3) Show how you would pick a decision threshold that maximizes expected utility when false positives cost $2, false negatives cost $50, and review capacity is 0.5% of traffic; provide the utility formula and describe the validation-time procedure. 4) List the minimal logging and monitoring you would add at deployment to catch drift and data quality problems within a week. Overview: This question tests a data scientist's ability to diagnose and fix underperforming binary classifiers under severe class imbalance, including validation diagnostics, calibration, threshold selection under operational review limits, cost-sensitive utility reasoning, and basic deployment monitoring for drift.