Waymo · Statistics & Data Analysis
Compare two rare-event detection models statistically
TrueInterview
October 7, 2026 · 1 min read
You are evaluating two models, Model A and Model B, for a rare-event detection task—fraud, abuse, or medical adverse events, for example. Positive cases are extremely scarce.
You have only limited evaluation results for each model—perhaps a few aggregate counts such as TP, FP, FN, and TN, or precision and recall at a chosen threshold—and you may assume these are sufficient to reconstruct confusion-matrix counts, but the number of positive cases is small.
Questions
- Which metrics are best suited to comparing models when positives are rare, and why are accuracy or ROC-AUC by themselves insufficient?
- How would you compare Model A and Model B while accounting for statistical uncertainty?
- State the relevant distributions or assumptions, such as the binomial, and describe how you would calculate confidence intervals.
- How would you test whether one model is significantly better than the other?
- If you are working only with a small number of results, how would you make a decision responsibly? Address thresholds, calibration, and cost or alert-budget trade-offs.
Be explicit about your assumptions, such as paired versus unpaired evaluation and fixed threshold versus full curve.
Overview: The question probes a candidate’s ability to statistically evaluate models for rare-event detection: it covers choosing suitable metrics for imbalanced data, computing confidence intervals and running hypothesis tests with small samples, and weighing calibration, thresholding, paired versus unpaired comparisons, and cost or alert-budget trade-offs.
It appears often in machine-learning and data-science interviews and tests both conceptual grasp of statistical assumptions and the practical use of limited aggregate results to quantify uncertainty and guide decisions; it sits under machine learning and statistical inference, with emphasis on both conceptual and applied reasoning.