Stripe · Statistics & Data Analysis
Choose threshold under costs and uncertainty
TrueInterview
October 7, 2026 · 1 min read
Suppose a deployment sends an incentive to users predicted to be positive. The base rate is for purchases without an incentive. The benefit per true positive (incremental profit) is $50. The cost per false positive (incentive plus email) is $1.00. Three candidate operating points from validation are A: TPR = 0.70, FPR = 0.12; B: TPR = 0.55, FPR = 0.05; C: TPR = 0.80, FPR = 0.20. (1) For a cohort of 100,000 users, compute the expected incremental profit for A, B, and C using . Which threshold is best? (2) Give a 95% confidence interval for the chosen threshold’s profit using a delta method or a nonparametric bootstrap; specify which distributional components you resample and why. (3) The model outputs probabilities; describe and compute two calibration diagnostics you would include, such as Brier score and a reliability curve with ECE. (4) Outline a monthly population drift test to ensure the chosen threshold remains optimal as shifts.
Overview: This question assesses cost-sensitive threshold selection, uncertainty quantification for expected profit, calibration diagnostics, and population drift detection in a data science context.
This question is drawn from a data science interview.