Meta · ML System Design
Choose metrics for fake-user classifier
TrueInterview
October 7, 2026 · 2 min read
You believe a large number of fake accounts are artificially inflating comment counts. Your job is to build a classifier that flags fake accounts for review. Propose and justify evaluation metrics and threshold choices under two operational constraints, then complete the required calculations:
Context: There are 10,000,000 daily active users; the true fake rate is about 1%; the review team can handle 50,000 accounts per day. Two candidate models give these validation metrics at their selected thresholds:
- Model A: precision of 0.60 and recall of 0.20 at threshold τA.
- Model B: precision of 0.20 and recall of 0.80 at threshold τB.
Tasks:
- Choose offline metrics: Explain when PR-AUC is preferable to ROC-AUC. Name the primary metrics, including precision@K, recall@K, PR-AUC, calibrated Brier score, and cost-weighted utility. Justify these choices given the severe class imbalance and limited review capacity.
- Capacity feasibility: For each model at its given threshold, calculate the expected true positives and false positives per day if applied to the full population. Indicate whether each model stays within the 50,000/day capacity and, if it does not, how you would set K or raise the threshold to satisfy capacity while maximizing expected true positives.
- Business trade-offs: With a false positive cost of $2 (review cost) and a false negative cost of $100 (missed abuse), choose an Fβ score with an appropriate β and justify it. Show the expected daily cost for Model A and Model B at their current thresholds.
- Thresholding and calibration: Describe how you would select τ using a precision-recall curve under the constraint precision at least 0.7 or false positives at most 20,000 per day; explain how you would apply probability calibration (Platt or isotonic) before thresholding.
- Validation protocol: Describe time-based cross-validation to prevent leakage, offline-to-online guardrails (such as CUPED or an AA test), and which online metrics you would monitor (reported precision among reviewed accounts, review throughput, and downstream abuse reduction).
Overview: This question tests a candidate's ability to choose and interpret evaluation metrics for severely imbalanced classification problems, carry out thresholding and probability calibration, and quantify capacity- and cost-constrained trade-offs between precision and recall.
Read the full interview experience this question came from.