Meta · Statistics & Data Analysis
Design measurement to detect fake accounts
TrueInterview
October 7, 2026 · 2 min read
Context
You're on a social platform team. Friend requests—sending, receiving, accepting, declining—are the only product surface you can depend on. Start from the assumption that there is no anti-fake model, no rule set, and no metrics already in place.
Task
- Give an operational definition of a “fake account.”
- Which behaviors count: spam, scam, bot activity, account farming?
- How would you treat ambiguous or gray-area accounts?
- Design the data and instrumentation.
- Which events and fields would you record for friend requests and the actions that follow?
- Which joins or identifiers do you need to follow outcomes across time?
- Propose an initial detection approach that does not use a model.
- What heuristic signals or risk score would you begin with—rate limits, graph patterns, acceptance ratios, burstiness, messaging after acceptance if available, and so on?
- How would you set thresholds while avoiding harm to legitimate users?
- Measurement and evaluation plan.
- How would you get labels—manual review, user reports, enforcement actions—and handle labels that are delayed or biased?
- What are the primary, diagnostic, and guardrail metrics?
- Platform-level reporting.
- How would you estimate and report the platform’s fake-account problem over time—prevalence and incidence—when you can only see partial ground truth?
- What would you present to executives versus an operational team?
Overview: The question tests a data scientist’s ability in fraud detection measurement, event instrumentation, labeling strategy, metric design, and reporting under constrained product signals.
Community answers
Answer by SS
Q3: Detection Heuristics (No Model) Risk score is a weighted sum of: Volume score (weight 0.4): requests sent compared with the 7-day P95 Acceptance score (weight 0.4): acceptance rate compared with the 7-day P25 Account age score (weight 0.2): under 7 days counts as high risk Thresholds: 0.9 or above → auto-ban 0.8 to 0.9 → manual review 0.7 to 0.8 → rate limit below 0.7 → no action Protect real users: use a 7-day observation window before flagging.
Q4: Measurement and Labels Labels come from: User reports, such as “Report as spam” Manual review queue Post-enforcement actions, such as accounts banned later for violations Metrics: Primary: Precision: flagged accounts that are actually fake Recall: true fake accounts that were caught Diagnostic: Overall platform acceptance rate Time to detection, in days until flag Manual review queue size Guardrails: False positive rate below 1%—real users flagged DAU/WAU stable User retention unchanged
Q5: Platform Reporting For executives: “We detect X% of fake accounts daily, affecting Y million users” “Estimated true rate: X–Z% (confidence interval)” “Cost: $W in support plus churn” Show the trend: is it improving or worsening? For ops teams: Precision and recall by threshold Detection latency distribution False positive examples Manual review backlog Key point: you only observe detected fakes, not all fakes, so use sampling to estimate true prevalence.
Bonus: One More Data Source? Messaging volume—it correlates with friend request spam and strengthens detection.
Remember: Stay focused on the constraints—friend requests only Use standard metrics such as precision and recall