Meta · Statistics & Data Analysis
How to measure harmful-content severity and run experiments
TrueInterview
October 7, 2026 · 2 min read
Question
You work as a Data Scientist on content integrity / harmful content at a large social media platform, covering areas such as hate and harassment, self-harm, graphic violence, sexual exploitation, spam, and misinformation. Harmful content varies widely in severity. The team aims to reduce the harm users face and is considering an intervention, such as a ranking demotion, a revised removal or enforcement policy, or a new ML classifier paired with an enforcement workflow. Design a measurement and experimentation framework for this problem. Cover the following points:
- Define "severity" of harmful content.
- Suggest a severity framework that can support both measurement and decision-making.
- Which signals would you rely on (policy labels, human review, user reports, downstream user harm, virality, repeat exposure, content type, viewer vulnerability)?
- Should severity be represented as a binary label, ordinal levels, or a continuous score — and why? Compare the pros and cons of each.
- Discuss the tradeoffs (interpretability vs. sensitivity, policy alignment, subjectivity, multilingual / cross-region considerations).
- Design metrics (primary / diagnostic / guardrails).
- Provide one clearly defined primary metric for harmful-content impact, several diagnostic metrics that explain why it moves, and several guardrail metrics that flag unintended harm.
- Separate prevalence (creation-side), exposure (distribution-side), severity-weighted exposure, enforcement accuracy, and user-experience side effects. State the pros and cons of each.
- Give exact definitions (numerators/denominators) and any weighting (for example, by exposure or severity).
- Which denominator is appropriate — content created, content viewed/impressions, active users, or sessions — and how does that choice depend on the intervention?
- Design an experiment to evaluate the intervention.
- Pick an appropriate randomization unit (viewer/user, viewer-session, content item, author/creator, community/network cluster, or geo) and justify it. Discuss the tradeoffs of each option.
- State the primary success metric, guardrail metrics, and long-term metrics.
- Discuss the pitfalls: interference / spillover (content spreads across users and social graphs), network effects, contamination, novelty effects, delayed outcomes, and measurement error. How should you address interference?
- Describe the analysis plan (intent-to-treat vs. per-protocol, variance reduction, segmentation, multiple testing).
- Biases, pitfalls, and edge cases.
- List sources of selection bias (reporting bias / brigading), labeling bias and reviewer drift, delayed feedback, and Simpson's paradox / subgroup regressions.
- How should you handle rare-but-severe harms versus common low-severity harms (the base-rate problem)?
- How would you stop the team from "improving" the chosen metric while making the platform worse overall (metric gaming)?
- How would you arrive at the final launch recommendation?
Assumptions
- You have access to logs for impressions/views, engagement, reports, enforcement actions, and model outputs.
- "Harmful content" is identified through a mix of policy rules, human review, and ML signals, all of which are imperfect. Overview: A Meta Data Scientist analytics and experimentation question: design a measurement and experimentation framework for harmful content where severity varies. It covers defining severity (binary vs. ordinal vs. continuous), a metric stack (prevalence, exposure, severity-weighted exposure, enforcement accuracy, guardrails) with the right denominator, A/B test design and randomization-unit choice, interference/spillover, and biases including reporting bias, Simpson's paradox, the rare-severe base-rate problem, and metric gaming.
Loading comments…