Meta · Statistics & Data Analysis
Design metrics for violating content exposure
TrueInterview
October 7, 2026 · 1 min read
You are building metrics for a UGC product with a mix of automated and human moderation. Create a measurement framework for user exposure to policy-violating content.
-
Lay out exact formulas for at least three daily metrics and their 7-day rolling equivalents: , , and .
-
Give precise inclusion and exclusion rules for counting a view as violating under two timing schemes: ex-ante, where only violations known when the view happened should count, versus ex-post, where the final review decision applies; also cover late-arriving labels, appeals, deleted items, repeated views, and traffic from bots.
-
Should view_prevalence serve as the north-star metric? Contrast it with and ; discuss the relevant tradeoffs, such as detection lag, denominator gaming, changes in precision and recall, Simpson’s paradox across regions or surfaces, and Goodhart’s law.
-
Suggest a weekly alerting approach whose thresholds rely on uncertainty intervals, such as Wilson intervals or Bayesian beta-binomial intervals, and describe guardrails like false-positive exposure, creator churn, and review-queue SLA.
-
Outline an A/B test aimed at lowering view_prevalence: specify the primary metric, the main segments—country, surface, creator cohort—power assumptions, and how you would adjust for label latency and selection bias when violations are only found after a user has already been exposed.
Overview: The question tests metric-design and measurement skills for content safety, covering exposure metric definitions, inclusion and exclusion rules, label latency and appeals, uncertainty-aware alerting, and A/B test design in an Analytics & Experimentation Data Scientist context.