Meta · Statistics & Data Analysis
Choose robust metrics for skewed comments
TrueInterview
October 7, 2026 · 2 min read
Daily comment counts per user on a website are heavily skewed and contain many zeros. A backend optimization is deployed with the expectation of boosting engagement.
(a) For data like this, explain when the mean, median, 10% trimmed mean, 95/5 winsorized mean, and geometric mean of are the better choices for central tendency. Discuss the bias–variance trade-offs when tails are heavy (for example, Pareto) and how interpretable each is for product decisions.
(b) Assume the control group’s daily per-user counts are [0,0,0,1,1,2,2,3,20,50] and the treatment group’s are [0,0,1,1,1,2,2,3,5,10]. Compute the mean, median, 10% trimmed mean, and winsorized mean for each group, then decide which estimator most reliably picks up a practically meaningful improvement in this case. Justify your answer rigorously.
(c) Explain how you would construct a 95% confidence interval for the estimator you selected using a nonparametric bootstrap stratified by user activity buckets. List the assumptions and describe how you would verify them.
(d) If you need to report an effect size that is robust yet comparable across experiments, suggest a transformation and effect metric—for example, a log1p-based percent change or a quantile treatment effect at —and justify your choice.
Overview: This question tests knowledge of robust estimation and inference for zero-inflated, heavy-tailed count data, covering choices of central tendency (mean, median, trimmed and winsorized means, geometric mean), nonparametric bootstrap confidence intervals, and robust effect-size transformations.
Community responses
Response by SS
(a) (a) Which estimator should you use?
Mean Use when: total impact matters (overall comment volume) Problem: a few highly active users can dominate it → It can be noisy under heavy tails
Median Use when: the typical user is what matters Problem: with many zeros, the median can remain 0 even when engagement improves → Too insensitive for this setting
10% Trimmed Mean Use when: you want to discard extreme low and high values Problem: dropping data can distort results, especially when zeros are common → A reasonable balance, but it sacrifices data
Winsorized Mean (95/5) Use when: you want robustness while retaining all observations How: cap extreme values rather than deleting them → Often the best practical option
Geometric Mean of Use when: you want to smoothly reduce the influence of large outliers Problem: harder to explain and not linked to total counts → Useful, but less intuitive
Response by SS
Mean Control: sum = → mean = Treatment: sum = → mean = → Control appears higher than treatment, driven by the 20 and 50.
Median Control: middle = Treatment: middle = → No difference.
10% Trimmed Mean
(remove 1 smallest and 1 largest)
Control:
Trim → [0, 0, 1, 1, 2, 2, 3, 20]
Sum = →
Treatment:
Trim → [0, 1, 1, 1, 2, 2, 3, 5]
Sum = →
→ Control is still higher, but the gap is smaller.
Winsorized Mean (95/5)
(replace min with 2nd smallest, max with 2nd largest)
Control:
Winsorized → [0, 0, 0, 1, 1, 2, 2, 3, 20, 20]
Sum = →
Treatment:
Winsorized → [0, 0, 1, 1, 1, 2, 2, 3, 5, 5]
Sum = →
→ Control is still higher, but the effect of extremes is reduced.