Meta · Statistics & Data Analysis
Analyze DAU comments distribution and resampling
TrueInterview
October 7, 2026 · 3 min read
Consider the metric comments_per_DAU (the number of comments a daily active user makes in a day).
a) Shape: Describe and justify the distribution you would expect for comments_per_DAU across users on a given day (e.g., zero-inflation, skew/heavy tail). Is the variable discrete or continuous? What parametric families are reasonable to consider (e.g., Poisson vs Negative Binomial), and why might Poisson be inadequate?
b) Bootstrapping: You repeatedly resample users with replacement from that day’s user list and compute the sample mean, doing this 100,000 times. Describe the bootstrap distribution’s shape and center. Under what conditions will it be approximately normal, and when might it remain skewed? What is the relationship between its standard deviation and the population variance ?
c) Scaling n: If you increase from 10,000 to 20,000, how (quantitatively) does the width of the bootstrap distribution of the mean change? State the factor and the intuition.
d) Summary stats: For this metric, compare mean, median, mode, and p95. Which is most stable, which is most decision-relevant, and why might the mode be 0? How do you interpret and compute p95 for a discrete count variable (e.g., tie handling, integer vs real thresholds)?
e) Data types and aggregation: The per-user value is an integer, but the mean across users is a real number. Explain pitfalls from storing as integer vs float at different aggregation levels (e.g., truncation, rounding bias, overflow) and how you’d ensure numeric stability when computing large-day aggregates.
f) Estimation: Suppose the per-user variance is overdispersed (). Write the approximate standard error of the sample mean and discuss when you’d prefer robust estimators (trimmed mean, Winsorization) or variance reduction techniques (CUPED with a prior-day covariate).
Overview: This question evaluates a candidate's competency in statistical modeling of count data, resampling and bootstrap inference, summary-statistic interpretation, and numeric aggregation/stability considerations within the Statistics & Math domain for a data scientist role.
See the full Meta Data Scientist interview experience from which this question came.
Community answers
Answer by SS
(a) Shape of comments_per_DAU
Type of variable A discrete integer count variable (0, 1, 2, 3, … comments)
Expected distribution shape Strong right skew Most users make: 0 comments 1 comment Few power users make: 10, 50, 100+ comments Thus the distribution has a long right tail.
Zero inflation A large fraction of users may have exactly 0 comments, so we often observe a spike at zero followed by rapid decay and a long tail.
Good candidate distributions
Poisson (baseline)
Its appeal: it models count data and is simple. Its limitation: it assumes , but real data often has heavy tails, overdispersion, and many zeros.
Negative Binomial (better fit)
Why better: it permits and accommodates heavy-tailed users.
Zero-inflated models (best realistic) A mixture of two components: “inactive users” (always 0) “active users” (Poisson or negative binomial)
(b) Bootstrap distribution
You draw 10,000 users, compute the mean, and repeat 100,000 times.
Shape of bootstrap distribution
Center:
So it is centered at the true sample mean.
Shape: depends on the original data.
Case 1: well-behaved distribution large n mild skew In this case, the bootstrap mean is approximately normal.
Case 2: heavy skew or zero inflation long-tail users rare viral commenters In this case, the bootstrap mean: still remains centered correctly, but may be slightly skewed, heavy-tailed, and not perfectly normal.
When does the central limit theorem apply?
The bootstrap mean is approximately normal if: n is large (10,000 helps a lot), finite variance, and no extreme dominance.