Meta · Statistics & Data Analysis
Model comment count distribution and validate assumptions
TrueInterview
October 7, 2026 · 1 min read
Suppose you are looking at daily comment counts per post for a large social app, and the distribution is heavily skewed with many zeros. a) Pick a suitable discrete model from Poisson, Negative Binomial, Poisson–lognormal, and discrete power-law (with xmin), and justify your choice through overdispersion, tail weight, and zero inflation. b) Describe a step-by-step model selection plan: examine the mean–variance relationship; run a dispersion test; fit Poisson and NB by maximum likelihood and compare them with a likelihood-ratio test; fit Poisson–lognormal and compare it with NB using AIC/BIC and Vuong's test; for the upper tail, estimate xmin by minimizing the Kolmogorov–Smirnov statistic and compare power-law versus lognormal tails; use QQ-plots and posterior predictive checks. c) Using your chosen model, estimate and the 95th percentile of comments per post with bootstrap 95% confidence intervals; explain how you would compute standard errors and guard against small-sample bias. d) Explain how left-truncation (for example, posts with 0 comments sometimes being missing), right-censoring (for example, late-arriving comments), and mixture segments (bots versus humans) would bias the estimates, and how to correct for them (for example, zero-inflated NB or a truncated likelihood). e) If threaded replies launch next month, predict qualitatively how the parameters would change (for example, higher variance and a heavier tail), and describe how you would revalidate model fit one week after launch.
Overview: The question assesses skill in statistical modeling of discrete count data—in particular overdispersion, zero inflation, heavy tails, model selection, and inference with uncertainty quantification—within the Statistics & Math domain, and it calls for both conceptual understanding and practical application.