Google · Statistics & Data Analysis
Estimate population singletons from a 10% log
TrueInterview
October 7, 2026 · 1 min read
Each row of a day's search log corresponds to a single query string. From those rows you take a 10% simple random sample without replacement. Call a query a "unique query" (or singleton) when it occurs exactly once across the entire day's log.
a) Why is it biased to estimate the singleton count by tallying singletons in the 10% sample and scaling that tally by 10? State which direction the bias goes and explain the reasoning behind it.
b) Build a superior estimator from a frequency-of-frequencies model: connect the sampled counts to the population counts via binomial thinning, then put forward a Poisson or negative-binomial mixture, or a Good–Turing / Chao-style estimator, for .
c) Sketch how you would obtain standard errors (delta method, bootstrap) and how you would detect model misspecification when query frequencies are heavy-tailed.
d) Lay out a simulation plan for comparing estimators over traffic distributions that resemble real data.
Overview: The question tests command of statistical estimation and sampling theory — spotting bias introduced by subsampling, frequency-of-frequencies modeling for rare events, quantifying uncertainty when counts are heavy-tailed, and designing simulation studies.