Back to problems

Design long-tail search evaluation under label budget

System Design · Google · Hard

Suppose you are responsible for evaluating two ranking functions for a large web search engine that receives about 80 million queries per day. Query traffic is extremely skewed: it follows a Pareto distribution with shape parameter $$\alpha = 1.1$$; the highest-frequency 1% of queries account for roughly 70% of traffic, and the remaining 99% make up the long tail. Your metric is NDCG@10, and relevance labels are graded on a 0–3 scale. Each week, human raters can label at…

Checking your access…