Meta · Statistics & Data Analysis
Define composite success for search and test it
TrueInterview
October 7, 2026 · 1 min read
Each query for a new search feature receives two binary labels: relevancy (1/0) and accuracy (1/0).
-
Suggest a composite success metric built from the two labels. Specify the exact scoring rule (for example AND, a weighted score, or lexicographic ordering), defend that choice under different error costs, and describe how you would calibrate the weights against business outcomes.
-
State the aggregation level (query, session, user, or day), and explain how to handle multiple queries per user, missing labels, and correlated outcomes.
-
Design an online experiment covering the randomization unit, primary and guardrail metrics, power and sample-size inputs, pre-registration of the analysis, and a plan for limiting p-hacking across many slices.
-
Propose an offline evaluation pipeline (label collection, inter-rater agreement, golden sets), and explain how you would monitor label drift and Simpson’s paradox when segmenting by intent or locale.
Overview: The question assesses ability to design composite success metrics, experiments, and evaluation pipelines for search features, covering metric formulation, calibration to business outcomes, aggregation choices, handling missing or correlated labels, statistical power and sample-size reasoning, and label collection and monitoring in the Analytics & Experimentation area for Data Scientist roles. It is often used to evaluate whether measurement aligns with business objectives and whether candidates can reason about trade-offs and biases; it tests both conceptual understanding of trade-offs and practical skill in running online experiments and building offline evaluation pipelines.