Airbnb · Statistics & Data Analysis
Test conversion difference and adjust for clustering
TrueInterview
October 7, 2026 · 2 min read
For the aggregated results over the 7-day window 2025-08-26 through 2025-09-01, evaluate statistical significance and power for conversion uplift while accounting for day-level clustering: Given totals: Control (C): visits , bookings ; Treatment (T): visits , bookings .
- Point estimates: calculate , , the absolute lift (, in percentage points), and the relative lift.
- Significance: run a two-sided test for the difference in proportions using the unpooled standard error. Report the statistic, the -value, and a 95% confidence interval for . Mention any continuity correction you use.
- Clustering: adjust for day-level clustering with and 7 days per variant. Use the design effect , where . Recompute effective sample sizes as and provide an adjusted -value and confidence interval. Explain the assumptions and limitations of this correction.
- Power and sample size: What total visits per variant are needed to detect a 0.30 percentage-point absolute lift from a 3.00% baseline at 80% power and using an unpooled -test? Show the formula and final per variant. Then recompute with the design effect from to give a clustered per variant and the implied experiment duration if each variant receives 2,000,000 visits/day.
- Robustness: briefly describe how you would check day-to-day heterogeneity (e.g., -test or interaction with weekday) and how that influences the decision to launch. Overview: The question measures competence in statistical inference for A/B testing, including estimating and comparing conversion proportions, running two-sided hypothesis tests, correcting for day-level clustering via ICC and design effects, and computing power and sample size; it sits in the Statistics & Math domain for a Data Scientist role and mixes conceptual knowledge with practical use. It is frequently used to evaluate a candidate's ability to interpret conversion uplift under realistic experimental constraints, account for intra-cluster correlation when estimating effective sample sizes and uncertainty, and reason through experiment duration and robustness checks.
Loading comments…