Thumbtack · Statistics & Data Analysis
Explain power drivers and resolve unexpected A/B results
TrueInterview
October 7, 2026 · 1 min read
Give concise answers to every part, including calculations where they are asked for. (a) Define statistical power in a two-proportion A/B test, then list the main levers that raise power, ordered by typical real-world impact from greatest to least, with a short explanation of each trade-off. Cover: effect size (MDE), variance or metric volatility, sample size, allocation ratio, alpha, variance reduction (such as CUPED), bucketing/stratification, and test duration/seasonality. (b) Scenario: baseline conversion . Target relative lift is +7%, so . Use two-sided , desired power 0.80, equal allocation, independent users, and no clustering. Compute the required sample size per variant and the minimum test duration in days when 80,000 eligible users arrive per day and 10% post-randomization attrition is expected. Show formulas and numeric results. (c) Recompute part (b) under CUPED with , meaning a 30% relative variance reduction. What are the new per-variant sample size and duration? (d) How does a 90/10 allocation (90% control, 10% treatment) change power when total traffic is held constant? Give the intuition and, where possible, a quantitative comparison with equal allocation. (e) Your test, run for the duration computed in (b), comes back with a statistically significant -2% lift (treatment worse), opposite your prior expectation of +7%. Outline a step-by-step diagnostic plan before reaching conclusions; include SRM checks and why they matter, audits of instrumentation and metric definitions, bot/geo/device imbalance, novelty or learning effects, outlier clipping, Simpson’s paradox through key segments, guardrail metrics, and peeking/stopping risk. (f) After diagnostics, propose an evidence-based decision tree for when to (i) ship, (ii) iterate with a follow-up test and specify one design change, or (iii) rerun, stating the precise condition that would justify a rerun. Overview: This question tests a data scientist's command of A/B testing fundamentals: statistical power and sample-size calculations, effect-size and variance considerations, allocation strategies, variance-reduction methods such as CUPED, and experiment diagnostics including SRM, instrumentation audits, imbalance, and segmentation checks.