Meta · Statistics & Data Analysis
Quantify launch decision with tests and guardrails
TrueInterview
October 7, 2026 · 2 min read
You are asked to formalize the statistical decision rules for the Instagram button experiment described above.
Assume a baseline exploration rate of per user-week, a desired absolute MDE of percentage points, a two-sided , and power of . Randomization happens at the cluster level, with average cluster size users and intra-cluster correlation .
(a) Compute the required number of users and clusters per arm from a proportions test adjusted for clustering, using the design effect . Show the formulas and the resulting final counts. (b) State the exact hypothesis test you would use for the primary metric, such as a cluster-robust z-test on cluster means or a user-level test with cluster-robust standard errors. Explain when a nonparametric alternative would be preferable. (c) Define the confidence interval you will report and how you will interpret it jointly with practical significance. (d) You will review the primary metric each week for four weeks. Choose and justify a sequential testing plan, such as O’Brien–Fleming alpha-spending, and give the adjusted per-look alphas. (e) You track three guardrails: p95 latency, crash rate, and add-to-cart rate. Describe a multiple-testing control that preserves power on the primary metric, such as hierarchical testing or Holm–Bonferroni, and write the decision logic that combines the primary metric with the guardrails. (f) Suppose contamination causes 10% of control users to see the button. State the direction of the resulting ITT bias and outline a correction, such as CACE with assignment as an instrument.
Overview: The question assesses a data scientist’s competence in experimental design, statistical inference, and causal-effect estimation for clustered A/B tests. It spans sample size calculation with a design effect, choice of cluster-robust hypothesis tests, sequential alpha-spending, multiple-testing control for guardrails, and quantifying contamination bias. It is commonly asked in the Statistics & Math domain to evaluate the ability to formalize decision rules that control type I and type II error rates and to interpret confidence intervals under clustering and interim looks, covering both conceptual understanding of statistical principles and practical experiment governance.
Read the full data scientist interview experience this question came from.
Community answers
Answer by SS
(a) Given: Baseline rate: . Absolute MDE: , so . Significance: two-sided, giving . Power: , giving . Cluster size: . ICC: .
Step 1: Sample size ignoring clustering Variance = . MDE = . Users per arm = . Users per arm without clustering: about .
Step 2: Adjust for clustering with the design effect
Step 3: Effective sample size with clustering Users with clustering = . Adjusted users per arm: about .
Step 4: Convert users to clusters .
Final answer Per arm: Users required: about . Clusters required: about . Total experiment: Users: about . Clusters: about .
Key insight The roughly design effect sharply increases the required sample size because of clustering. Even a modest ICC of becomes expensive when the cluster size is as large as .