ByteDance · Statistics & Data Analysis
Compute cluster-aware significance and sequential corrections
TrueInterview
October 7, 2026 · 1 min read
Suppose we run a randomized experiment on the tipping UI with creators as the randomization units. Each arm contains 10,000 creators, and a creator averages viewer sessions during the analysis window. At the viewer level, the purchase rate is 5.00% for control and 5.20% for treatment. Purchases within the same creator have an intra-cluster correlation of . 1) Calculate the design effect and the effective viewer-level sample size for each arm; then obtain the z-statistic and two-sided p-value from the cluster-robust standard errors implied by that design effect. 2) Suppose there are 4 interim looks plus one final analysis; approximate an O’Brien–Fleming-style spend of overall by specifying a conservative per-look alpha, and compare that with a plain Bonferroni correction. Explain how these alternatives affect power and the required duration. 3) Given four guardrail metrics, outline a Holm–Bonferroni adjustment, and discuss the situations in which you would instead report Bayesian posterior intervals with a ROPE to assess practical significance.
Overview: The question tests skill in analyzing clustered randomized experiments: computing the design effect and effective sample size, performing cluster-robust inference for a difference in proportions, applying sequential alpha spending in the style of O’Brien–Fleming and comparing it with Bonferroni, using Holm–Bonferroni adjustments for several guardrail metrics, and interpreting Bayesian ROPE results. It belongs to the Statistics & Math domain and is often used to see whether candidates can handle intra-cluster correlation, keep Type I error controlled across interim looks and multiple metrics, and show both conceptual understanding and practical application of the trade-offs among power, duration, and multiplicity.