Meta · Statistics & Data Analysis
Design and analyze a group-calls experiment
TrueInterview
October 7, 2026 · 2 min read
You are evaluating whether to launch Group Video Calls. Answer each part precisely; justify your choices with pros and cons, and use formulas where relevant.
-
Clarify the CTR comparison: For a new entry-point button, should the primary CTR be (a) Treatment vs Control measured at the same time (between-subjects), or (b) Treatment users' Month 1 compared with their own Month 6 baseline (within-subjects)? State when each approach is valid, the bias risks (seasonality, maturation), and recommend a preferred estimator. Give formulas for: (i) the simple difference versus control, (ii) the within-user pre/post difference, and (iii) the difference-in-differences estimator that combines both. Specify the exact windows you would use (e.g., M1 = 2025-09-01 to 2025-09-30, M6 = 2025-03-01 to 2025-03-31) and how you would treat users who lack a full baseline.
-
Experiment design: Propose a complete A/B test (or cluster test) for Group Calls that accounts for interference/network effects (users call each other). Specify the randomization unit (user, household, call-graph clusters, geography), allocation, stratification (e.g., country, device), and ramp plan. Define the primary success metric(s) and justify them (e.g., incremental video-call minutes per DAU, completed-call rate), and list guardrails (e.g., app crashes, latency, churn). Provide decision thresholds and a stopping plan (power, MDE, alpha, sequential corrections).
-
Interference handling: Calls involve multiple users who may be assigned to different variants. Describe a design that prevents contamination (e.g., cluster-by-ego-network, geo switchback). Explain how you will attribute a group call to a variant and how you will analyze spillovers; include at least one robustness check (e.g., exposure reweighting, CUPED, cluster-robust SEs).
-
Launch decision resources: In addition to the experiment, list external/internal resources you would consult before approving group calls (market research, competitor benchmarks, support tickets, qualitative UX studies, capacity/SRE constraints, cost models) and explain what each would change in your launch threshold.
-
Set participant limit: Propose a data-driven method to choose a maximum participants-per-call limit (e.g., 4/8/16) under infrastructure constraints. Outline how you would simulate expected concurrency using historical call distributions, model QoS (latency, failure rate) as a function of , and run a multi-arm experiment to select the limit. Include guardrails and rollback criteria.
-
Edge cases: Explicitly address novelty effects, day-of-week effects, France-only versus global rollouts, and how you would ensure the analysis uses correct absolute dates (e.g., yesterday = 2025-08-31) and local time handling.
Overview: This question tests competency in experimental design, causal inference, metric specification, interference handling, and infrastructure-aware analytics within the Analytics & Experimentation domain for a Data Scientist role.