Meta · Statistics & Data Analysis
Design analysis to test social vs game engagement
TrueInterview
October 7, 2026 · 9 min read
Question
Hypothesis: Users of Oculus (Meta Quest) who engage with social features are more consistently active than those who engage with game features. With activity data spanning roughly 8 weeks (for instance, 2025-07-01 through 2025-08-31), outline a rigorous analysis plan to test this claim. Address each of the following:
- Define the outcome. Specify exactly what "regularly engaged" means (for example, at least 3 active days per week across the 8-week window, or at least 10 active days within any 28-day period). Select one primary metric and 2–3 guardrail metrics. Explain the rationale for each and give a metric definition suitable for analysis that is resistant to outliers and seasonal effects.
- Ideal randomized experiment. If randomization is possible, describe the precise experiment: the randomization unit, the treatment (such as an onboarding prompt that steers first-week exposure toward social or game features), the primary metric, guardrails, stratification factors, and the steps to avoid contamination and novelty effects. Give the power or minimum detectable effect target at .
- Observational fallback (causal inference). When randomization cannot be done, suggest an observational design that contrasts social-only and game-only users. State the inclusion and exclusion rules (minimum account age, region, device), the null and alternative hypotheses, and a causal method (propensity score matching or weighting, inverse probability weighting, or exact matching on tenure buckets). Enumerate the covariates to adjust for (signup date or tenure, baseline activity, device, country, acquisition channel, content availability, weekday mix) and the checks to confirm common support and covariate balance.
- Power / sample size. Assume that among game-only users the baseline proportion who are "regularly engaged" is . You wish to detect an absolute increase of 3 percentage points (MDE = 0.03) with (two-sided) and power . Calculate the necessary sample size per arm for a two-proportion Z-test and write out all formulas and assumptions.
- Estimation and inference. Specify the main estimator and statistical test (for instance, a difference in proportions with cluster-robust standard errors when randomization is at the user level; a two-proportion z-test for the share of regular weeks; Welch's t-test or Mann–Whitney for mean weekly active days; or logistic regression with covariates and robust standard errors). Describe how you would manage multiple comparisons and interim analyses.
- Bias and robustness checks. Describe the bias and robustness procedures: pre-trend verification, difference-in-differences on users who change categories, placebo outcomes, sensitivity analysis for unmeasured confounding (such as Rosenbaum bounds), and multiplicity control for secondary metrics. Identify at least five specific validity threats (reverse causality — more engaged users choose social features; category misclassification; taxonomy drift; bots or multi-accounts; seasonality; geographic shocks) and explain how you would detect and address each.
- Decision-making and communication. Set a clear decision rule that requires statistical significance, a practical significance threshold, and intact guardrails (for example, and lift at least X% with a confidence interval excluding 0). Explain how you would present results, risks, and assumptions to product stakeholders, and what follow-up you would conduct if the effect differs across tenure cohorts.
Overview: This is a Meta (Oculus) data science onsite question about designing a rigorous study to test whether users of social features are more regularly engaged than users of game features. It spans metric definition, a randomized experiment, an observational causal-inference alternative (PSM/IPW), a two-proportion power calculation, estimation and inference, robustness and bias checks, and a decision rule.
Solution
1) Defining "regularly engaged" and the metrics Oculus / Quest is a VR hardware platform, so "engagement" refers to repeated headset sessions rather than web visits. Set the outcome as a clear binary defined over a fixed time window so that it can drive a two-proportion test:
- Primary outcome:
regularly_engaged= the user has at least 3 active days per week in at least 6 of the 8 weeks (or equivalently at least 10 active days in any 28-day sub-window). A binary, windowed definition is resistant to isolated heavy-use days and is the simplest to power. - Guardrails: (a) median session length or total minutes (detects a scenario where "social" boosts day counts through very short sessions), (b) 4-week retention or churn, (c) crash or comfort-related opt-outs (a VR-specific health and safety guardrail).
- Robustness: winsorize continuous metrics (such as minutes) at the 99th percentile; define active days using the user's local time zone; compare matching calendar weeks to cancel out weekday or holiday seasonality; require a minimum account age so the onboarding spike of brand-new users does not dominate.
2) Ideal randomized experiment
- Unit of randomization: the user (account). Randomizing at the user level prevents spillover between users and aligns with how the exposure is actually delivered.
- Treatment: an onboarding or home-screen prompt that pushes first-week exposure toward social features (treatment) versus toward games (control) — this changes the category of early exposure, which is the lever we can actually manipulate. (You cannot randomize a user's pre-existing preference, so you randomize the nudge, not the trait.)
- Primary metric:
regularly_engagedmeasured during weeks 2–8 (after exposure), so that the nudged sessions themselves are not mechanically included. - Stratification: device generation (Quest 2 vs 3 vs Pro), country or region, and baseline activity tier — stratified randomization improves precision and ensures balance on the strongest predictors.
- Guardrails: session length, retention, and comfort opt-outs, as described above.
- Contamination / novelty: randomize at the user level and hold the assignment fixed for the entire window to avoid cross-arm leakage; set aside a hold-out and examine the time series so a temporary novelty bump is not confused with a lasting lift; pre-register the analysis to prevent peeking.
3) Observational fallback — causal inference When randomization is not possible, compare social-only and game-only users while explicitly controlling for confounders.
- Inclusion/exclusion: require a minimum account age (for example, headset at least 30 days old before the window so onboarding is excluded), include only supported regions or locales, exclude flagged bots and shared family accounts, and restrict to active devices.
- Hypotheses: ; (one-sided if the direction is specified).
- Causal method: estimate a propensity score for "chooses social" using pre-window covariates — signup date or tenure, baseline activity in a pre-period, device generation, country, acquisition channel, content library size or supply, weekday mix — then apply propensity score matching or inverse probability weighting, optionally with exact matching on tenure buckets.
- Diagnostics: assess common support or overlap (trim propensity regions with no overlap), confirm covariate balance after weighting or matching (standardized mean differences below 0.1), and report the effective sample size after weighting.
4) Power / sample size (two-proportion Z-test) For a two-sided two-proportion test with equal group sizes, , , , and power :
where , , and .
- , so ; multiplied by 1.95996 gives about 1.33444.
- ; ; their sum is 0.4631, so ; multiplied by 0.84162 gives about 0.57273.
- Numerator .
- Denominator .
- , i.e. about 4,050 per arm (about 8,100 total). Assumptions: independent users (one observation per user), equal allocation, a fixed effect size, and no interim peeking; if you randomize by user but analyze user-weeks, inflate the sample size by a design effect for clustering. A simpler pooled approximation, , yields about 4,015 and is acceptable as a sanity check.
5) Estimation and inference
- Primary estimator: the difference in the
regularly_engagedproportions; test it with a two-proportion z-test (or logistic regression on the treatment indicator). With user-level randomization and one row per user, standard errors are sufficient; if you analyze repeated user-weeks, use cluster-robust standard errors by user. - Continuous secondary metrics (mean weekly active days, minutes): use Welch's t-test (unequal variances) or Mann–Whitney if the distribution is heavily skewed.
- Covariate adjustment (for observational designs or to improve precision): logistic regression with the propensity covariates and robust standard errors, or inverse probability weighted estimation.
- Multiple comparisons: control the family-wise error rate or false discovery rate (Bonferroni or Benjamini–Hochberg) across guardrails and secondary metrics.
- Interim looks: use alpha-spending (O'Brien–Fleming or Pocock) or sequential testing; do not peek with a fixed .
6) Bias and robustness
- Pre-trend checks: verify that social and game cohorts had parallel engagement before the window (a pre-treatment divergence invalidates the causal interpretation).
- Difference-in-differences on switchers: for users who change categories, compare before and after, differencing out fixed user characteristics.
- Placebo outcomes: test an outcome that should not be affected (such as account-settings visits); a "significant" placebo effect indicates residual confounding.
- Sensitivity analysis: Rosenbaum bounds — how large an unobserved confounder would need to be to overturn the result.
- Threats to validity (at least 5): (i) reverse causality — already-engaged users self-select into social features, not the other way around; address this with the randomized design or pre-period matching. (ii) category misclassification or taxonomy drift — audit the social/game labels and freeze the taxonomy for the window. (iii) bots, multi-accounts, or family-shared headsets — filter using device fingerprints and activity heuristics. (iv) seasonality and content shocks (such as a hit game launch) — align calendar weeks, add time fixed effects, and inspect week-by-week trends. (v) geographic or supply shocks — stratify by region and check robustness after excluding affected markets. (vi) novelty effects — measure the durable post-exposure window, not the nudge week itself.
7) Decision-making and communication
- Decision rule: proceed or conclude only if the primary effect is both statistically significant (, confidence interval excluding 0) and practically significant (lift at least a pre-specified threshold, e.g., +3 percentage points), and no guardrail metric regresses beyond its bound.
- Communication: start with the estimated lift and its confidence interval in plain language, state the design (randomized versus observational) and its key assumptions, and be explicit about residual confounding risk if the design is observational.
- Heterogeneity follow-up: if the effect differs by tenure cohort (for example, strong for new users, flat for veterans), report the interaction, avoid over-generalizing the average, and propose a targeted follow-up experiment on the responsive segment.
Explanation Rubric: a strong answer approaches this as a causal question, not a correlational one. It (1) fixes "regularly engaged" as a windowed binary that can be powered; (2) proposes randomizing the exposure nudge (you cannot randomize a pre-existing preference) with user-level assignment and stratification; (3) falls back to PSM/IPW with named confounders and overlap and balance diagnostics; (4) correctly applies the two-proportion sample-size formula (about 4,000–4,050 per arm here); (5) selects tests that match the data type and controls multiplicity and interim looks; (6) defends against the dominant threat — reverse causality or self-selection — plus misclassification, bots, and seasonality, with placebo and sensitivity checks; and (7) sets a decision rule requiring both statistical and practical significance with intact guardrails. Red flags: claiming causation from a raw social-versus-game comparison, ignoring self-selection, or omitting the power calculation.
Community answers Answer by fm Am I really expected to know all of that???