Pinterest · Statistics & Data Analysis
Design and assess a video-pin increase experiment
TrueInterview
October 7, 2026 · 9 min read
Question
Pinterest aims to grow the portion of video pins shown in the Home Feed — lifting video share from a baseline near 30% up to a 45% goal, i.e. about +10 to +15 percentage points — in order to raise engagement. Lay out a rigorous evaluation, then read the supplied results and give a recommendation.
A) Experiment design
- Unit of randomization & exposure control. Say whether you randomize at the user level or the session level, and defend the choice in light of network/content-supply interference and feed-ranking spillovers. Describe how you would limit each user's exposure to the new mix (for instance moving from a 30% baseline video share to a 45% target) and how the ramp would proceed.
- Metrics. Give one clear primary hypothesis and at least two alternatives (for example substitution effects, novelty effects, reliability cost). Specify a primary success metric plus guardrails, each with an exact formula (such as
saves_per_impression,clicks_per_impression,time_spent_per_user_day,complaint_rate = complaints / impressions,session_end_rate, creator churn/follows, crash rate, bandwidth cost per user). Explain the rationale for each and give the win/loss direction. - Power, duration & sequential analysis. Describe your power / MDE and duration assumptions (alpha, two-sided test, allocation, variance source), and how sequential looks or peeking will be handled (for example group sequential boundaries, or CUPED for variance reduction). Include an A/A check, a manipulation check (did video share actually move?), and a novelty/fatigue plan (minimum run and long-term holdout).
- Quasi-experiment. If an RCT cannot be run, propose a credible quasi-experiment (for example staggered-rollout difference-in-differences with user fixed effects plus inverse-propensity weighting, or synthetic control). List the identifying assumptions, diagnostics, and sensitivity checks.
B) Interpret the readout(s) — judge statistical significance, directionality, and practical significance; flag red flags (SRM, multiple testing, heterogeneous effects, cross-unit metrics). Then recommend ship, ramp/iterate, or stop, justifying the call with the metrics, multiple-testing/guardrail considerations, and possible mitigations (for example capping video share for sensitive cohorts, rank-quality filters), and state what further data or follow-up analysis you would want before a full rollout.
Readout 1 — 14 days, user-level randomization, robust SEs (N ≈ 2.0M users)
| metric | control | treatment | lift_% | p_value |
|---|---|---|---|---|
| CTR | 3.00% | 3.60% | +20.0 | 0.010 |
| avg_session_sec | 310 | 340 | +9.7 | 0.040 |
| 7d_retention | 28.0% | 27.0% | -3.6 | 0.070 |
| complaint_rate | 0.50% | 0.65% | +30.0 | 0.030 |
Readout 2 — 14 days (alternate run, with absolute lifts and CIs)
| metric | control | treatment | lift vs ctrl | p-value | 95% CI |
|---|---|---|---|---|---|
| CTR (clicks/impressions) | 4.70% | 4.85% | +0.15 pp | 0.040 | [+0.01, +0.29] pp |
| Saves per impression | 0.92% | 0.87% | -0.05 pp | 0.090 | [-0.11, +0.01] pp |
| Avg session time | 12.0 m | 12.1 m | +0.8% | 0.200 | [-0.3%, +1.9%] |
| Session crash rate | 1.20% | 1.32% | +0.12 pp | 0.010 | [+0.03, +0.21] pp |
| Exposed users | 200,300 | 199,700 | — | SRM p=0.62 | — |
In both cases: does the CTR win offset the negative guardrail (retention/complaints in Readout 1, crash rate in Readout 2)? If not, what mitigations or follow-ups would you require before full rollout?
Overview: A Pinterest Data Scientist technical-screen question on Analytics & Experimentation: design a rigorous A/B test for increasing video-pin share in the Home Feed — unit of randomization, primary metric and guardrails, power/MDE, sequential analysis, and a quasi-experiment fallback — then interpret two 14-day readouts and recommend ship, iterate, or stop. It rewards balancing engagement lifts (CTR) against guardrail regressions (retention, complaints, crash rate), multiple-testing robustness, and cross-unit (per-impression vs per-session) trade-offs.
Solution
A) Experiment design
1) Unit of randomization & exposure control
Randomize at the user level, not the session level. The Home Feed is personalized and carries state across sessions, so a user assigned to "more video" ought to receive that treatment consistently over time; this blocks within-user contamination (the same person oscillating between arms from one session to the next) and makes cross-session outcomes such as 7-day retention measurable. Randomizing by session would distort retention and other longitudinal metrics and produce mixed exposure within a single user.
Interference / spillovers to control for:
- Content-supply / two-sided interference: Exposing treatment users to more video can alter creator behavior and the training signals of the shared ranking model, which leaks into control. Counter this by confining the surfacing change to the serving layer for treatment users only, freezing model retraining for the duration of the test (or running a separate model), and considering cluster/geo-level randomization when supply effects are large.
- Ranking spillovers: When the feed reranks one shared candidate pool, raising video share for one cohort can change the inventory left for others. Adopt a treatment-only allocation policy for this purpose.
Exposure capping & ramp: Define the treatment as a target video share (say 45%) enforced as a per-session cap, so no single feed turns entirely into video (guarding against degenerate sessions). Ramp through 1% → 5% → 20% → 50% of users, checking guardrails (crash rate, complaints, latency, bandwidth) at every stage behind an automated kill switch.
2) Hypotheses, metrics & guardrails
Primary hypothesis (H1): Pushing video share toward the target raises high-quality engagement (saves per impression and/or time spent) without damaging reliability or trust.
Alternative hypotheses:
- H2 (Substitution): Video crowds out high-intent static pins, so clicks rise while saves/impression and downstream value fall.
- H3 (Novelty/fatigue): Video is novel and lifts CTR temporarily; the effect fades, which only a long-term holdout can reveal.
- H4 (Reliability cost): A heavier video load increases crash rate, jank, and bandwidth, offsetting engagement gains and hurting retention/complaints.
Primary success metric: saves per impression = saves / impressions. A save is a strong intent/quality signal on Pinterest and is tied to long-term value, whereas CTR invites clickbait and time-spent can be inflated by autoplay. Normalizing per impression keeps the metric robust to shifts in traffic volume. (saves_per_user_day is a reasonable alternative when a user-level rather than impression-level denominator is wanted; pre-register whichever you pick.)
Secondary / diagnostic metrics: CTR (clicks/impressions), time spent per user-day, video play starts and completion rate, creator follows.
Guardrails (hard, must-pass), with directions:
crash_rate(crashed sessions / sessions) — must not rise past a pre-set bound (e.g., absolute Δ ≤ +0.05 pp or relative ≤ +5%). Lower is better.complaint_rate = complaints / impressions— must not rise. Lower is better.7d_retention(returning users / exposed users) — must not drop. Higher is better.session_end_rate/ bounce — must not worsen. Lower is better.- Bandwidth/serving cost per user — watch for cost regressions. Lower is better.
Manipulation check: confirm the treatment really moved video share to the target (share of video impressions). If it did not, the readout is not measuring the intended change.
3) Power / MDE, duration, sequential analysis
- Fix alpha = 0.05 (two-sided), power = 0.80, 50/50 allocation. Derive per-arm sample size from the variance of the primary metric (use historical data; shrink variance with CUPED on a pre-period covariate). Solve for the MDE detectable at expected daily traffic, and choose a duration that (a) reaches that sample size and (b) spans at least one to two full weekly cycles to absorb weekday/weekend seasonality — commonly 14–28 days.
- Peeking / sequential looks: do not read p-values continuously against a fixed 0.05. Either use a group-sequential design (e.g., O'Brien–Fleming alpha-spending) or always-valid mSPRT confidence sequences when monitoring daily, so early looks do not inflate Type-I error.
- A/A test before launch to validate randomization and SE estimation (expect no significant differences).
- Novelty/fatigue: plan a minimum run and a long-term holdout (a small fraction kept at baseline for weeks/months) to separate a durable lift from a novelty spike.
4) Quasi-experiment (if RCT infeasible)
Use a staggered/phased rollout difference-in-differences with user (and time) fixed effects, comparing not-yet-treated against treated cohorts; add inverse-propensity weighting to balance covariates, or a synthetic control at the market/geo level.
- Identifying assumptions: parallel trends (treated and control move together absent treatment), no anticipation, stable composition (SUTVA at the chosen unit).
- Diagnostics: event-study/pre-trend plots, placebo (in-time and in-space) tests, covariate-balance checks, and sensitivity to the comparison set. Prefer modern staggered-adoption estimators (e.g., Callaway–Sant'Anna) over naive two-way fixed effects, which can be biased under heterogeneous treatment timing.
B) Interpreting the readouts
Readout 1 (lift-% table, N ≈ 2.0M)
- CTR +20.0% (p=0.010): significant and large — a clear lift in the engagement funnel.
- avg_session_sec +9.7% (p=0.040): significant, though possibly autoplay-inflated; corroborate with completion rate and saves.
- 7d_retention −3.6% (p=0.070): not significant at 0.05, yet the point estimate is negative — a serious early warning for a longitudinal metric. Underpowered retention is common at 14 days; do not wave it away.
- complaint_rate +30.0% (p=0.030): a significant deterioration in a trust guardrail.
Multiple testing: four metrics were inspected. Under Holm/Benjamini–Hochberg control (or Bonferroni α ≈ 0.0125 for 4 tests), CTR (0.010) survives, session-time (0.040) and complaints (0.030) are borderline, and retention (0.070) does not reach significance. The CTR win is robust; the complaint regression is borderline-robust; the retention drop is suggestive but underpowered.
Recommendation: do NOT ship as-is; iterate. A strong CTR lift is undercut by a significant rise in complaints and a negative (if not-yet-significant) retention trend — exactly the substitution/quality-degradation pattern (H2). Shipping a CTR win that costs trust and possibly retention is a bad trade. Mitigations and follow-ups below.
Readout 2 (absolute lifts + CIs)
- CTR +0.15 pp (p=0.040, CI [+0.01,+0.29]): significant but small (~+3.2% relative).
- Saves/impression −0.05 pp (p=0.090, CI crosses 0): not significant; negative trend (~−5.4% relative) — the primary quality metric is moving the wrong way.
- Avg session time +0.8% (p=0.200): not significant.
- Crash rate +0.12 pp (p=0.010, CI [+0.03,+0.21]): significant deterioration (~+10% relative) — a hard reliability guardrail breached.
- SRM p=0.62: no sample-ratio mismatch (good; the split is trustworthy).
Multiple testing (Bonferroni α ≈ 0.01 for ~5 metrics): CTR (0.040) would not survive correction; crash (0.010) sits at the threshold and stays concerning. So the CTR win is fragile under multiplicity while the reliability regression is robust.
Cross-unit caution: CTR and saves are per-impression while crash rate is per-session; you cannot net them directly. Frame the trade-off as expected value per session: EV ≈ w_click·clicks + w_save·saves + w_time·minutes − w_crash·crashes − w_complaint·complaints. The crash weight is large (crashes drive immediate abandonment, app-store rating damage, and churn). With saves trending negative and only a fragile CTR gain, the small CTR uptick does not offset a robust +10% relative crash increase.
Recommendation: do NOT ship; stop the current implementation and fix reliability, then rerun. A reliability guardrail failure supersedes a small, non-robust engagement gain.
Required mitigations & follow-ups (both readouts)
- Reliability (Readout 2): triage top crash signatures (OOM during decode, player lifecycle, GPU surfaces); lower-bitrate/shorter previews and deferred off-screen autoplay on older/low-memory devices; cap concurrent decodes; device/OS gating; real-time crash monitoring + kill switch.
- Quality / trust (both): rank-quality filters to surface high-quality video and exclude clickbaity/low-quality video that drives clicks without saves; cap video share for sensitive cohorts; investigate the complaint drivers (autoplay sound? irrelevant video?).
- Measurement: confirm the manipulation check (video share actually moved); pre-register the primary metric and guardrail thresholds; extend to 28 days + long-term holdout to test novelty decay and let retention reach power.
- Dose-finding: multi-cell test (0 / +5pp / +10pp / +20pp) to find the share that maximizes saves while keeping crash, complaints, and retention within guardrails.
- Heterogeneity & downstream: slice by device/OS/network, new vs returning, heavy- vs light-video users, content vertical, geo; measure day-1/day-7 retention and saves-to-return attribution for video vs static.
- Stats hygiene: use a pre-specified analysis plan and Holm/BH correction for secondaries; respect the sequential-testing boundaries.
Bottom line: In both readouts a CTR lift is paired with a guardrail regression (complaints/retention in Readout 1; crash rate in Readout 2) and no convincing high-quality-engagement gain. Do not ship: iterate on quality and reliability, add the missing checks, run a dose-finding test with a longer horizon, and re-evaluate against the pre-registered primary metric and guardrails.
Explanation Rubric: (1) chooses user-level randomization and reasons about two-sided/content-supply interference and r