SoFi · Statistics & Data Analysis
Plan and validate ranking experiment
TrueInterview
October 7, 2026 · 1 min read
You are tasked with a new home-page ranking algorithm and need to validate it without introducing risk. Create a three-stage evaluation plan: offline replay with IPS/DR, small-scale interleaving using team-draft, and then a full A/B test. Be specific: (1) Define the exposure unit—impression-level versus session-level—and the bucketing approach to avoid contamination across sessions and devices. (2) The primary metric is 30-day funded-account conversion per 1,000 impressions, with a baseline of 1.20%, a target relative uplift of +5%, power of 0.8, and alpha of 0.05. Compute the per-arm sample size assuming independent impressions, then discuss inflation due to repeated exposures and cluster-robust variance. (3) List guardrails such as p95 latency, app crash rate, CS tickets, and decline rate, and describe how you would set sequential boundaries—for example, alpha spending or SPRT—to allow early stopping without inflating Type I error. (4) Explain how you would mitigate novelty effects, carryover, and seasonality; specify the ramp policy and duration for capturing 30-day outcomes while using proxy metrics for early reads with CUPED or covariate adjustment. (5) Describe the heterogeneous treatment effect analysis across new and existing users and credit tiers, and how you would control false discovery using BH or Holm. (6) Provide a plan for detecting p-hacking and Simpson's paradox, and define ship criteria for cases where the primary and guardrail metrics disagree.
Overview: This question tests experimental design and analytics skills, spanning offline counterfactual replay, interleaving and A/B testing, sample-size and power calculations, sequential testing and alpha spending, guardrail monitoring and ramp policies, proxy metrics and covariate adjustment, heterogeneous treatment effect analysis, and governance issues such as p-hacking and Simpson's paradox in the Analytics and Experimentation area for Data Scientist roles. It is often used to probe how well candidates can rigorously validate ranking changes while balancing statistical error, operational risk, and bias mitigation, and it prioritizes practical application of applied statistical concepts and experiment governance rather than purely theoretical knowledge.