Roblox · Behavioral Stories
Describe resolving revenue–UX metric conflict
TrueInterview
October 7, 2026 · 7 min read
Describe a time when you were in charge of a high-stakes decision where advertising revenue goals conflicted with user experience metrics. Include specifics: 1) The precise metrics that were in opposition (for instance, RPM per session versus session duration, bounce rate, D7 retention) along with their starting values and targets; 2) The explicit thresholds or guardrails you established and the reasoning behind them; 3) The decision-making structure you used (stakeholders, alignment approach, who owned the decision, timeline); 4) The experiment or analysis you conducted, including how you mitigated risk and your plan if early guardrails were violated; 5) The ultimate decision and the measurable impact over at least two time periods; 6) One error you made and how you would alter your process for next time.
Overview: This question assesses a data scientist's ability in metrics-driven product leadership, particularly around balancing monetization with user experience, performing quantitative trade-off analyses, designing experiments, managing risk, and making decisions with cross-functional stakeholders.
Solution
A STAR-format model response: Balancing monetization and user experience on a user-generated content gaming/social platform
Situation
The goal was to boost ad revenue before an important quarter without harming core engagement. The product team proposed raising ad load on the home feed and introducing one interstitial ad during transitions between worlds. The danger was that this would harm short-session users, new users, and teens, which are critical groups for long-term retention.
Task
As the senior data scientist responsible for ads, I was accountable for the decision framework: specifying success metrics and safety thresholds, creating the experiment design, executing a phased rollout with ongoing monitoring, and providing a go/no-go recommendation along with targeting guidance.
Action
1) Conflicting metrics (with starting values and goals)
- Monetization metrics
- Revenue per ad session (ARPS): starting value $0.21; goal increase of 6–8%.
- Ad impressions per session: baseline 5.2; proposed increase of 10–15%.
- Fill rate and RPM (revenue per thousand impressions) were tracked to detect changes in budget or ad quality.
- UX/safety metrics
- Session duration: starting point 14.2 minutes; tolerable change at least −2.0%.
- Bounce rate (sessions under 60 seconds): baseline 17.5%; acceptable change at most +0.5 percentage points.
- D1 retention: starting at 41.8%; acceptable change no lower than −0.3 percentage points.
- D7 retention: baseline 13.2%; acceptable change no lower than −0.2 percentage points.
- User complaints per 10,000 sessions: baseline 8.1; acceptable increase at most 10%.
- App crash rate: baseline 0.33%; tolerable increase at most +0.03 percentage points. Definitions
- ARPS is total ad revenue divided by total sessions:
- Dd retention is the probability that a user is active on day d given they installed or were active on day 0:
- A 30-day ad LTV estimate is given by:
2) Guardrails and rationale
- Bounce rate: limit of +0.5 percentage points. Rationale: historically, a +0.5pp increase leads to a −0.15pp change in D7 and a 1–2% drop in LTV over 90 days.
- Session length: a maximum decrease of 2% to protect the creator economy and recommendation quality.
- D1 and D7 retention: allowable decreases of 0.3pp and 0.2pp respectively, since these are significant at our scale and closely tied to LTV.
- Complaints: no more than +10% to prevent erosion of user trust and avoid spikes in moderation workload.
- Crash rate: cap of +0.03pp to ensure technical stability.
- Policy: completely exclude users under 13 from the new interstitial format for compliance and brand safety.
3) Decision structure
- Stakeholders
- Monetization product manager (who proposed the change)
- Core user experience product manager (responsible for session metrics)
- Trust and safety team (policy and complaint handling)
- Ads operations and brand safety (creative quality assurance)
- Data engineering and experimentation platform (instrumentation)
- Finance (forecasting)
- Alignment and ownership
- Decision owner: General Manager of Engagement and Monetization.
- RACI: Monetization PM as Responsible, GM as Accountable, Data Science/Trust & Safety as Consulted, Engineering/Ads Ops as Informed.
- Artifacts: a six-page document with pre-read materials, risk register, and pre-registered guardrails and stopping rules.
- Timeline
- Weeks 0–1: define metrics and guardrails, conduct power analysis, write specification.
- Weeks 2–3: instrumentation and dark launch.
- Weeks 4–6: staged ramp-up with sequential monitoring.
- Week 7: decision and rollout plan.
4) Experiment and analysis
- Design
- Treatment groups: control; variant A with a 15% increase in feed ad load; variant B with the same feed increase plus one interstitial during teleportation (displayed once after 60 seconds, capped at one per 10 minutes, and a maximum of six ad exposures per session).
- Randomization was at the user level, stratified by platform (mobile/desktop), geography, and account age. We applied CUPED (controlled pre-experiment data) using the prior 14 days of session metrics to reduce variance.
- Power/variance
- Minimum detectable effect for ARPS was 3%; for bounce rate and D7 retention it was 0.15pp. The required sample size was approximately 1.5 to 2.0 million users per arm to achieve 80–90% statistical power.
- Ramp plan and sequential monitoring
- We used a phased rollout: 1% dark launch (instrumentation only), then 5%, 10%, 25%, and 50%, with 24 to 48 hour hold periods between stages.
- We applied group-sequential boundaries (O'Brien-Fleming) to detect harm on guardrails, and mSPRT to monitor ARPS uplift.
- Risk mitigation
- Enabled creative quality assurance and blocklists; applied policy filters for sensitive user segments; excluded under-13 users from the interstitial.
- Implemented a kill switch per cohort (new users, teens, specific geographies).
- Maintained a permanent 1% safety holdout for backtesting.
- Breach playbook (pre-registered)
- If bounce rate increased by +0.5pp in any protected cohort (new users defined as account age less than 7 days, or teens), we would immediately turn off the interstitial for that cohort and continue monitoring the feed-only variant.
- If D7 retention was projected to drop by 0.2pp, we would pause the ramp and require retuning of frequency caps and creative mix.
5) What happened during the test
- During the 10% ramp stage (within 24 hours), we observed a cohort-specific breach.
- New users experienced a +0.9pp increase in bounce rate in variant B (with interstitial). We executed the playbook: disabled the interstitial for users with account age less than 7 days and kept the feed-only variant A for them. Within 48 hours, the bounce rate difference normalized to +0.2pp.
- Main results at the 50% ramp stage (days 7–14) were as follows.
- Variant A (feed +15%): ARPS increased by 4.1% (95% confidence interval: +3.4 to +4.8), session length decreased by 0.6%, bounce rate increased by 0.1pp, and D7 retention decreased by 0.05pp (not statistically significant).
- Variant B (feed + interstitial, excluding new users): ARPS increased by 8.3% (CI: +7.5 to +9.1), session length decreased by 1.3%, bounce rate increased by 0.3pp, and D7 retention decreased by 0.15pp (borderline but within the −0.2pp guardrail). Complaints increased by 6% (within cap); crash rate unchanged.
- LTV signal: Using retention multiplied by ARPDAU, we estimated a 30-day LTV uplift of +2.2% for variant A and +3.0% for variant B among eligible users (account age 7 days or more, and age 13 or older).
Result
Final decision
- Deploy variant B to eligible users (13+ and account age at least 7 days) with caps: no more than one interstitial per session, a maximum of six total ad exposures per session, and dynamic throttling for short sessions (under 5 minutes) to protect user experience.
- Deploy variant A to new users and short-session users; no interstitial for these cohorts.
Quantified impact (two horizons)
- Short-term (first 30 days after rollout)
- Monetization: blended ARPS increased by 5.9% across all eligible users; incremental revenue was +$3.8 million compared to control forecast (this is scale-dependent and normalized by traffic).
- UX: session length decreased by 1.1%; bounce rate increased by 0.2pp; D7 retention decreased by 0.08pp. All within guardrails.
- Medium-term (90 days)
- Monetization: blended ARPS increased by 5.1% (a slight regression due to creative fatigue, mitigated by rotating creatives), and 90-day LTV increased by 2.6% on eligible cohorts.
- UX: D7 retention stabilized at −0.05pp; complaints returned to +3% with improved creative quality assurance; no impact on crash rate. Business takeaway: The hybrid approach secured the majority of the revenue uplift (+5–6% ARPS) while safeguarding long-term engagement through cohort-specific eligibility and frequency limits.
Mistake and what I’d change
- Mistake: I initially monitored guardrails only at the overall level and by a simple new vs. existing user split, which obscured a larger bounce increase among 13–17-year-old short-session users who fell into the 'existing' bucket. We detected it at the 10% ramp, but it could have been flagged earlier.
- Changes for next time:
- Predefine segment-level guardrails for protected cohorts (teens, short-session users) with hierarchical monitoring and alerting.
- Require no material harm in each key cohort before ramping beyond 10%.
- Add a pre-experiment observational backtest to tune interstitial timing for short sessions.
Why this approach generalizes (guardrails and validation)
- Putting guardrails first forces explicit clarity on acceptable user experience risk.
- Cohort-specific rollout preserves value while avoiding harm to sensitive user groups.
- Sequential testing with pre-registered stopping rules prevents p-hacking and limits potential negative impact.
- CUPED and stratification reduce variance and enable faster decisions without compromising rigor. Formulas used
- Common pitfalls to avoid
- Watch for Simpson’s paradox across cohorts; always monitor results by segment.
- Be aware of ad novelty and creative fatigue; plan creative rotation and re-measure at 60–90 days.
- Interference: prevent cross-over by using user-level randomization and frequency caps.
- Seasonality or budget shifts: use holdouts and RPM/fill rate diagnostics to distinguish demand shocks from UX effects.