Meta · Statistics & Data Analysis
Estimate ads ranking revenue impact
TrueInterview
October 7, 2026 · 8 min read
You work as the data scientist on an ads ranking team at a major social platform. The team has developed a fresh ranking algorithm for feed ads. Compared with today's production ranker, it reorders ads by weighting bid, predicted click-through rate (pCTR), predicted conversion rate (pCVR), and ad quality in a different way. An early ramp points to higher revenue per daily active user (DAU), yet the team fears the short-term gain may not reflect medium-term impact: users could adjust to the changed ad mix, advertisers could revise bids or budgets, and auction dynamics could move. Your job is to design a way to estimate the medium-term revenue impact of shipping the new ads ranking algorithm across a 4- to 8-week horizon, and to recommend whether to launch.
Constraints & Assumptions
- Platform size: a large eligible DAU base (Part 5 supplies the exact headcount for the scale-up); revenue comes from an ad auction where advertisers set daily or lifetime budgets that are paced out over time.
- The intervention is a feed ranking change that users experience individually, yet it operates inside a shared marketplace (advertiser budgets are pooled across users).
- A brief (say, 2-day) ramp already produced a positive revenue-per-DAU signal; what remains open is whether the lift holds, fades, or flips sign over weeks.
- Standard experimentation tooling is available to you (A/B tests, geo holdouts), along with pre-period covariates and advertiser-side delivery and pacing data.
- "Medium-term" here means capturing user adaptation, advertiser reactions in bids and budgets, and at least one complete advertiser budget cycle.
Clarifying Questions to Ask
- Are you estimating revenue at full rollout (the launch decision) or only the effect inside the experiment population? That choice decides whether spillover counts as bias or as part of the estimand.
- How long is the relevant advertiser budget cycle, and about what share of revenue comes from advertisers that are budget-constrained?
- Which guardrail thresholds (retention, negative feedback, advertiser ROAS) block a launch, and which are simply monitored?
- Is a usable geo/market holdout available, plus pre-period covariates for variance reduction (CUPED, for example)?
- Does the business already have a revenue or LTV model for valuing medium-term user and advertiser effects?
- What statistical power and minimum detectable effect (MDE) can each candidate randomization unit deliver?
Part 1 — Causal estimand
Pin down the exact causal quantity you want to estimate: the population, the comparison, the time horizon, and the outcome(s) it should cover.
Hint — What an "estimand" is here: Spell out all four components of an estimand: population (eligible users or markets), treatment versus counterfactual (new ranker against keeping the current ranker, evaluated at full rollout), horizon (4-8 weeks), and outcome. Ask which fits the business question better — a per-impression metric or a per-user / cumulative one. Hint — What to encompass: A ranking change's "revenue impact" is more than the auction revenue it produces right away. Think about whether the estimand should also absorb medium-term marketplace effects — advertiser budget reallocation, future ad inventory from users who stay, and advertiser value (conversions / ROAS) — since those loop back into revenue.
What This Part Should Cover
- A precise estimand that names the population, the counterfactual (full rollout versus status quo), the 4-8 week horizon, and the outcome.
- Acknowledgement that a launch decision calls for a full-equilibrium estimand rather than the partial-equilibrium effect on a small treated slice.
- An outcome wider than per-impression revenue, with an understanding of how per-impression revenue can be gamed.
Part 2 — Experiment or quasi-experiment design
Lay out the design you would run: the randomization unit, the reason for picking it, and how long the test has to run. Explain the tradeoff between user-level and market-level (geo) randomization when the product shares one auction.
Hint — Randomization unit tradeoff: Begin with the no-interference (SUTVA) assumption and ask whether it survives when treated and control users pull from the same advertiser budgets. Balance the statistical power of finer units against the contamination they permit. Hint — Duration: The horizon has to span the feedback loops that concern you: weekly seasonality, advertiser budget and pacing cycles, bid re-optimization, and user adaptation. A 2-day ramp catches none of them.
What This Part Should Cover
- A randomization choice justified through the interference tradeoff, not a generic "run an A/B test."
- A clear comparison between user-level (more power, a clean UX read, but spillover-biased) and geo/market-level (absorbs the marketplace, but underpowered).
- A duration covering at least one complete advertiser budget cycle, with weight on the later, post-transient weeks.
Part 3 — Primary, secondary, and guardrail metrics
Propose a metric tree: one primary metric (or co-primary pair), secondary and diagnostic metrics that explain the mechanism, and guardrail metrics covering both users and advertisers.
Hint — Picking the primary metric: Favor a metric that holds up under impression-mix shifts. Revenue per impression can climb even as total revenue per user drops (fewer or worse impressions). A per-eligible-user or cumulative-over-window revenue metric is tougher to game. Hint — Diagnostics vs. guardrails: Secondary metrics should let you explain a revenue movement (eCPM, ad load, fill rate, win price, pCTR/pCVR calibration, budget exhaustion time). Guardrails shield you from regressions you would not accept in exchange for revenue (engagement, retention, negative feedback, advertiser ROAS/CPA, latency).
What This Part Should Cover
- A primary metric that resists impression-mix gaming (per-eligible-user or cumulative revenue, not per-impression).
- Diagnostic metrics picked to explain mechanism (price versus volume, calibration, pacing) rather than merely describe.
- Guardrails on both the user side (engagement, retention, negative feedback) and the advertiser side (ROAS, CPA, churn), along with system health.
Part 4 — Auction interference, budgets, seasonality, and heterogeneity
Explain how you would deal with each of the four validity threats below, and how each one could bias a naive user-level read:
- Auction interference / spillover (advertiser budgets shared between treatment and control).
- Advertiser budget constraints and pull-forward (spending faster today does not mean spending more across the window).
- Seasonality (day of week, holidays, campaign cycles).
- User-level heterogeneity (the effect differs by market, tenure, engagement, ad load).
Hint — Interference: Ask which advertisers actually transmit the spillover, and whether a design that absorbs the marketplace (a geo holdout, say) could sanity-check the user-level number. Consider the direction of the bias. Hint — Pull-forward and seasonality: Track the trajectory across the window rather than a single day, and interpret what the budget-pacing diagnostics report. Concurrent randomization and pre-period covariates help as well. Hint — Heterogeneity: Choose the segments that matter before looking, so a heterogeneous effect becomes a pre-registered finding instead of a fishing expedition.
What This Part Should Cover
- For each threat: the mechanism, the direction of bias it creates in a naive user-level read, and a concrete mitigation.
- Spillover treated as the headline threat, with constrained versus unconstrained advertiser slices and a geo check.
- Pull-forward separated from real lift using cumulative revenue and pacing/exhaustion diagnostics; a named variance-reduction technique for seasonality; pre-registered segments for heterogeneity.
Part 5 — Translating to company-level revenue impact
Starting from a per-user effect estimate with a confidence interval, show how you would scale it to a company-level revenue figure for the window, and which corrections or caveats you would attach. For the scale-up, use an eligible population of 200M DAU and a 28-day window.
Hint — Scaling and caveats: Multiply the per-user-per-day lift by eligible DAU and by the number of days, and carry the confidence interval through that same multiplication (round only at the end). Then discount or flag the point estimate for known biases — budget pull-forward, advertiser ROI harm, UX-driven inventory loss, and any measured experiment spillover.
What This Part Should Cover
- Correct scale-up arithmetic:
- A confidence interval carried through the same multiplier (not just a point estimate).
- Honest caveats that shift the headline number: pull-forward, ROI harm, inventory loss, and spillover (reconciling geo with user-level).
Part 6 — Launch recommendation under a UX-guardrail conflict
State the decision logic you would apply when short-term revenue is positive but certain user-experience guardrails deteriorate. Make the tradeoff explicit instead of defaulting to "ship" or "kill".
Hint — Make the tradeoff quantitative: Separate a cosmetic guardrail shift (a small uptick in ad hides, say) from a load-bearing one (a measurable 28-day retention loss). Frame the decision in terms of long-term user and advertiser lifetime value, and treat ramp or segment-targeting as middle options rather than only full launch versus no launch.
What This Part Should Cover
- An explicit decision rule with three outcomes (launch / ramp-or-target / do-not-launch), not a binary.
- The conflict settled against long-term value (user and advertiser LTV), not the dollars in this window.
- A concrete contrast between an acceptable guardrail move and one that blocks launch.
What a Strong Answer Covers
These dimensions run across every part and should show up throughout the answer:
- Keeping interference central. The candidate never loses sight of the fact that a per-user intervention priced inside a shared, budget-constrained auction breaks SUTVA, and lets that drive the design, the metrics, and the company-level correction.
- Cumulative-over-window thinking. Revenue is interpreted as a trajectory spanning a full budget cycle, never as a one-day or per-impression snapshot.
- Quantitative honesty. Estimates come with intervals, caveats are subtractive (they move the number), and the final recommendation rests on long-term user and advertiser lifetime value rather than the immediate lift.
Follow-up Questions
- Say the user-level A/B reports +2.4% revenue while a concurrent geo holdout shows roughly 0%. How do you reconcile the two, and which do you trust for the launch decision?
- Week 1 shows a +5% lift but week 4 shows +1%, still positive. How do you separate budget pull-forward from genuine user adaptation, and how does that change your company-level estimate?
- Advertiser ROAS dips slightly in treatment. Why might that eat away at the revenue lift over a horizon longer than 8 weeks, and how would you watch it after launch?
- How would you set up a long-term holdout that keeps measuring the shipped change after the experiment closes?
Overview: This question tests competency in causal inference, experimentation design, metric construction, and marketplace economics for measuring ad ranking revenue within the Analytics & Experimentation domain.