Meta · Behavioral Stories
Demonstrate leadership under ambiguity
TrueInterview
October 7, 2026 · 5 min read
Talk about a high-stakes initiative where the priorities shifted partway through and you needed to sway others without formal authority. In your response: 1) Provide concrete background, the people involved, and the limitations; 2) Describe the hardest constructive feedback you gave (to a manager or a peer), the precise words you chose, and how you kept the working relationship intact; 3) Explain how you established trust fast with a new team (specific steps taken in the first two weeks); 4) Detail a disagreement you settled, including the compromises you made; 5) Put numbers on the results (business and engineering/science) and identify one measure that declined and why; 6) Looking back, what would you alter to get a better result?
Overview: This prompt assesses leadership amid uncertainty, persuasion without positional power, managing stakeholders, giving difficult feedback, resolving disputes, building trust quickly, and the capacity to measure and review project results for a Data Scientist in the Behavioral & Leadership area.
Example Answer
Sample, Organized Response (with instructional notes)
1) Background, people involved, and limits
- Scenario: I served as the data scientist for Notifications Relevance on three surfaces (social interactions, group posts, and event reminders) at a large consumer app. The objective was to cut notification fatigue (mute, unsubscribe, and complaint rates) by while keeping sessions initiated from notifications within .
- Midway change: After two weeks, a privacy/audit discovery mandated tighter consent gating for social-context features within three weeks. Leadership redirected priorities to: 1) deploy compliant suppression right away, and 2) prevent a drop in session starts. Our initial plan (a more complex ML re-ranker) became risky because of feature gaps and inference capacity.
- Stakeholders: the Notifications PM, the server Engineering Lead, the ML Lead, Integrity/Privacy legal counsel, the Infra capacity owner, and two partner PMs for Groups and Events. I held no direct authority over any of them.
- Constraints: an experiment freeze at quarter-end in week 7; constrained online inference capacity; missing feature logging; significant reputational risk if compliance failed; decision guardrails: session starts , complaint rate down, opt-out rate down. Teaching note: State the stakes, timeline, and guardrails clearly so that later trade-offs sound believable.
2) Hardest constructive feedback (precise wording) and preserving the relationship
- Scenario: The PM suggested launching a blended variant that used open rate as the success measure, with a two-week A/B test that, under the new consent gating, would lack power for our guardrail (session starts). I had to push back to a senior colleague.
- The exact words I used in a one-on-one and then repeated in the team document: "I worry that we are optimizing for the wrong metric and risking a false positive. Given the new consent limits, open rate is a weaker stand-in for session starts. Our power calculation indicates only about power to detect a shift in sessions over two weeks. I cannot endorse shipping on this design. I suggest we change the north star to net session starts with opt-outs as a guardrail, lengthen the runtime or broaden markets to reach power, and run a smaller parallel canary for compliance."
- How I preserved the relationship:
- I led with data: a one-pager containing the power analysis, sensitivity analysis, and a pre-registered decision rule.
- I provided a plan, not only criticism: I laid out two scope choices with timelines and risks.
- I gave credit and demonstrated flexibility: I retained the PM's qualitative insights in feature selection and asked her to present the revised plan, while I served as the technical backup. Teaching note: Quote brief, respectful wording, then demonstrate how you combined critique with solutions and credit.
3) How I established trust fast (actions in the first two weeks)
- Held fifteen 30-minute one-on-ones (PMs, Engineering, Integrity, Infra) to chart incentives, constraints, and decision-makers.
- Authored and circulated a two-page "Metrics & Guardrails" document: definitions for session starts, opt-outs, complaint rate, delivery latency, and a pre-registered decision rule.
- Delivered a quick win: corrected a logging inconsistency in notification open attribution that was overstating opens by roughly percentage points; released a shared dashboard with cohort breakdowns and alerts.
- Built a prototype offline-to-online backtest comparing simple rules against a lightweight re-ranker under consent gating; shared the trade-offs between lift and latency.
- Established weekly office hours and a decision log so that disagreements were recorded and settled openly. Teaching note: Actions in the first two weeks should blend relationship-building, instrumentation cleanup, and a concrete analytical win.
4) Disagreement settled and trade-offs accepted
- Disagreement: Infra wanted an immediate hard throttle ( sends) to ensure compliance before the freeze. PM/ML favored the new re-ranker, which needed scarce online feature joins and would increase p95 latency.
- My proposal (a compromise):
- Phase 1 (2 weeks): rule-based suppression for sensitive categories to satisfy compliance and lower fatigue; no online joins.
- Phase 2 (canary to ): a lightweight re-ranker using only existing features (no new joins), with an explicit guardrail.
- Trade-offs accepted:
- We sacrificed about estimated extra lift compared with the heavier model to remain within latency and capacity.
- We narrowed scope to two surfaces at first, postponing Events to the next quarter. Teaching note: Identify the opposing sides, present the compromise, quantify what you surrendered, and explain why it was sensible given the constraints.
5) Measured impact (business and engineering/science), plus one metric that declined and why
- Business results (90 days after rollout):
- Total notifications sent: (target ).
- Opt-out rate: relative.
- Complaint rate: relative.
- Sessions started from notifications: (inside the guardrail).
- Reactivation (30-day dormant users): no meaningful change (, ) after adding a small exception list.
- Engineering/science results:
- Model AUC: absolute versus baseline rules (offline), which translated to sessions in the treatment where applied.
- Experiment design: pre-registered analysis, CUPED used for variance reduction (about variance reduction on sessions).
- Latency: p95 delivery latency rose by (from to ) because of online scoring; still under the guardrail.
- Metric that worsened and why:
- p95 delivery latency worsened () because even the lightweight re-ranker added feature fetch and scoring time. We accepted this trade-off within a set budget to meet compliance and fatigue targets. Teaching note: Distinguish business impact from engineering/science impact. Choose a specific, defensible downside and explain its cause and guardrails.
6) Looking back: what I would change to get a better result
- Negotiate capacity earlier: escalate for temporary inference capacity two weeks sooner; that likely would have enabled the fuller-featured model and an additional roughly sessions lift without exceeding latency.
- Automate power checks: release a template that calculates MDE and required duration when an experiment is created, preventing underpowered designs.
- Segment-specific guardrail: establish a stricter reactivation guardrail in advance (for example, for 30-day dormant users), which would have cut the iteration cycles needed to adjust exceptions.
- Instrumentation parity test: run a formal offline-to-online consistency suite before the canary to reduce the debugging time spent on feature drift. Teaching note: Hindsight should be actionable (process, tooling, or earlier escalation), not merely "do more."
How to adapt this pattern for your own responses:
- Use STAR but map it explicitly to the six prompts above.
- Put numbers on both the upside and the controlled downside using guardrails.
- Demonstrate influence tactics: clear memos, pre-reads with options and trade-offs, decision logs, and credit-sharing.
- Verify rigor: power analysis, pre-registration, variance reduction (such as CUPED), guardrail metrics, and latency/capacity budgets.