ByteDance · Project Deep Dive
Reflect on a challenging project you led
TrueInterview
October 7, 2026 · 5 min read
Walk through a project you owned from start to finish that had a real effect on a product choice. Include specifics: (a) the unclear problem and the success measures you defined; (b) the riskiest assumption you ruled out, how, and when; (c) a concrete change of direction you made after seeing new data; (d) a high-stakes conflict and how you built alignment; (e) the trade-offs you accepted under time/resource limits and why; (f) what you would change if you began again on 2025-09-01.
Overview This question tests leadership, cross-functional influence, and applied data science ability—specifically ambiguity handling, causal measurement, hypothesis testing, and moving product decisions.
Solution
Worked example for teaching (end-to-end, product-changing)
Project: Cold-start feed seeding for new users to lower early churn. Summary: We assessed whether to add an “interest selection” onboarding step or keep silent personalization for the first-session feed. The data indicated that the extra friction hurt activation and retention for most users. We shifted to silent personalization with light exploration, and reserved explicit interest selection for a small, high-intent segment. This moved the product plan away from a global interest-selection launch toward a segmented approach.
(a) Ambiguous problem and success criteria
- Ambiguity: New users were leaving during their first session. Product managers wanted an interest selection screen to improve personalization; design was concerned about added friction; engineering warned about cold-start latency. The unresolved question was whether explicit interests would improve personalization enough to justify the extra steps.
- Decision options: (1) Launch interest selection for all new users; (2) Keep the current silent personalization; (3) Use a hybrid or segmented approach.
- Primary success metric: D7 retention lift. Baseline around 20%; minimum detectable effect (MDE) = +0.5 percentage points (pp).
- Secondary/guardrails: D1 retention, new-user session length, hide/report rates, push opt-out, creator complaint rate, app uninstalls within 24 hours, and system health (latency/crashes).
- Sample sizing (illustrative): For a two-arm A/B with baseline , , , power : new users per arm.
- Experiment design: User-level randomization; pre-registered analysis; SRM check; CUPED (if available) to reduce variance; 14-day run to observe D7 and early D14 trends.
(b) Riskiest assumption invalidated (how and when)
- Riskiest assumption: “Explicit interest selection improves short- and medium-term retention after accounting for the added friction.”
- How we tested:
- Offline replay: We replayed historical new-user sessions to simulate explicit interest picks by mapping early swipes to topic clusters, and estimated an upper-bound benefit of “perfect” picks on feed relevance.
- Rapid funnel experiment (Week 1): We randomized 20% of new installs into a prototype interest-selection step requiring 3 picks. We measured activation completion and time-to-first-video.
- Full A/B (Weeks 2–3): Interest selection versus control. Primary metric: D7. Guardrails: D1, hides/reports, latency.
- Finding: The activation drop (−1.8pp completion, +6s time-to-first-video) outweighed the personalization gains for most traffic. Early D1 was −0.3pp; D7 ATE stayed around −0.2pp, with HTE showing gains only in a small high-intent segment (for example, users arriving through creator-linked acquisition).
- Timing: invalidated by Day 10, after 7 days of accumulation and a pre-registered interim analysis.
(c) Concrete pivot based on new data
- Pivot: We dropped the global interest-selection rollout. We shipped silent personalization plus lightweight exploration for all new users, and limited explicit interest selection to a narrow high-intent slice.
- What changed specifically:
- Feed seeding became a hybrid: a trending-but-diverse starter set, stronger early exploration, and fast adaptation to the first 5–10 interactions.
- We removed the mandatory “pick 3 interests” step for general traffic; interest selection stayed as an optional interstitial only for high-intent referrals.
- Result (illustrative): +0.6% increase in quality-adjusted watch time in the first session, +0.4pp D1, +0.5pp D7 versus baseline, no significant increase in hides/reports, and stable latency. This cleared our pre-set launch bar.
(d) High-stakes disagreement and how I earned alignment
- Disagreement: PM and Marketing wanted a global interest-selection launch tied to a brand campaign timeline. Eng and Data were worried about friction and latency risk.
- Alignment approach:
- Pre-registered metrics and decision thresholds to prevent moving the goalposts.
- Transparent interim readouts, SRM checks, and HTE by acquisition channel/locale.
- Simulations showing that even “perfect” interest picks could not recover the activation losses for most cohorts.
- Proposed a compromise: a segment-specific rollout where the effect was positive, plus silent personalization elsewhere.
- Outcome: Cross-functional sign-off to shift the product plan from global to segmented rollout, with updated experiments to monitor long-term retention.
(e) Trade-offs under time/resource constraints (and rationale)
- Time constraint: 3 weeks to decide before campaign creative locked.
- Trade-offs made:
- Shipped a minimal seed model (logistic regression on simple signals such as locale, device language, time-of-day) instead of a deeper model, to meet the timeline and keep latency below the p95 target.
- Limited experiment scope: 20% traffic for the prototype funnel test; then a 50/50 split only for new installs in the top 3 markets to reach sample size quickly, deferring long-tail locales.
- Reduced instrumentation to the critical events to avoid client release delays; deferred nice-to-have surveys to post-launch.
- Accepted modest infra cost for early prefetch during the first session to preserve experience quality.
- Rationale: Optimize for decision quality on the primary metric (D7) while staying inside latency and infra guardrails.
(f) What I’d do differently if starting again on 2025-09-01
- Plan a 28-day holdout for long-horizon retention and satisfaction stability, with sequential testing/alpha spending to allow earlier safe stops.
- Use off-policy evaluation (IPS/DR) from logged bandit data to down-select seeders before any user-facing test.
- Apply CUPED and variance reduction by default to reduce sample/time needs.
- Build a self-serve dashboard with pre-registered metrics and HTE slices to speed alignment.
- Introduce fairness/diversity guardrails earlier (for example, content category and creator exposure diversity in cold-start).
- Start privacy/compliance review earlier for any new signals considered for personalization.
Validation and guardrails applied
- Randomization: user_id hashing; monitored SRM; no leak from acquisition channels.
- Metrics health: Monitored D1, D7, hides/reports, app crashes, latency p95; creator-side impact (new creator exposure share).
- Heterogeneous treatment effects: Segments by acquisition source, locale, device class; only the high-intent cohort benefited from explicit interest selection.
- Power analysis and MDE pre-registration to avoid overfitting to noise.
- Post-launch monitoring: 4-week KPI drift checks; rollback plan documented.
Key takeaways
- Ambiguity resolution requires pre-registered success criteria and HTE, not just topline ATE.
- The highest risk was friction versus personalization benefit; we invalidated it quickly with a funnel test and a short A/B.
- Data supported a pivot to a segmented strategy and silent personalization, materially changing the original product decision.