Snowflake · Project Deep Dive
Present an end-to-end project and defend decisions
TrueInterview
October 7, 2026 · 6 min read
You have ten minutes, and no more than five slides, to present a project you led from start to finish that reached production users. Be sure to cover the problem setting, what stakeholders wanted, the data you used, the modeling or analysis, the key decisions, outcomes, and trade-offs. Then address the following:
- What was the one metric you optimized for, what guardrails did you define, and why? Tell me about an occasion when that metric was at odds with a metric another stakeholder cared about, and how you settled the disagreement.
- Where did things break? Offer one specific mistake—say, a bad metric trade-off or an incorrect assumption—and what you changed in response.
- Suppose leadership rejects the plan because of metric worries (for instance, retention improves but revenue declines). Outline a revised follow-up experiment or rollout that handles those worries without changing the schedule.
- How did you work with data engineering, product, and design? Give a concrete example of bargaining over scope or data model changes under deadline pressure.
Overview: This question tests end-to-end project leadership for a data scientist, along with product and experiment design, metric choices, trade-off analysis, and cross-functional negotiation.
This prompt comes from a Snowflake Data Scientist interview experience.
Solution
Sample Five-Slide Talk Track: Adaptive Query Acceleration (AQA) for a B2B Data Platform
Context: Users complained that analytics queries were slow during peak times. We developed and shipped an "Adaptive Query Acceleration" feature that automatically adjusts compute sizing and applies conservative optimizations to heavy queries to bring down tail latency.
Slide 1 — Problem & Stakeholders
- Problem: P95 query latency rose sharply in peak hours, generating support tickets and increasing churn risk for mid-market accounts.
- Why now: Seasonal traffic growth led to more frequent SLO violations, and competitors were pushing "instant analytics."
- Stakeholders & goals:
- Users/CS: quicker queries, fewer timeouts, fewer tickets.
- Product/PM: increase adoption and retention for analytics workloads.
- Finance/RevOps: prevent any significant loss of consumption revenue.
- Infra/DE: maintain stable error rates and avoid capacity thrash.
- Success criteria (at launch):
- Primary: cut P95 latency by at least 15% while keeping error rate flat.
- Guardrails: error rate rise of no more than 0.05 percentage points; queue wait time no worse; credits per 1,000 queries must not fall more than 15% to protect revenue.
Slide 2 — Data Sources & Instrumentation
- Data sources:
- Query logs:
query_id, start and end, bytes scanned, spills, retries, error code. - Warehouse telemetry: size, concurrency, queue wait time, cache hit rate.
- Billing/usage: credits burned per query and per account-day.
- Support tickets: topic, account, timestamp used to correlate incidents.
- Account metadata: segment, commitment tier, prior churn signals.
- Query logs:
- Instrumentation added:
- Durable query-to-warehouse join keys; optimization decisions tagged with feature flags, chosen action, and confidence.
- P50/P95 latency and queue time aggregated per account-day; pre/post baselines to support CUPED variance reduction.
- Experiment design:
- Randomization unit: account×warehouse, chosen to reduce interference.
- 50/50 split, four-week duration, with holdouts for high-value accounts.
Slide 3 — Modeling & Policy
- Goal: Choose an action that minimizes tail latency while leaving guardrails intact.
- Predictive modeling:
- Features: time-of-day, historical log-latency, concurrency, query complexity (bytes scanned and joins), and spill signals.
- Model: gradient-boosted trees for log-latency prediction, with quantile loss aimed at the tail (p95); a separate model for credits per query.
- Decision policy (cost-aware optimization):
- Objective: minimize subject to , where equals baseline credits times .
- Implementation: , with tuned through offline replay; reject outright when predicted error rate increases.
- Exploration:
- Gentle perturbations ( resize step) and an automatic kill switch if guardrails are violated for an account-day.
Slide 4 — Key Decisions & Results
- Key product decisions:
- Rolled out at account×warehouse granularity to avoid noisy cross-traffic effects.
- Optimized for P95 tail rather than P50 to match how users feel performance.
- Auto-applied only safe actions; all other actions appeared as recommendations requiring user confirmation.
- Results (roughly 600 account×warehouse units and 15M queries over 4 weeks):
- P95 latency dropped from 8.3s to 6.4s, a 23% reduction with a 95% CI of 20% to 26%.
- Queue wait time fell 15%.
- Error rate moved from 0.19% to 0.21%, a non-significant rise of 0.02 percentage points.
- Credits per 1,000 queries fell 11%; Finance flagged the near-term revenue impact.
- Downstream business: 90-day retention improved 1.5 percentage points as an early leading indicator, and treated accounts had 18% fewer support tickets.
- Trade-offs:
- Better user experience and stability at the cost of lower compute consumption; the kill switch and strict guardrails reduced the risk of underprovisioning at peak.
Slide 5 — Postmortem, Plan B, and Collaboration
- What went wrong (concrete mistake): At first we optimized P50 latency, which helped medians but made P95 worse for some bursty workloads. We changed the objective to P95, switched to quantile models, and introduced a tail penalty inside . We also began using CUPED with pre-experiment baselines to make the estimates more stable.
- Metric conflict & resolution: Product cared most about P95 latency, while Finance raised a red flag because pilot credits/query fell 11%. We settled it by rolling out selectively to churn-risk and high-ticket accounts with net-positive NDR, capping savings through a per-account "compute floor" of no more than 10% daily credits reduction, and adding a paid "Performance" entitlement for wider rollout.
- If leadership says no because of revenue concerns (retention up but revenue down):
- Revised follow-up experiment with no timeline reset:
- Keep existing code paths; flip the config to run a three-cell test using current flags:
- Control: no AQA.
- A: AQA unlimited as built.
- B: AQA with a compute floor built in, allowing at most 5–10% credits reduction, and limited to churn-risk accounts.
- For part of cell B, add a pricing/packaging variant using existing entitlements, with no new UI: a "Performance" toggle that requires a higher-commit tier.
- Metrics: primary is P95; guardrails are error rate and queue time; business indicators are credits/account-day and an NDR proxy from expansion signals. Stop-loss rule: if overall credits fall more than 0.5%, pause the expansion.
- Keep existing code paths; flip the config to run a three-cell test using current flags:
- Rationale: It answers the revenue concern through floors and targeting while still proving user value, and it uses feature flags already in place to avoid slipping timelines.
- Revised follow-up experiment with no timeline reset:
- Collaboration under time pressure:
- DE: We needed warehouse-level queue wait time. Since a new pipeline would delay the timeline, we agreed on a minimal schema change—adding
warehouse_idandqueue_wait_msto the existing query log—and computed aggregates downstream. We also aligned on how to handle late-arriving rows so the daily p95 would not be biased. - PM: To meet the quarter, we limited "auto-apply" to warehouse resizing only; query rewrites went out as recommendations. We defined clear success gates for later re-enabling auto-apply on rewrites.
- Design: The UI shrank from a multi-chart dashboard to a simple "Before/After P95 and credits" card with one line of explanation: "We resized during peaks; predicted tail reduction 22%."
- DE: We needed warehouse-level queue wait time. Since a new pipeline would delay the timeline, we agreed on a minimal schema change—adding
How to adapt this pattern to your own project
- Pick a specific feature that actually shipped. Make the north-star metric clear and user-centered; list two or three guardrails with thresholds.
- Explain the experiment unit and why (interference, spillovers). Use a variance-reduction method such as CUPED and a tail-focused metric if the experience is spiky.
- Put numbers on at least one trade-off. Commit to a stop-loss upfront.
- Have a Plan B that can be enabled through flags—segmentation, caps or floors, or packaging—so leadership's concerns can be addressed without slipping.
- Have one concrete mistake ready and describe the exact process fix you implemented.