DoorDash · Behavioral Stories
Own Production Incidents and Validate Large Pipeline Migrations
TrueInterview
October 7, 2026 · 2 min read
Overview: Lay out an approach for validating a migration at scale, catching silent failures in which a migrated job stops producing anything, and settling who is accountable for incident recovery and ongoing monitoring when no dedicated QA function exists.
Part 1 — Take ownership of a production incident. Describe an actual production incident you helped cause; if you have never caused one, say plainly that your example is hypothetical. Explain what went wrong, how large the impact was, how you reacted, and what you changed afterwards. Cover your own part in the failure and its consequences without pushing blame onto users or neighbouring teams; the immediate mitigation, the communication you sent, and the evidence that showed recovery; and the corrective work that removes both the failure mechanism and the detection gap that allowed it to reach production.
Part 2 — Validate a large migration. Explain how you would build confidence in a large population of migrated jobs when reviewing each one by hand is impractical, including the situation where a job's output or activity falls to nothing after the cutover. Cover automated checks over configuration, execution and output behaviour across the whole migrated population; comparisons or invariants that expose silent breakage, together with a small targeted manual sample; health alerts that understand each job's workload and separate genuine idleness from missing expected output; and who receives alerts, which gates the rollout must clear, and who owns health after handoff given that there is no QA team.
Clarifying questions you should ask
- For a given job, what does it normally emit, how frequently, and how do you observe that a run succeeded?
- Is it safe to run the old and new versions side by side, or do side effects force a different approach?
- Which jobs are permitted to be idle, and which data tells you their expected schedule or whether their input has arrived?
- During the migration, who gets paged — and who owns the job once it has been handed over?
Example 1:
Input: 500 batch jobs migrated to a new scheduler; one job that previously wrote a daily file now writes nothing, yet its run reports success.
Output: Name the check that would have caught this before a downstream consumer did, and say who should be paged.
Example 2:
Input: A job with a seasonal schedule is legitimately idle over a holiday period while other jobs continue producing.
Output: Explain how your alerting distinguishes this expected gap from a real failure.
Example 3:
Input: An upstream feed never arrives, so a job emits zero records and that is the correct outcome.
Output: Describe the rule or signal that keeps this case from being reported as a fault.
Follow-up questions
- Once automated checks pass, how do you decide which migrated jobs still deserve a human look?
- What should happen when there is no input at all and producing zero output is genuinely correct?
- Why should the person leading the migration be paged directly instead of waiting for business users to discover the problem?
Constraints:
- The migrated population is far larger than can be inspected one job at a time, and no separate QA function is available.
- You may assume exit status, emitted artifacts, and configured schedules are all observable.
- Both parts must be answered; label your incident as hypothetical if you have not experienced one.
- The response should combine personal accountability with scalable validation and continuing health checks.