Tesla · Production Troubleshooting
Investigate Recurring Bugs and Prevent the Same Failure from Returning
TrueInterview
October 7, 2026 · 1 min read
A production bug keeps coming back even after it looks fixed. How would you investigate it and change the engineering process or system so the same failure stops returning?
Constraints
You are not given a specific bug, language, or architecture. Use a real example with accurate ownership details, or clearly label the example as hypothetical. Distinguish repeated instances of the same underlying defect from separate defects that merely produce the same symptom.
Clarifying Questions
- Do the incidents share a signature, triggering conditions, and code path?
- What did each previous fix change, and what evidence justified closing the incident?
- Did the bug recur on the same version, after a rollback, or after a later change?
Hint — Compare the previous fixes: A pattern in what earlier fixes missed can expose the missing invariant, test, or ownership boundary.
What a Strong Answer Covers
- Reproduce and compare the incidents before choosing a permanent fix.
- Provide a causal explanation, targeted regression coverage, and a safe rollout.
- Address ownership, monitoring, and criteria for verifying that the likelihood of recurrence has gone down.
Follow-up Questions
- How would you handle a recurrence that cannot be reproduced locally?
- When is a design change justified instead of another localized patch?
Overview: Investigate recurring production bugs by comparing incidents, identifying missing invariants, adding targeted regression coverage, and verifying permanent fixes.