Meta · Production Troubleshooting
Diagnose a Single-Server Crash with Concrete Linux Evidence
TrueInterview
October 7, 2026 · 1 min read
One server has gone down. Describe how you would figure out what broke, bring the service back without causing more harm, and find the root cause. For each step, name the specific commands or evidence you would rely on and explain how the outcome shapes your next move.
Constraints
For this scenario, suppose you have a Linux host running systemd, a management console you can reach, and authorization to look at logs and metrics. These assumptions exist only for diagnosis; they do not imply any particular cause, application, redundancy setup, or permission to erase evidence. Separate the cases of an application that crashed, a host that cannot be reached, and an OS that has rebooted or stopped.
Clarifying Questions
- Can you not reach the host at all, or is a single service down while the host itself still answers?
- When did the outage begin, what is the user-facing impact, and was there a deployment or infrastructure change just before it?
- Is another replica healthy, and are durable data or in-flight writes at risk?
Hint — Branch on what you observe: Failing to connect over SSH does not by itself prove the kernel crashed. Treat network reachability, host health, and application health as separate checks.
What a Strong Answer Covers
- Assess the impact, mitigate safely, and preserve volatile evidence where possible.
- Investigate host, service, and resource state in a step-by-step way using specific Linux commands.
- Connect the failure to recent changes, verify the fix after recovery, and tie prevention to the cause you actually proved.
Follow-up Questions
- What evidence would tell you whether the process was killed by an out-of-memory kill rather than throwing an application exception?
- If the server cannot boot far enough to accept SSH, what can you examine?
Overview: Walk through an investigation of a single-server crash using Linux service, kernel, memory, disk, and boot evidence while planning a safe recovery.