Back to problems

Debug and mitigate a CPU spike incident

System Design · Discord · Medium

You're on call for a backend service when an alert reports a sharp, sustained increase in CPU utilization across its instances or pods. Your answer should walk through four areas: the immediate steps you take to protect users, the questions and observability data (metrics, logs, traces, recent deploys and configuration changes) you inspect, how you narrow possible causes to a likely root cause, and how you validate the fix and guard against recurrence. Assume a typical…

Checking your access…