Back to problems

Debug distributed-system performance problems

System Design · Microsoft · Hard

I’d approach it as a structured performance investigation, moving from broad observability signals to a specific root cause, then validating the fix safely and preventing recurrence. 1. Start with high-level signals and metrics First, define the regression clearly: Which service, endpoint, or user segment is affected? When did the regression start? What changed around that time? How far are current latency/throughput from the baseline or SLO? Then check the most informative…

Checking your access…