System Design · Microsoft · Hard
I’d approach it as a structured performance investigation, moving from broad observability signals to a specific root cause, then validating the fix safely and preventing recurrence. 1. Start with high-level signals and metrics First, define the regression clearly: Which service, endpoint, or user segment is affected? When did the regression start? What changed around that time? How far are current latency/throughput from the baseline or SLO? Then check the most informative…
Checking your access…