System Design · Microsoft · Medium
Incident Response: Post-Deployment Latency/Reliability Regression 1. Detection – Immediate Signals I would first confirm the alert and scope the problem using monitoring dashboards and anomaly detection: Latency distributions: Compare p50/p95/p99 latencies for the affected endpoint(s) over the last 15 minutes against a pre-deployment baseline. A sudden shift in the tail (p99) often indicates contention or slow paths specific to the new code. Error rate and saturation: Track…
Checking your access…