Back to problems

Design a Server Metrics Monitor

System Design · Palantir · Medium

Design a centralized pull-based monitoring system that periodically collects metrics from a fleet of servers and makes them available for dashboards and alerting. The system must reach each of 1,000 servers every 10 minutes, pull a small set of gauges (CPU usage, memory usage, disk usage, and application health), and persist the readings so they can be queried and evaluated by alert rules. The core of this design is the worker that actually performs the collection jobs; you…

Checking your access…