System Design · Oracle · Hard
Design a centralized GPU health monitoring and remediation service for a fleet of 100,000 servers. The service must distinguish between a reachable host and a host whose GPUs are actually usable, and it must trigger automated repair actions from accelerator-level evidence rather than network reachability. Describe how GPU health observations are produced, accepted, and ingested. Explain why a successful host-level ping or SSH login is insufficient to conclude that its GPUs…
Checking your access…