System Design · Waymo · Medium
Requirements Inference serving architecture and system flow for 100M daily users Create an inference-serving architecture for a model used by roughly 100 million people each day. Build a first-principles capacity estimate: calculate the model's memory use, bytes that must be read or written for one inference, and expected latency on the selected accelerator. Focus on the two failure cases emphasized in this exercise: excessive latency, especially at the tail during traffic…
Checking your access…