Back to problems

Review and Scale an Inference Service Design

System Design · Anthropic · Hard

An inference service is proposed with this request path: a load balancer forwards traffic to a dynamic batching layer, which then sends batches to model servers. The platform must handle a sustained load of 10,000 requests per second, and every batch must complete in at most 100 ms. Assess this architecture for concrete flaws. Then explain: how batching should be designed when request shapes and model execution times differ; how to estimate the required number of model…

Checking your access…