Back to problems

Route and Batch Inference with Eight GPUs

System Design · Anthropic · Hard

Since the original design document is not supplied, I’ll evaluate the design space directly. The first step is to clarify four context questions: Model type: Are these autoregressive generative models such as LLMs, or non-autoregressive encoders/classifiers? This determines whether batching is stateful and whether output length is variable. Traffic mix: What fraction is online interactive traffic versus offline batch traffic? Online traffic needs strict tail-latency control;…

Checking your access…