Back to problems

Trace an LLM Request Through a Paged-KV Inference Engine

System Design · Amd · Hard

We’ll walk through the entire lifetime of a request in a high‑throughput LLM serving system built on paged attention (the seminal idea from vLLM). The model is a standard autoregressive transformer decoder (e.g., 13B parameters, 40 layers, hidden size 5120, 40 attention heads, max sequence length 8192). We assume an 80 GB GPU memory budget, typical prompt sizes of 100–2000 tokens, and generated outputs of up to 200 tokens. Latency targets: time‑to‑first‑token (TTFT) under…

Checking your access…