Amd · Behavioral
Use the Roofline Model to Analyze Prefill and Decoding
TrueInterview
September 26, 2026 · 1 min read
Apply the roofline model to reason through an LLM inference workload. How would you decide if a kernel is memory-bound or compute-bound, and in what ways do prefill and decoding differ?
Constraints & Assumptions
State the memory tier and the operation-counting convention you use for arithmetic intensity. Treat prefill and decoding bottlenecks as dependent on the workload, not as fixed labels.
Clarifying Questions
What are the batch size, sequence length, precision, hardware compute ceiling, and sustainable memory bandwidth? Are the measurements taken at the kernel level or end-to-end? Which bytes are actually moved?
What a Strong Answer Covers
The roofline equation, arithmetic intensity, measured utilization, data reuse, and how batching and KV-cache reads influence the two inference phases.
Follow-up Questions
Can a kernel sit well below both ceilings? How do kernel launch overhead, synchronization, and small tensor shapes affect the diagnosis? What shifts when decode batch size or context length increases?
Overview: Use roofline analysis to examine LLM prefill and decoding through arithmetic intensity, sustainable bandwidth, batching, KV-cache traffic, and measured kernel behavior.
Loading comments…