skip to content

Why is LLM prefill compute-bound while token-by-token decode is bandwidth-bound?

level: middleimportance: must knowfreq 70%

answer

  1. two phases, two different limits
  2. many tokens per weight read, or one
  3. FLOPs performed per byte read
  4. matrix-matrix versus matrix-vector
  5. math units idle while memory streams

basics

~20 s

Prefill pushes every prompt token through the model in one pass, so each weight read is amortized over many tokens and the math units saturate. Decode produces one token per pass, re-reading all weights and the cache for that single token, so memory bandwidth is the ceiling.

solid answer

~50 s

Inference splits into two phases with different hardware bottlenecks. **Prefill** processes the whole prompt at once: every weight matrix is loaded once and multiplied against a tall matrix of hundreds or thousands of token vectors, so arithmetic intensity — FLOPs performed per byte read — is high and the accelerator's tensor cores are the limit. **Decode** generates one token per forward pass: the same weights are read again, but each element is used for a couple of FLOPs against a single vector, and the growing KV cache must also be streamed in. Arithmetic intensity collapses to roughly one to two FLOPs per byte, far below what a modern GPU needs to stay busy, so the math units idle while memory delivers bytes. The practical consequence is that the two phases respond to different fixes: prefill wants FLOPs and cheaper attention, decode wants fewer bytes moved.

go deeper

for a junior

Know that the prompt is processed in one pass before generation starts, and that tokens then come out one at a time — and that this is why the wait before the first word differs from the speed of the words after it.

for a middle

Explain arithmetic intensity: prefill is a matrix-matrix multiply that amortizes each weight read over many tokens, decode is a matrix-vector multiply that reads the same weights for one token. Name which hardware resource saturates in each.

for a senior

Demonstrate that you tune the phases separately: read utilization counters to confirm the regime, and pick remedies that match — attention cost and prefix reuse for prefill, bytes-per-token reduction for decode.

for a principal

Own the capacity argument. Characterize the workload's prompt-to-output ratio, then justify hardware and topology choices — bandwidth versus FLOPs, co-located versus separated phases — against that mix rather than against a single throughput number.

## The two phases A request is served in two structurally different phases. **Prefill** runs the entire prompt through the model once, producing the KV cache for every prompt position and the first output token. **Decode** then runs one forward pass per generated token, each pass appending a single entry to the cache. Both phases execute exactly the same layers and the same math; what differs is the shape of the operands, and that shape decides which piece of hardware runs out first. ## Arithmetic intensity is the whole story Roofline analysis asks a single question of a kernel: how many FLOPs does it perform per byte it reads from memory? An accelerator has a peak math rate and a peak memory bandwidth; their ratio — typically in the hundreds of FLOPs per byte on current data-center parts — is the break-even intensity. Below it you are bandwidth-bound, above it compute-bound. In prefill, a projection is a matrix-matrix multiply: a weight matrix of d x d is read once and multiplied against P token vectors. The bytes read stay proportional to d^2 while the FLOPs scale with P x d^2, so intensity rises with prompt length. A few hundred tokens is already enough to sit comfortably on the compute side of the roofline. In decode, P is 1. The same weight matrix is read in full and each element participates in about two FLOPs. Intensity is roughly one to two FLOPs per byte — one to two orders of magnitude below break-even. The tensor cores spend most of their time waiting for weights to arrive. On top of that, attention in decode must read the entire KV cache for the sequence, and that read grows with context length, so long sessions move even more bytes per generated token. ## What each phase's cost scales with Prefill cost scales with prompt length for the projections and feed-forward blocks, and with the square of prompt length for full attention scores — which is why long-prompt systems reach for sparse or linear attention variants and for reusing previously computed prefixes. Decode cost per step is roughly constant in weight FLOPs but grows in bytes read as the cache lengthens, so inter-token latency drifts upward across a long session even though the model is doing the same amount of arithmetic each step. ## Why this matters for tuning Because the bottlenecks differ, so do the remedies, and applying the wrong one wastes money. A prefill-heavy service — long documents, short answers — benefits from raw FLOPs, better attention kernels, cheaper attention math and avoiding re-encoding of content the system has already processed. A decode-heavy service — short prompts, long generations — barely notices extra FLOPs; it improves when you reduce bytes moved per token: quantizing weights, quantizing or shrinking the KV cache, choosing architectures with smaller per-token cache entries, or having more sequences share each weight sweep so a single weight read serves many tokens at once. ## A concrete split Consider an IDE inline-completion service. Each keystroke sends a 6,000-token file prefix and the model returns a 20-token completion. The prefill is one parallel pass over 6,000 tokens — a large slab of arithmetic that the hardware is well-shaped to do. The decode is 20 sequential passes, each re-reading the full weight set and the 6,020-entry cache to emit one token. The prefill does far more total arithmetic; the decode does far more sequential trips to memory. Which one dominates the wall clock depends on model size, prompt length and hardware, so measure rather than assume — but they will never respond to the same optimization. ## Telling which regime you are in Instrument, do not guess. During prefill you should see high tensor-core utilization and FLOPs near a healthy fraction of peak. During decode you should see memory throughput near peak with math utilization low. If decode shows neither, you are limited by something else — kernel launch overhead, small batches, host-side scheduling or synchronization — and adding hardware will not help. ## The serving consequence Because the two phases want different things, production stacks increasingly run them separately: a long prefill sharing a device with active decodes stalls those decodes and spikes their inter-token latency, so engines either chop prefill into chunks that interleave with decode steps, or place the phases on different pools of hardware and transfer the cache between them. The lesson at the method level is the same either way: prefill and decode are two workloads wearing one model, and treating them as one is how capacity plans go wrong.

  • How does a decode step's cost change as the session gets longer?
    Weight FLOPs per step stay flat, but attention must read the whole KV cache, so bytes moved grow linearly with context. At long contexts the cache read can rival or exceed the weight read, and per-token latency creeps upward even though the arithmetic per step is unchanged.
  • Would a GPU with twice the FLOPs and identical bandwidth speed up generation?
    Prefill would improve substantially, since it is compute-bound. Decode would barely move, because it is limited by how fast weights and cache stream out of memory, not by math throughput. For a generation-heavy workload, the money is better spent on bandwidth, capacity, or on moving fewer bytes.
  • Why do serving stacks sometimes run prefill and decode on separate hardware?
    They want opposite things. A long prefill monopolizes compute and stalls in-flight decode steps, spiking inter-token latency for everyone on that device. Splitting them lets each pool be sized and scheduled for its own bottleneck, at the cost of transferring the cache between pools.

Prefill is reading a whole chapter in one sitting; decode is writing one word at a time and flipping through the entire dictionary before each word.

saying these in an interview costs you the question

  • Says decode is slow because attention is quadratic per generated token
  • Claims more FLOPs will fix per-token decode latency
  • Thinks prefill and decode execute different math or different layers
  • Assumes the prompt is re-encoded on every generated token
  • States decode cost per token is independent of context length

context