Why is single-stream LLM decoding limited by memory bandwidth rather than FLOPs?
answer
- Bytes moved, not maths done
- Roughly two operations per byte read
- Bandwidth divided by bytes per token
- A faster arithmetic unit changes nothing
- Batching amortises one weight read
basics
~20 sGenerating one token for one request reads every parameter that token's forward pass needs out of GPU memory but does only a couple of arithmetic operations per parameter. The accelerator therefore sits waiting on memory, so tokens per second track memory bandwidth, not peak FLOPs.
solid answer
~40 sAt batch size one, each decode step multiplies the model's weight matrices by a single vector. That means reading a huge number of bytes — every parameter involved in that step, plus the cached attention state — to perform roughly two arithmetic operations per parameter. Arithmetic intensity is therefore about one operation per byte, while modern accelerators need hundreds of operations per byte to keep their arithmetic units busy. On a roofline plot the workload sits far out on the memory-bound side, so the practical ceiling is `memory bandwidth / bytes read per token`: roughly 40 tokens per second for a 70B model at 8-bit weights on a 3 TB/s device. Doubling FLOPs moves that ceiling not at all. The lever is batching, which reuses one weight read across many sequences.
code
python · 12 linesbandwidth_bytes_per_s = 3.0e12 # ~3 TB/s of device memory bandwidth
active_params = 70e9 # parameters read per generated token
bytes_per_param = 1.0 # 8-bit weights
bytes_per_token = active_params * bytes_per_param
single_stream = bandwidth_bytes_per_s / bytes_per_token
print(round(single_stream, 1)) # ~42.9 tokens/s ceiling at batch 1
batch = 32 # one weight read now serves 32 sequences
print(round(batch * single_stream, 1)) # ~1371 tokens/s aggregate upper bound
# Upper bounds only: real systems also read cached attention state,
# never hit peak bandwidth, and become compute-bound at large batch.go deeper
Be able to say that generating a token requires reading a very large amount of data out of GPU memory, so memory speed and model size determine how fast text streams.
Explain arithmetic intensity concretely: a matrix-vector product does about two operations per byte read, far below what the hardware needs, and state the bandwidth-over-bytes-per-token estimate out loud.
Show you can diagnose it in production — high GPU utilization at a low fraction of peak FLOPs, token rate tracking byte footprint — and that you know batching raises throughput without raising per-user speed.
Own the procurement and architecture consequence: interactive generation is bought with memory bandwidth and with reductions in bytes read per token, which is why active-parameter count drives serving economics more than total size.
## What a single decode step actually does Autoregressive generation produces one token at a time. For one request in isolation, that step takes a single vector — the representation of the last token — and pushes it through every layer of the model. Each layer's weight matrix is multiplied by that one vector. The result is a matrix-vector product, and matrix-vector products are the worst possible shape for a modern accelerator: the hardware must read the entire matrix out of high-bandwidth memory and then perform only about two floating-point operations (a multiply and an add) per element it read. The ratio of arithmetic performed to bytes moved is called **arithmetic intensity**. Here it is roughly one operation per byte. A contemporary datacenter GPU has an arithmetic-to-bandwidth ratio in the hundreds — it can perform hundreds of operations in the time it takes to fetch one byte. The consequence is unavoidable: the arithmetic units are idle almost all the time, waiting on memory. The workload is **memory-bandwidth-bound**. ## The roofline estimate Because the step is bandwidth-bound, you can estimate the token rate with arithmetic you can do in an interview: tokens per second (ceiling) = memory bandwidth / bytes read per token bytes read per token ~= active parameters x bytes per parameter A 70-billion-parameter dense model with 8-bit weights reads roughly 70 GB per token. On a device with about 3 TB/s of memory bandwidth that is a ceiling near 43 tokens per second, and real systems land below it because the cached attention state must be read too and no kernel achieves 100% of peak bandwidth. The important property of this formula is what is missing from it: peak FLOPs does not appear. A device with twice the arithmetic throughput and the same memory system generates tokens at the same speed. This is also why the number people quote for a model on a laptop is so predictable. Bandwidth per dollar, not FLOPs per dollar, is what sets interactive generation speed. ## Why batching is the lever that works Run B requests together and the step becomes a matrix-**matrix** product: the same weight matrix now multiplies B vectors. The bytes read are essentially unchanged — you read each weight once — while the arithmetic performed is B times larger. Arithmetic intensity rises roughly B-fold, and aggregate throughput rises with it, almost for free, until intensity crosses the roofline knee and the workload becomes genuinely compute-bound. Past that point extra batching buys little throughput and costs latency. Two caveats keep this from being unlimited. First, the per-request token rate does not improve; batching raises tokens per accelerator-hour, not tokens per user, and beyond the knee it makes each user slower. Second, every concurrent sequence needs its own cached attention state in the same memory the weights live in, so available memory caps how many sequences can share a step long before arithmetic does. ## What changes the byte count Only the parameters a token's forward pass actually touches are read. That distinction became central once mixture-of-experts architectures became the norm for frontier models by the mid-2020s: such a model routes each token to a small subset of its expert blocks, so bytes read per token reflect the **active** parameter count, which can be an order of magnitude below the total. That is precisely why very large mixture-of-experts models can decode faster than a much smaller dense model. The full parameter set still has to be resident in memory, though, and at large batch sizes different tokens in the batch route to different experts, so the union of experts read per step grows and some of the advantage erodes. The other term is bytes per parameter, which is simply the numeric format the weights are stored in. Fewer bytes per parameter means fewer bytes read per token and a proportionally higher bandwidth ceiling; the accuracy consequences of that choice are a separate subject. ## What does not help A faster arithmetic unit does not help a memory-bound step. Neither does a larger accelerator with the same memory system, nor running the same single request across more devices unless that split also multiplies aggregate bandwidth. And the phase of the request matters: processing a long prompt is a different shape entirely, because thousands of prompt positions go through the weights together and one weight read serves all of them — that phase is compute-bound and behaves nothing like this one. ## How to recognise it in production The signature is a GPU reporting high utilization while achieving a small fraction of its advertised FLOPs, with token rate that tracks the model's byte footprint almost linearly and refuses to improve when you move to a compute-heavier device. If token rate scales when you add concurrency but not when you add arithmetic, you are on the bandwidth roof.
- Does a GPU with twice the FLOPs but identical memory bandwidth speed up batch-one generation?Essentially not at all. The step is limited by how fast weights can be pulled out of memory, and the extra arithmetic units simply idle for longer. You would see the same tokens per second and a lower fraction of peak FLOPs achieved. The device that helps is the one with more bandwidth, or the change that reduces bytes read per token.
- How does a mixture-of-experts architecture change this arithmetic?Only the routed-active parameters are read per token, so bytes per token can be far below the total parameter count and decoding is correspondingly faster than the model's size suggests. The trade is that all experts must still be resident in memory, and at larger batch sizes different tokens select different experts, so the union of expert weights read per step grows.
- What happens to the ceiling as you keep increasing batch size?Arithmetic intensity rises roughly linearly with batch size until it crosses the roofline knee, after which the step is compute-bound and further batching adds latency without adding throughput. In practice memory for per-sequence cached attention state usually caps batch size before the compute roof does.
At batch one the accelerator is a truck driving the entire warehouse across town to deliver a single parcel. Batching loads many parcels onto the same trip; buying a faster engine does not help when the trip itself is the cost.
saying these in an interview costs you the question
- Says a GPU with more FLOPs will fix slow token generation
- Thinks model weights stay resident in on-chip cache between tokens
- Believes token rate depends on parameter count with no bandwidth term
- Assumes batching speeds up an individual user's stream
- Confuses aggregate fleet throughput with per-user tokens per second