skip to content

Beyond weights, what else consumes GPU memory when serving an LLM, and when does it dominate?

level: seniorimportance: should knowfreq 55%

answer

  1. Weights are the floor, not the total
  2. Second term scales with tokens in flight
  3. Length times concurrency, not either alone
  4. Prefill spikes where decode is calm
  5. Halving the wrong term buys little

basics

~20 s

Three consumers share the card: weights, which are fixed once loaded; the KV cache, which grows with tokens in flight across all live requests; and transient activations plus runtime overhead. At long context and high concurrency the cache routinely exceeds the weights.

solid answer

~60 s

Weights are a constant resident cost — parameter count times bytes per parameter, decided at load. The KV cache is the variable term: it grows linearly with the number of tokens being attended to, meaning sequence length multiplied by concurrent requests, so a deployment that is comfortable at 4K context and batch 2 can OOM at 128K and batch 16 with identical weights. Activations are the third term, small during token-by-token decoding but spiking during prefill of a long prompt, alongside allocator fragmentation and framework workspace. The practical planning move is to treat weights plus overhead as the floor and everything left as a budget you spend on concurrency times context — those two multiply, so you cannot have both. The important consequence for quantization is that shrinking the weights only enlarges the budget; if the cache is what is filling the card, a smaller weight format buys you headroom but does not change the shape of the problem, and you should be quantizing or compressing the cache too.

code

python · 7 lines
python
params, bytes_per_param = 70e9, 4.25 / 8      # 70B at a 4-bit block-scaled width
weights = params * bytes_per_param
kv_bytes_per_token = 100e3                    # measured for this model and cache precision

budget = 80e9 - weights - 4e9                 # 80 GB card, minus weights and overhead
print(round(weights / 1e9, 1), "GB weights")
print(int(budget / kv_bytes_per_token), "cached tokens affordable")

go deeper

for a junior

Know the three consumers by name — weights, KV cache, activations — and that weights are only the starting floor of a deployment's memory, not the whole requirement.

for a middle

Explain that the cache grows with sequence length times concurrency while weights stay fixed, and be able to turn leftover VRAM into a rough token budget.

for a senior

Diagnose from the regime: say which term dominates for a given workload, name the matching lever, and describe how you would measure the split rather than estimate it.

for a principal

Own the capacity plan: set admission limits and context caps as product decisions, decide where quantization spend goes across weights and cache, and be explicit about the concurrency-versus-context tradeoff the hardware forces.

## Three consumers, three different behaviours Ask anyone to size an LLM deployment and they will compute weight memory. That is necessary and insufficient. A serving GPU holds: **1. Weights — fixed, resident, known at load time.** Parameter count times bytes per parameter, adjusted for whatever scales a quantized format stores. Once the process is up this number does not move. It is the floor. **2. KV cache — variable, proportional to tokens in flight.** Every token that has been processed and is still part of a live conversation contributes cached key and value tensors so that later tokens do not recompute attention over the whole history. The total is (bytes per token) x (tokens across all active sequences). The per-token size depends on the model's layer count and attention design and is a topic of its own; what matters for budgeting is the shape of the growth: **linear in sequence length and linear in concurrency, so quadratic in the product you are trying to increase.** **3. Activations, workspace and overhead — transient but real.** During decoding, only one token per sequence is in flight, so activation memory is modest. During prefill of a long prompt, the model processes thousands of tokens at once and the intermediate tensors and attention workspace balloon. On top of that sit the driver/runtime context, communication buffers when the model is sharded across GPUs, and allocator fragmentation — memory that is free but not contiguous enough to use. ## The regime question Which term dominates depends entirely on the workload: - **Short prompts, low concurrency** (a batch job classifying single sentences): weights dominate overwhelmingly, and the cache is noise. Quantizing weights is the whole game. - **Long documents, high concurrency** (a chat product with long histories, or document analysis): the cache dominates. It is entirely normal for cache to exceed weight memory several times over, and at that point a cheaper weight format buys headroom but does not address the bottleneck. - **Long prefills with bursty arrivals**: the activation spike becomes the failure mode. The server sits comfortably for hours and then OOMs when two long prompts arrive together. The last case is the one that catches people, because average utilisation looks fine and the failure is a tail event. ## A concrete budget Take a 70B model on a 4-bit format on a single 80 GB accelerator: - Weights at ~4.25 effective bits: ~37 GB. - Runtime, context and communication buffers: call it 3-5 GB. - That leaves roughly 38-40 GB for cache and activations. If the measured cache cost is on the order of 100 KB per token for this model and cache precision, 38 GB buys about 380K cached tokens. That is one request at 380K context, or twelve concurrent 32K sessions, or a hundred concurrent 4K chats — not all three. Stating the budget that way, rather than as "it fits", is what an interviewer is listening for. ## The levers When the budget does not close, the options are: - **Reduce weight bytes** — a lower width frees space one-for-one, and is the right lever when weights dominate. - **Reduce cache bytes per token** — store the cache at a lower precision, or choose a model whose attention design caches less. This is the right lever when the cache dominates, and it is the reason low-precision cache storage became routine. - **Cap the variable term** — limit maximum context, cap concurrent sequences, or evict and recompute inactive sessions rather than holding their cache resident. - **Bound the prefill spike** — process long prompts in bounded chunks rather than all at once, trading a little latency for a much lower peak. - **Add hardware** — the honest answer when the arithmetic simply does not close. ## Why quantization people must think this way It is easy to present quantization as "the model got smaller, problem solved". The budget view corrects that. Halving weight bytes on a cache-dominated deployment might raise achievable concurrency by a modest fraction, whereas halving cache bytes per token might double it. Knowing which term dominates *before* choosing a format is the difference between an engineering decision and a habit. ## Measuring rather than guessing Every credible answer ends here. Load the model, observe resident memory with zero traffic to confirm the floor, then ramp real traffic and watch peak allocation against context length and concurrency. Serving runtimes typically pre-reserve the cache region up front, which makes the split visible directly. Reserve a safety margin — fragmentation means the last few percent of a card is not usable — and set an admission limit that refuses work rather than crashing the process.

  • A service runs fine for hours, then OOMs. Weight memory has not changed. What do you suspect first?
    The variable terms. Either concurrent long sessions pushed cached tokens past the budget, or two long prompts arrived together and their prefill activation spikes coincided. Both are tail events invisible in average utilisation. Check peak allocation against concurrency and context length, then cap admitted concurrency, cap maximum context, or bound the prefill chunk size — and leave headroom for fragmentation.
  • If the cache dominates your budget, is moving from 8-bit to 4-bit weights worth doing?
    It helps, but it is the smaller lever. Freeing weight bytes enlarges the cache budget one-for-one, so the gain in achievable concurrency is proportional to the fraction of the card the weights held. If weights were 40% of the card and the cache is the binding constraint, halving them raises the budget by a third — real, but less than halving cache bytes per token would give. Measure which term binds before choosing.
  • Why does peak memory spike during prefill but not during decoding?
    Because prefill processes the whole prompt at once, so intermediate activations and attention workspace scale with prompt length times batch. Decoding advances one token per sequence, so those same buffers are tiny. The fix is to process long prompts in bounded chunks, which caps the peak at the chunk size and costs a little latency rather than a crashed process.

Weights are the furniture already in the room; the cache is the guests, and every guest brings luggage proportional to how long they stay. Buying smaller furniture helps only if furniture was the crowding problem.

saying these in an interview costs you the question

  • Sizing a deployment from weight memory alone
  • Treating the KV cache as small or negligible
  • Assuming smaller weights automatically raise concurrency proportionally
  • Ignoring the prefill activation spike because decode looks calm
  • Planning to fill the card with no fragmentation headroom

context