skip to content

Why does the KV cache, not model weights, cap how many long sessions fit on a GPU?

level: seniorimportance: must knowfreq 56%

answer

  1. what grows per user, not per model
  2. bytes per token times tokens resident
  3. weights are fixed, cache is not
  4. 2 x kv_heads x head_dim x width x layers
  5. budget aggregate resident tokens

basics

~20 s

Model weights are loaded once and shared by every request, so they are a fixed cost. The KV cache is per session and grows with every token, so on long contexts it quickly dwarfs the weights and consumes whatever memory is left.

solid answer

~40 s

Weights are a one-time allocation: an 8B-parameter model in bf16 occupies about 16 GB whether one user or a hundred are connected. The KV cache is charged per resident token per session. For a 32-layer model with 8 KV heads of dim 128 in bf16, one token costs 2 (K and V) x 8 x 128 x 2 bytes = 4 KB per layer, so 128 KB across the stack. A 100K-token session therefore holds about 13 GB — comparable to the entire model — and on an 80 GB card with 16 GB of weights only four or five such sessions fit. Sixty-four concurrent editors at that length would need roughly 840 GB. Capacity planning is therefore a budget over **aggregate resident tokens**, not over user count.

code

python · 7 lines
python
def kv_cache_bytes(tokens, layers=32, kv_heads=8, head_dim=128, width=2):
    per_token_per_layer = 2 * kv_heads * head_dim * width  # one key, one value
    return tokens * layers * per_token_per_layer

print(kv_cache_bytes(1))                    # 131072 bytes = 128 KB per token
print(kv_cache_bytes(100_000) / 1e9)        # ~13.1 GB for one 100K session
print(kv_cache_bytes(100_000) * 64 / 1e9)   # ~839 GB for 64 such sessions

go deeper

for a junior

Know that each active conversation holds its own memory on the GPU and that longer conversations hold more, so a service can run out of memory even though the model itself fits comfortably.

for a middle

Be able to derive bytes per token from layer count, KV heads, head dimension and precision, and multiply it out to a per-session and per-fleet figure without reaching for notes.

for a senior

Show you diagnose from the number: compute the resident-token budget, compare it against free memory after weights, and pick a lever — precision, architecture, retained context or hardware — with its cost stated.

for a principal

Own the capacity model. Set the product's retained-context policy and admission limits against a token budget and a tail-length distribution, and be explicit about which is cheaper for the business: shorter memory, lower precision, or more silicon.

## Two very different memory line items GPU memory during serving splits into a fixed part and a variable part. Weights, and the activation scratch space a forward pass needs, are fixed: they are allocated once and shared by every concurrent request. The KV cache is variable and per-session: every token that a session still holds in context owns a key and a value entry in every layer, and nobody else can use those bytes. Once contexts get long, the variable part dominates, and that is what actually sets concurrency. ## The arithmetic, done once Bytes per token = 2 (one key, one value) x kv_heads x head_dim x bytes_per_value x layers. For a 32-layer model with 8 KV heads of head_dim 128 in bf16: 2 x 8 x 128 x 2 = 4,096 bytes per layer, times 32 layers = 131,072 bytes, or 128 KB per token. Memorize the shape of that formula rather than any one number — you will be asked to redo it live for whatever model is on the table. Scale it up: a 100,000-token session holds 100,000 x 128 KB, about 13 GB. That is roughly the size of the whole 8B model. On an 80 GB accelerator carrying 16 GB of weights and some working space, perhaps 60 GB remains, so four sessions at full length fit, maybe five. Sixty-four editors each holding a 100K-token working set would demand about 840 GB — over ten cards' worth of memory, for a model that fits on one. ## Why this surprises people The instinct from other services is that a model's memory is its size, and concurrency is a CPU or queue question. Here the marginal user brings their own multi-gigabyte allocation, and it grows the longer they stay. Three specific traps follow. First, the footprint scales with resident tokens, not with request rate, so a quiet service with a few very long sessions can be tighter than a busy one with short ones. Second, it grows *during* a request as tokens are generated, so a request that was admitted safely can fail later. Third, it scales with the advertised context limit only if sessions actually use it — provisioning for the maximum window per user is usually a very expensive mistake. ## The right unit of capacity Plan against aggregate resident tokens. Multiply your expected concurrent sessions by their realistic median and tail context length, multiply by bytes per token, and compare that against free memory after weights. Then decide what happens when the budget is exceeded — that is an admission-control question, and the answer has to be a deliberate policy rather than an out-of-memory error at token 40,000. Also watch the tail: a handful of sessions at the maximum length can consume the budget that hundreds of ordinary ones would have shared. ## Levers, and what each one costs There are only four ways to move the equation, and they are worth naming explicitly. **Fewer bytes per entry.** Quantize the cache to a narrower format, which is the cheapest lever in memory terms and carries a quality question you have to measure. **Fewer entries per token.** Choose a model whose attention design stores less per token — more query heads sharing each KV head, a compressed latent per token, sliding-window layers whose cache is capped, or recurrent layers that carry a fixed-size state instead of a per-token entry. **Fewer tokens resident.** Trim or summarize history, cap the context you retain per session, or drop stale content instead of carrying it forever. This is the lever with the most direct product consequence, because it changes what the model can remember. **More memory.** Shard across devices, route long sessions to larger nodes, or offload cold cache to host memory — bearing in mind that decode is bandwidth-bound, so reading offloaded cache back over a slower link directly costs per-token latency. ## The diagnostic When a service starts failing under load, resist blaming the model size. Compute bytes per token from the architecture, multiply by the token counts actually resident at the moment of failure, and compare with free memory. In most long-context incidents the number lands within a few percent of the memory you have, and the fix is a context or precision decision rather than a bigger GPU.

  • A service is fine at 200 short sessions but dies at 12 long ones. What does that tell you?
    That memory is charged per resident token, not per session. Twelve sessions near the context limit can hold more tokens than two hundred short ones, so the cache budget blows through while request rate looks trivial. Capacity must be planned on aggregate token residency, with admission control on that number.
  • Why is offloading the cache to host memory not a free fix?
    Because decode is memory-bandwidth-bound. Every generated token must read the sequence's cache, so entries living across a slower host link add directly to inter-token latency. It can rescue capacity for idle or rarely-touched sessions, but it degrades the sessions that are actively generating.
  • Does a longer advertised context window automatically cost more memory?
    Only if sessions actually fill it. The cache is proportional to tokens resident, not to the declared limit — unless the engine preallocates for the maximum, which is exactly why preallocation is a poor default. The real risk of a large window is that it invites usage patterns that do consume it.

saying these in an interview costs you the question

  • Assumes memory need equals the model's parameter size
  • Plans capacity by concurrent users rather than resident tokens
  • Thinks the cache is allocated once up front and stays flat
  • Believes a larger context window is free until it is used
  • Reaches for a bigger GPU before computing bytes per token

context