How much GPU memory does a long shared KV prefix pin, and what does it displace?
answer
- tokens times layers times head width
- two tensors, sixteen bits each value
- hundreds of kilobytes per token
- the budget is shared with concurrency
basics
~20 sRoughly 2 x layers x key/value heads x head width x tokens x bytes per value. For a 70B-class model that is about 320 KB per token — near 9 GB for a 30,000-token prefix — and it comes out of the same pool serving in-flight requests, so retention trades hit rate against concurrency.
solid answer
~60 sThe arithmetic is simple and worth doing out loud: two tensors (K and V) per layer, each about `kv_heads x head_dim` values per token. For 80 layers, 8 key/value heads of width 128, in 16-bit, that is ~320 KB per token, so a 30,000-token shared preamble pins roughly 9 GB and a 200,000-token one pins over 60 GB. On a self-hosted server that memory sits beside the model weights in the same budget that holds every running request's own KV, so every gigabyte retained is a gigabyte unavailable for concurrency — fewer simultaneous sequences, a smaller batch, lower throughput. There is no universally right setting: the judgement is whether the prefill you avoid is worth the batch size you give up, which depends on how many distinct hot prefixes you have, how often each is hit, and how long your outputs are. The levers are grouped-query or latent attention to shrink bytes per token, eviction policy, and spilling cold entries to host memory or NVMe — where reload transfer time replaces prefill compute, a trade that only pays above some prefix length.
code
python · 7 linesdef kv_bytes(tokens, layers=80, kv_heads=8, head_dim=128, bytes_per_value=2):
# two tensors per layer (keys and values), per token
return 2 * layers * kv_heads * head_dim * tokens * bytes_per_value
print("per token:", kv_bytes(1) / 1024, "KiB")
for n in (1_000, 30_000, 200_000):
print(f"{n:>7} tokens: {kv_bytes(n) / 1e9:.1f} GB")go deeper
Know that a cached prompt is real memory on the accelerator, sized by how many tokens it covers, not a small bookkeeping entry — and that memory is finite.
Be able to produce the formula and a number: two tensors per layer, key/value heads times head width per token, which lands in the hundreds of kilobytes per token for a large model.
Show the operating consequence — retained prefix memory competes with in-flight requests' KV, so it lowers achievable batch size and throughput — and know the levers: attention architecture, cache precision, eviction weighting, tiering to host memory.
Own the capacity argument end to end: value scales as tokens saved times hit count against memory held, so the decision hinges on prefix diversity and workload shape, and cache partitioning across tenants is a policy judgement you should make explicitly rather than inherit from a default.
## The arithmetic Per token, per layer, the cache holds one key vector and one value vector, each of about `num_kv_heads x head_dim` elements. So: `bytes = 2 x layers x kv_heads x head_dim x tokens x bytes_per_value` With a 70B-class configuration — 80 layers, 8 key/value heads, head width 128, 16-bit values — that is 327,680 bytes per token, roughly 320 KB. Multiply through: - 1,000 tokens: ~0.33 GB - 30,000 tokens: ~9.8 GB - 200,000 tokens: ~65 GB Those are not rounding errors on an 80 GB accelerator. A single very long shared preamble can consume a meaningful fraction of a device. ## What it displaces On a self-hosted server, device memory is partitioned between model weights (fixed), activation workspace (small), and KV space (everything left). Every in-flight request occupies KV space proportional to its own context length, and the number of requests a server can run concurrently is set by how many fit. A retained shared prefix is drawn from that same pool. So the tradeoff is not "cache versus nothing" — it is **hit rate versus batch size**. Retain aggressively and you skip prefill but run fewer concurrent sequences, which on a decode-heavy workload directly reduces throughput, because decode is bandwidth-bound and batching is what amortises the weight reads. Retain conservatively and you rebuild prefixes more often, raising time-to-first-token. Note also that a shared prefix is stored **once** and read by every request that matches it. This is the property that makes retention attractive: one 9 GB preamble serving two hundred concurrent sessions is dramatically better than nothing, whereas two hundred distinct 9 GB prefixes is simply impossible. The design question is therefore about *prefix diversity* as much as prefix length. ## Levers that change the numbers - **Attention architecture.** Grouped-query attention shares a small number of key/value heads across many query heads; going from 64 query heads to 8 key/value heads cuts bytes per token roughly eightfold. Architectures that compress the cached state into a lower-dimensional latent representation cut it further. This is a model-selection decision with a direct capacity consequence. - **Precision.** Storing cached keys and values at lower precision than the compute path halves or quarters the footprint, at some accuracy risk that has to be measured on your own evals rather than assumed. - **Eviction policy.** LRU is the default. Weighting by build cost or hit frequency favours long, hot prefixes over short, cold ones, which is usually what you want, since the value of an entry scales with tokens saved times hits. - **Tiering.** Spilling cold entries to host RAM or NVMe converts a memory problem into a transfer problem. Reloading pays interconnect time instead of prefill FLOPs, so it wins above a crossover prefix length and loses below it — measure the crossover on your hardware rather than adopting someone else's number. ## Multi-tenancy Sharing one prefix across tenants is functionally sound: the state is derived only from tokens they have in common, and anything user-specific enters the sequence after the divergence point, so no tenant's content leaks into another's attention state. The genuine concern is observational — whether a request hits or misses is visible in its latency, which in principle lets one party probe whether a particular prefix has been seen. Where prompts themselves are sensitive, partitioning the cache per tenant is the conservative answer, and it costs you exactly the deduplication you were hoping for. That is a policy call, not a tuning knob. ## What is not in your hands On a hosted API none of this is yours to configure — the provider sizes and evicts. What transfers across both settings is the reasoning: the cache is a real memory object measured in hundreds of KB per token, its value scales with tokens saved times hits, and it is always competing with something. Candidates who can produce the per-token figure from the model's shape, and then say what that memory is being taken from, are demonstrating capacity thinking rather than feature recall. ## How to decide There is no consensus formula and no default worth memorising. Characterise the traffic first: how many distinct prefixes are hot, how long they are, how often each recurs, and whether the workload is prefill-heavy or decode-heavy. A handful of long, extremely hot prefixes justifies aggressive retention. A long tail of near-unique prompts does not — you will pay the memory and get the misses anyway. Measure both time-to-first-token and sustained throughput when you change the retention budget; optimising one while ignoring the other is the classic failure.
- When is reloading an offloaded KV prefix from host memory worse than recomputing it?Below a crossover length. Reload cost is bytes divided by interconnect bandwidth and scales linearly with tokens; prefill cost is FLOPs on the GPU and also scales with tokens but with a much larger constant on long sequences. Short prefixes favour recompute, long ones favour transfer. The crossover is hardware-specific — measure it rather than importing a number.
- How does grouped-query attention change this budget?Cache size scales with the key/value head count, not the query head count, so sharing eight key/value heads across sixty-four query heads cuts bytes per token roughly eightfold. Architectures that compress cached state into a lower-dimensional latent form reduce it further. This makes attention architecture a capacity decision, not only a quality one.
- Is sharing one prefix cache across tenants a security problem?Not for correctness — the shared state derives only from the tokens they have in common, and user-specific content appears past the divergence point. The residual issue is timing: hit and miss latencies differ observably, which in principle lets one party test whether a prefix has been seen before. Where the prompts themselves are sensitive, partition the cache per tenant.
- How would you decide a retention budget for a real deployment?Characterise the traffic first — how many distinct prefixes are hot, how long they are, how often each recurs, and whether output lengths make the workload decode-heavy. A few long, very hot prefixes justify aggressive retention; a long tail of near-unique prompts does not. Then change the budget and watch both time-to-first-token and sustained throughput, never one alone.
saying these in an interview costs you the question
- Treats the cache as free metadata rather than gigabytes of device memory
- Sizes it from characters or query heads instead of tokens and key/value heads
- Ignores that retained memory reduces achievable concurrency
- Assumes offloading to host memory is always cheaper than recomputing
- Believes each concurrent request needs its own copy of a shared prefix