skip to content

When is quantizing the KV cache to FP8 the right way to buy serving capacity?

level: principalimportance: nice to knowfreq 30%

answer

  1. bytes per entry is a lever too
  2. half the bytes, twice the sessions
  3. it speeds decode as well as freeing memory
  4. keys carry the outliers
  5. damage shows up at long context first

basics

~20 s

Cache quantization is right when memory, not quality headroom, is the binding constraint and your own evals show the loss is tolerable at your longest contexts. Halving bytes per entry roughly doubles resident sessions and also speeds bandwidth-bound decode.

solid answer

~50 s

Narrowing the cache from bf16 to an 8-bit format halves bytes per token, and 4-bit formats quarter them, which converts almost linearly into concurrent sessions: on an 80 GB card holding a 16 GB model, a workload that fit four 100K-token sessions fits about nine. It is also a **latency** win, because decode is memory-bandwidth-bound and a smaller cache is fewer bytes read per generated token. The cost is accuracy, and it is unevenly distributed: keys carry per-channel outliers and are typically more sensitive than values, so quantizers scale the two differently; degradation shows up first on long contexts and precise-retrieval tasks, not on short prompts. Reach for it when memory is genuinely the binding constraint, validate on your own long-context eval rather than a public benchmark, and confirm the hardware and kernels handle the format natively so dequantization does not eat the win.

code

python · 7 lines
python
def sessions_per_gpu(gpu_gb, weights_gb, tokens, bytes_per_token):
    free = (gpu_gb - weights_gb) * 1e9
    return int(free // (tokens * bytes_per_token))

# 32 layers, 8 KV heads, head_dim 128 -> 128 KB/token at bf16, 64 KB at 8-bit
print(sessions_per_gpu(80, 16, 100_000, 131072))  # 4 sessions, bf16 cache
print(sessions_per_gpu(80, 16, 100_000, 65536))   # 9 sessions, 8-bit cache

go deeper

for a junior

Know that the cache can be stored in a narrower number format, that this roughly halves or quarters its memory, and that the tradeoff is some loss of accuracy rather than a free win.

for a middle

Explain both effects: fewer bytes per entry raises how many sessions fit and reduces the bytes read per decode step, which helps bandwidth-bound generation. Be able to state where the quality loss concentrates.

for a senior

Show you would gate it behind a workload-specific long-context eval against a baseline, know that keys and values need different treatment, and check that the format is handled natively rather than dequantized on every read.

for a principal

Own the choice among the levers. Frame it as a stated exchange rate — quality points against sessions per GPU — and decide where quantization, retained-context policy, model architecture and more hardware each earn their place across the fleet.

## What the lever actually does Every entry in the cache is a vector of numbers, and the format those numbers are stored in is a free parameter. Moving from 16-bit to 8-bit halves bytes per token; 4-bit formats quarter them. Nothing about the architecture changes — the same number of entries are stored, each one smaller. This is the most mechanical of the capacity levers, which is exactly why it is tempting and why it needs a quality gate. ## Two wins, not one Capacity is the obvious win. On an 80 GB accelerator carrying a 16 GB model, roughly 60 GB is available; at 128 KB per token a 100,000-token session costs about 13 GB, so four sessions fit. Halve the entry width and about nine fit. The second win is subtler and often larger in user-visible terms: decode is memory-bandwidth-bound, and at long context the cache read is a significant share of the bytes moved per generated token. Cutting that read in half reduces inter-token latency for exactly the long sessions that were slowest. Capacity and speed move together here, which is unusual among serving tradeoffs. ## The cost, and where it hides Quantization is lossy and the loss is not uniform. Two patterns matter. First, keys and values behave differently: key vectors exhibit strong per-channel outliers, which is why quantizers commonly scale keys along a different axis than values, or keep keys at higher precision when only one can be preserved. A naive uniform scheme applied to both usually degrades more than a scheme aware of that asymmetry. Second, degradation is context-length-dependent. Short prompts often look untouched, while errors compound across tens of thousands of accumulated entries and surface as weakened long-range retrieval — the model loses a detail from early in a document, or a completion stops matching an identifier defined 5,000 lines up. A pass on short-prompt benchmarks tells you nothing about this. ## How to decide Treat it as an experiment with a pre-declared metric, not a config flip. Establish the binding constraint first: if you are failing on memory at your target concurrency, this lever applies; if you are failing on prefill compute or on quality, it does not, and it will make the second worse. Then measure on a workload-specific eval at your realistic tail context length — for a code-completion service, acceptance rate on real long files; for a document assistant, multi-hop retrieval across the full document. Compare against the bf16 baseline on the same traffic, not against a leaderboard. Decide with the business quantity in hand: a fraction of a percent of acceptance rate against a doubling of sessions per GPU is a decision someone can actually make; "the benchmark barely moved" is not. ## Where it sits among the alternatives The competing levers each have a different bill. Retaining less context per session is free in memory and immediate, but changes what the product remembers, which is a user-visible regression rather than a numerical one. Choosing a model with a cheaper cache design — more query heads sharing each key/value head, a compressed latent, interleaved recurrent layers — is often the biggest single reduction available, but it means a model migration and a fresh quality evaluation. Adding or sharding across hardware always works and always costs money. Offloading cold cache to host memory buys capacity while directly harming the bandwidth-bound decode path, so it suits idle sessions rather than actively generating ones. Cache quantization is attractive precisely because it is orthogonal to all of these: it multiplies with the architecture choice rather than substituting for it. ## Practical cautions Confirm the format is supported natively along the whole path. On hardware and kernels that handle narrow formats in the attention computation directly, the saving is real end to end; where the cache must be widened back before use, you pay dequantization work on every read and keep only the capacity benefit. Check that both prefill-written entries and decode-written entries use the same scheme, so quality does not depend on how a session was assembled. And keep the ability to roll back per deployment: cache precision is one of the easier things to flip if a quality regression shows up in production, provided you did not build anything that assumes the narrower footprint. ## The honest summary As of mid-2026, 8-bit caches are routine in production serving and 4-bit cache formats run natively on current data-center hardware with reported quality costs small enough that many teams accept them. "Routine" is not "free": the loss is real, it concentrates in long-context recall, and the only evidence that matters is your own eval on your own traffic at your own context lengths.

  • Why are keys usually more sensitive to quantization than values?
    Key vectors show strong per-channel outliers, so a single scale across the vector wastes range on the outlier channels and crushes the rest. Values are better behaved. Practical schemes therefore scale keys along a different axis than values, or keep keys wider when only one can be preserved.
  • What eval would convince you a narrower cache is safe to ship?
    One built from your own traffic at your tail context length, scored on the metric the product is judged by — completion acceptance rate, multi-hop retrieval accuracy across a full document — against a bf16 baseline on the same inputs. Short-prompt public benchmarks routinely show no change while long-context recall has degraded.
  • Would you quantize the cache before switching to a model with a cheaper cache design?
    Usually the architecture change is the larger reduction, but it is also a migration with a full re-evaluation. Cache precision is a same-model, reversible flip, so it is the right first experiment; the two are multiplicative, so a cheap-cache model with a narrow cache is the eventual destination if memory stays binding.

saying these in an interview costs you the question

  • Treats cache quantization as lossless because short prompts look fine
  • Quantizes keys and values with one uniform scheme and no measurement
  • Validates only on public benchmarks rather than long-context traffic
  • Reaches for it when the bottleneck is prefill compute or quality
  • Ignores whether the kernels support the format without dequantizing

context