skip to content

In prompt caching, what is physically stored for a cached prefix?

level: middleimportance: must knowfreq 72%

answer

  1. state, not text
  2. two tensors per layer
  3. keys and values, never queries
  4. bound to exact token ids and positions

basics

~20 s

The per-layer attention key and value tensors for every token in the prefix — not the prompt text, not an embedding, and not the model's answer. A hit reloads those tensors and skips recomputing them.

solid answer

~50 s

A transformer computes, for every token and every layer, a query, a key and a value vector. Generation already reuses past keys and values within one request so earlier tokens are never reprocessed — that is the KV cache. Prompt caching is simply the decision to keep that state alive after the request ends and let a later request adopt it. So what is stored is a stack of K and V tensors, two per layer, roughly `kv_heads × head_dim` values each per token, bound to the exact token ids and their absolute positions. Queries are not stored, because a query is consumed by the single step that creates it, while keys and values are read again by every future token. Nothing semantic is kept: it is not a saved response and not a similarity index, so a hit changes latency, not the answer distribution.

go deeper

for a junior

Be able to say that the cache holds the model's internal attention state for the prompt's tokens, not the prompt text and not a saved answer, and that reusing it saves work before generation starts.

for a middle

Explain the mechanics: two tensors per layer, keys and values only, derived from the exact token ids and their positions, with size scaling as tokens times layers times key/value head width.

for a senior

Show you know where those bytes physically live and what they compete with — GPU memory shared with in-flight requests, sometimes spilled to host RAM — and that a hit is designed to be numerically equivalent, so it is a latency change, not a quality change.

for a principal

Own the framing: because the artifact is positional attention state rather than text, every architectural rule about caching is derivable rather than memorised, and any design that assumes semantic or fuzzy reuse is reaching for a different mechanism entirely.

## The problem the KV cache solves A transformer generates one token at a time. To produce the next token it runs a forward pass in which, at every layer, the current token's *query* vector is compared against the *key* vectors of all preceding tokens, and the resulting weights mix those tokens' *value* vectors into the current position's representation. Crucially, the keys and values of earlier tokens do not change as generation proceeds — they are a function of those earlier tokens and their positions alone. Recomputing them at every step would make generating an N-token answer cost O(N²) full passes. Storing them makes each step cost one pass over a single token. That store is the **KV cache**, and it exists inside every production LLM server whether or not a "prompt caching" feature is advertised. ## What prompt caching adds Ordinary KV caching lives and dies with one request. Prompt caching keeps the same tensors after the request finishes, keyed by the token sequence that produced them, so a *different* request beginning with the same tokens can adopt the state instead of rebuilding it. Nothing new is invented — the same object is simply given a longer lifetime and a lookup path. ## The shape of what is stored Per layer there are exactly two tensors, K and V. Per token per layer each holds about `num_kv_heads × head_dim` numbers. So the total is roughly: `2 (K and V) × layers × kv_heads × head_dim × tokens × bytes_per_value` For a 70B-class configuration — 80 layers, 8 key/value heads of width 128, 16-bit values — that is about 320 KB **per token**. This is why cache size is quoted in tokens rather than characters, and why a long shared prefix is a real memory object, not bookkeeping. ## What is not stored - **Not the prompt text.** Re-tokenizing text is microseconds; the expensive part is the attention state derived from it. - **Not queries.** A query is used once, at the step that produces it. Cached tokens never issue new queries, so there is nothing to reuse. - **Not the response.** A response cache returns a previously produced answer for an identical input; a KV cache returns intermediate state and the model still generates fresh output. Conflating the two is the single most common misunderstanding. - **Not an embedding.** There is no similarity search. Matching is exact over token ids; a semantically identical paraphrase is a complete miss. Because the cached tensors are numerically what recomputation would have produced, a hit is designed to be quality-neutral. It does not make sampling deterministic: temperature and seed still govern that, and floating-point results can differ marginally when a request is batched differently. ## Position is baked in Modern models apply rotary position information to queries and keys before the dot product. A cached key therefore carries its absolute position with it. That is why cached state cannot be relocated to a different offset in a new prompt, and it is one of the two reasons — the other being causal dependency on preceding tokens — that reuse is defined only on a *prefix*, never on an arbitrary middle chunk. ## Where the bytes live On a self-hosted server the tensors sit in GPU high-bandwidth memory alongside the model weights and the KV of every in-flight request; some stacks spill cold entries to host RAM or NVMe and pull them back on a hit. Either way, a "hit" means the attention path finds the tensors already materialised for the matched span and runs prefill only over the tokens that follow. ## Why this framing pays off in an interview Once you can say "the cache is per-layer key and value tensors bound to token ids and positions," every downstream rule follows without memorisation: reuse must be a contiguous prefix (later state depends on earlier tokens), matching must be exact at the token level (the tensors were derived from those exact ids), what improves is time-to-first-token and not streaming speed (only prefill is skipped), and retention costs real GPU memory (hundreds of KB per token). Candidates who memorise the rules without the mechanism get stuck the moment an interviewer asks "why?"

  • Does a cache hit make the model's response deterministic?
    No. The hit restores identical intermediate state, but decoding is still stochastic — temperature, top-p and the seed govern the sampled tokens. It is also not a response cache, so nothing pre-written is replayed. Providers aim for numerical equivalence with recomputation, though results can shift slightly when a request lands in a different batch.
  • Why are query vectors not cached alongside keys and values?
    A query is used exactly once, by the position that produced it, to score against earlier keys. It is never read again, so storing it buys nothing. Keys and values are the opposite: every future token in the sequence reads them at every layer, which is precisely what makes them worth keeping.
  • How does grouped-query attention change the size of what is stored?
    With grouped-query attention many query heads share a smaller number of key/value heads, and cache size scales with the key/value head count, not the query head count. Going from 64 query heads to 8 key/value heads cuts stored bytes per token roughly eightfold, which is a large part of why long shared prefixes are affordable at all.

saying these in an interview costs you the question

  • Says the cache stores the prompt text and re-tokenizes it faster
  • Confuses it with a response cache that replays a saved answer
  • Thinks prompts are embedded and matched by semantic similarity
  • Claims query vectors are cached along with keys and values
  • Assumes cache size depends on characters rather than tokens and layers

context