skip to content

Why does editing one word early in a long system prompt void the whole cache?

level: middleimportance: must knowfreq 78%

answer

  1. attention only looks backwards
  2. later state was conditioned on earlier tokens
  3. insertions shift every following position
  4. matching halts at the first differing token

basics

~20 s

Attention is causal: every later token's cached key and value tensors were computed from all the tokens before it. Change a token near the top and the stored state for everything after it is stale, so the server must rebuild the entire prompt.

solid answer

~50 s

Take a 30,000-token system prompt with a single word edited at token 12. Token 12's own key and value change immediately, and from the first attention layer onward every subsequent position mixes in information from token 12 — so tokens 13 through 30,000 have different hidden states and therefore different keys and values too. None of that stored state is valid any more. Two further effects compound it: if the edit changes the token count, every following token shifts position, and keys carry absolute position information, so even unchanged text lands in the wrong place; and matching is done on token ids, so a change as small as a trailing space can alter the tokenization. Matching therefore stops at the first differing token, which is why reuse is defined on a contiguous prefix and why appending is free while prepending is fatal.

go deeper

for a junior

Be able to state the rule and the reason in one breath: each token's stored state was built from the tokens before it, so an early change makes everything after it invalid.

for a middle

Walk the mechanism layer by layer — the edited token's keys change, later positions attend over them from the first attention layer, and the change propagates upward — and add that insertions also shift absolute positions.

for a senior

Bring the diagnosis: near-zero reuse on a big static preamble almost always traces to an early-position mutation such as an injected timestamp, an interpolated user field, or an unstable serialization order, and you should be able to reason to that from the mechanism alone.

for a principal

Own the asymmetry as a design constraint: the cost of a prompt edit scales with how much prompt follows it, which makes ordering of prompt content an architectural decision with a measurable throughput consequence rather than a stylistic one.

## Causal attention is the whole answer Decoder-only transformers use a causal mask: the representation of the token at position *i* may depend on positions 0..*i*, and never on anything later. Read that in the other direction and you get the caching rule. Because position *i* can see everything before it, its computed state is a function of that entire preceding span. Alter any token in the span and position *i*'s state is a function of different inputs — so the cached tensors for position *i* are simply wrong. Concretely, with a 30,000-token prompt and a one-word edit at token 12: 1. Token 12's own key and value vectors change at every layer, since they derive from its embedding. 2. At layer 1, tokens 13 onward attend over layer-0 output that now includes a different token 12, so their hidden states change. 3. Changed hidden states produce changed keys and values, which propagate upward through all remaining layers. (The pedantic footnote: layer-0 keys and values of later tokens depend only on their own embedding and position, so strictly they survive. From the first attention layer up, nothing does — and a KV entry is the whole per-layer stack, so in practice everything from position 12 onward is discarded.) ## Position binding compounds it Modern models fold rotary position information into queries and keys before the attention dot product. A cached key is therefore stamped with the absolute position it was computed at. If the edit changes the number of tokens — inserting a date, deleting a sentence — every following token now sits at a different index, and the cached keys, though derived from identical text, encode the wrong positions. Even a pure substitution that preserves length still fails on the causal-dependency ground above. ## Matching is on token ids, not on your intent The lookup compares token id sequences. Two prompts that look identical to a human can tokenize differently: byte-pair merges cross visual boundaries, so a trailing space, a different Unicode normalization, a swapped quote character or a re-serialized JSON blob can shift ids from that point onward. From the server's perspective this is not a near-miss to be forgiven; it is a divergence at token *k*, and everything from *k* on is recomputed. ## Why the middle can never be spliced in A natural question is why the server cannot keep the unchanged 29,988 tokens and just patch around the edit. Two reasons: those tokens' states were computed while attending to a version of the prefix that no longer exists, and their keys are position-stamped for offsets that may have moved. Splicing a cached middle chunk into a new context yields attention over state the model never actually produced. Research systems explore re-encoding or blending cached chunks to recover part of this, but standard serving offers exactly one guarantee — longest identical **prefix**, matched from token zero. ## What is cheap The asymmetry is the useful takeaway. Appending costs only the prefill of the appended tokens, because nothing before them changed: this is why a growing conversation keeps hitting the cache turn after turn. Prepending, inserting, or editing anywhere near the top costs a full rebuild of everything downstream. The cost of an edit is not proportional to its size; it is proportional to how much prompt follows it. ## How this shows up in practice The classic symptom is a system that reports near-zero reuse despite an enormous, apparently static preamble. The usual causes are all early-position mutations: a timestamp or request id injected at the top, a user's display name interpolated before the shared instructions, a tool list whose order is not stable across processes, or a serializer that emits map keys in hash order. Each one changes a token within the first few hundred positions and forfeits everything after it. The mechanism explains why the fix always takes the same shape — stabilise what comes first — without needing to memorise any vendor's rules. ## The one-line version Cached state is state *conditioned on everything before it*. Change a prefix token and you have changed the condition, so every cached entry downstream describes a prompt that no longer exists.

  • If the changed token is the very last one in the cached prefix, what is still reusable?
    Almost all of it. Matching walks from token zero and stops at the first divergence, so everything strictly before the edit stays valid. Two caveats: implementations that hash fixed-size token blocks round the match down to the last fully identical block, and some servers enforce a minimum matched length before a hit is worth recording.
  • Why can't a server splice a cached middle section into a different prompt?
    Because that section's keys and values were produced while attending to a specific preceding context, and they carry the absolute positions they were computed at. Dropping them into a new prompt gives the model attention state it never generated, at offsets that may have moved. Only a prefix match from token zero preserves both conditions.
  • Does appending a new user turn to a cached conversation cost anything?
    Only the prefill of the new tokens. Nothing before them changed, so the entire earlier prefix — instructions, tool definitions, all prior turns — stays valid and is reused. This is exactly why multi-turn conversations are the workload where caching pays best: each turn's miss is bounded by the size of that turn.

saying these in an interview costs you the question

  • Thinks only the edited token needs recomputing
  • Assumes text after the edit is unaffected because it was untouched
  • Believes same-length edits preserve the cache
  • Treats matching as fuzzy or semantic rather than exact on token ids
  • Expects a cached middle chunk to be spliceable into a new prompt

context