skip to content

In prompt caching, why must stable content come before per-request content?

level: middleimportance: must knowfreq 70%

answer

  1. matching starts at token zero
  2. stable content, then variable content
  3. first differing token ends reuse
  4. shared corpus before per-user profile
  5. tool definitions are prompt text too

basics

~20 s

Prompt caching matches a new request against a stored prompt from the very first token onward, and the first difference ends the match. Anything that varies per request must therefore sit after every block you want reused.

solid answer

~50 s

Caching is **prefix**-based, not content-based: the provider compares the incoming request to a stored one starting at token zero and stops at the first mismatch, so only the unchanged head is reused. That makes ordering the whole game. Sort blocks by how often they change — system prompt, tool definitions, long static documents and few-shot examples first, then conversation history, then the current user turn and anything derived per request. One early varying token — a timestamp, a request id, a freshly retrieved snippet — invalidates everything behind it. The classic own-goal is personalization: putting a per-user profile blob ahead of a shared corpus shrinks the shared prefix to nearly zero, because every user's prompt diverges before the corpus starts. Move the profile below the shared documents and the corpus becomes reusable across the entire user population.

code

yaml · 10 lines
yaml
prompt_layout:
  - system_prompt        # changes on deploy
  - tool_definitions     # changes on deploy
  - policy_documents     # changes on content update
  - few_shot_examples    # changes when prompt is tuned
  # ---- cacheable prefix ends here ----
  - user_profile         # changes per user
  - conversation_history # grows per turn
  - retrieved_snippets   # changes per request
  - user_message         # changes per request

go deeper

for a junior

Know the one-line rule and be able to say it plainly: unchanging content first, per-request content last, because reuse is measured from the start of the prompt.

for a middle

Be ready to explain that matching runs token-by-token from position zero and stops at the first difference, and to walk a real prompt block by block, labelling each with how often it changes.

for a senior

Show that you would audit an existing prompt for hidden variability — interpolated dates, user names, dynamically ordered examples — and demonstrate the bisect: move the suspect block to the tail and watch reuse recover.

for a principal

Own the layout as a contract. Define a fixed layer order for the prompt template, make "nothing new inserted above the boundary" a review rule, and be clear that in a multi-tenant system the shared-before-personal ordering is what makes a corpus reusable at population scale.

## What a cached prefix actually is Prompt caching lets a provider store the internal state it computed for the beginning of a prompt and reuse it when a later request *begins with exactly the same tokens*. The key word is *begins*. The comparison runs from the very first token of the request and continues while tokens match; at the first divergence the match ends, and everything from that point onward is processed from scratch. There is no search for repeated blocks in the middle of a prompt, and no similarity matching. Two prompts that are 99% identical but differ in their opening sentence share nothing at all. Everything the model sees is part of that token sequence: the system prompt, the serialized tool definitions, documents you pasted in, few-shot examples, prior conversation turns, and the current user message. Tool definitions surprise people most often — they are rendered into the prompt like any other text, so renaming one tool or editing one description shifts every token after it. ## The ordering rule Order blocks by rate of change, most stable first, and you have designed a good prefix. A typical stack, from top to bottom: system prompt and role instructions (changes on deploy), tool definitions (changes on deploy, or mid-session if you load tools dynamically), long static documents and policy text (changes on a content update), few-shot examples (changes when you tune the prompt), conversation history (grows each turn), the current user message and anything retrieved or computed for this request (changes every request). The rule has a corollary that catches people: **position matters even when content is identical**. Matching is token-by-token from the start, so swapping the order of two otherwise-unchanged blocks produces a different token sequence and a miss from the swap point down. ## Locating the boundary in a real prompt Ask of each block: *on what event does this text change?* Deploy? Weekly document refresh? Per user? Per session? Per turn? Per request? Write the answer next to each block and sort. The cacheable region ends where per-session-or-slower content ends and per-request content begins. The hard part is blocks that look static but are not. Watch for a system prompt that interpolates the current date or time, a header that greets the user by name, a template that embeds a request or trace id, a "recent examples" section pulled dynamically from a store, or an example set ordered by a query-dependent score. Each of these looks like boilerplate and each one, sitting near the top, reduces reuse to whatever precedes it. ## A worked example A tutoring bot answers statistics questions with a 30,000-token textbook excerpt and a grading rubric in context. Put the textbook and the rubric immediately after the system prompt, and the student's question last. Every student, in every session, now shares that 30k-token head; only the tail is new work. Invert it and the win evaporates. If the prompt opens with "Student: Ana. Level: intermediate. Session started 14:02." and *then* the textbook, the shared prefix is a couple of dozen tokens and the textbook is recomputed for every student on every turn. Nothing about the content changed — only the order. ## Personalization is where this usually goes wrong Per-user context feels like it belongs at the top, next to the persona, because that is where identity lives conceptually. But it is the single fastest-changing block in a multi-tenant system: it differs for every user, so placing it above shared material makes the shared material private to each user. Put shared, population-wide content first, and per-user content immediately before the per-request content. The same logic applies to per-tenant policy overlays and per-conversation preferences. ## What legitimately stays variable A well-ordered prompt still has an uncached tail — the user's turn, retrieved chunks, live tool results. That is fine and expected; the point is that the tail should be small relative to the head, not that it should vanish. If your variable content genuinely dominates (say a fresh 40k-token document per request with a 500-token system prompt), prefix design has little left to give and you should be honest about that rather than shuffling blocks. ## How you know you got it wrong The symptom is reuse stuck near zero despite a large static corpus. The diagnosis is mechanical: take the suspect block, move it to the tail, and see whether reuse jumps. Bisecting from the top is fast because the answer is always "the earliest thing that changes" — the first divergent token is the only one that matters, and everything downstream of it is collateral.

  • If two blocks both change occasionally, how do you decide which one goes first?
    Order by change frequency first, then by size. The slower-changing block goes on top, because anything after a changed block is lost regardless of its own stability. When frequencies are close, put the larger block earlier so that when the smaller one changes you still lose less. If two large blocks change on genuinely independent schedules, that is a case for a breakpoint between them rather than for agonizing over the order.
  • Does moving a block that has identical content, but to a different position, break reuse?
    Yes. Matching is positional and token-by-token from the start of the request, so reordering two identical blocks yields a different token sequence and a miss from the point where the sequences diverge. This is why prompt templates should have a fixed, reviewed layer order rather than being assembled by whatever code path happens to append first.
  • Do tool definitions really participate in the prefix, or are they handled separately?
    They participate. Tool schemas are serialized into the prompt the model sees, so they occupy a position in the token sequence like any other block. Editing a tool description, adding a tool, or reordering the tool list shifts every token after it and invalidates the rest of the prefix — which is why dynamic toolsets need care about where the tool block sits.

It works like autocomplete on a phone number: the system can only reuse the part you typed from the start, so a digit changed at the front makes the whole rest new.

saying these in an interview costs you the question

  • Thinks the cache matches repeated text anywhere in the prompt
  • Believes matching is semantic similarity rather than exact tokens
  • Puts a timestamp or request id in the system prompt
  • Places a per-user profile ahead of shared documents
  • Assumes reordering identical blocks is harmless

context