How do you stop just-in-time loading from invalidating a cached prompt prefix?
answer
- exact prefix match, nothing later survives
- order layers by volatility
- append-only, never splice
- edits buy space and cost a cold prefix
basics
~20 sKeep the prompt append-only. Put stable material — system instructions, tool definitions, the catalog — at the front, never reorder or edit it, and append everything fetched mid-session at the end. Anything inserted above a cache boundary invalidates every cached token after it.
solid answer
~50 sPrefix caching matches on an exact token prefix, so the cache survives only while everything before the changed point is byte-identical to the previous request. Just-in-time loading fights that if fetched documents are spliced into the middle of the prompt — say, injected next to the system instructions or grouped with the turn that requested them — because every token after the splice point stops matching and gets reprocessed. **The fix is ordering discipline.** Design the context as a stable prefix followed by an append-only tail: instructions, tool definitions and any preloaded catalog first, then the conversation, then each tool result appended in arrival order. Loaded content then lands after the cache boundary and the prefix survives. The same discipline explains why mid-session edits — deleting stale tool results, rewriting earlier turns, re-sorting retrieved chunks — are expensive: they buy window space at the price of a cold cache, and as of mid-2026 that is a first-class budgeting trade rather than a detail.
code
python · 15 linesSTABLE_PREFIX = [
{"role": "system", "content": "You are a warehouse analyst."},
{"role": "system", "content": "CATALOG: orders(12M) returns(400k) regions(120)"},
]
def build_request(history, fetched):
# stable prefix first, conversation next, fetched payloads appended last
return STABLE_PREFIX + history + [
{"role": "user", "content": f"[{ref}]\n{body}"} for ref, body in fetched
]
history = [{"role": "user", "content": "Weekly returns by region?"}]
fetched = [("ddl:returns", "returns(id, region_id, created_at, amount)")]
for m in build_request(history, fetched):
print(m["role"], "|", m["content"][:40])go deeper
Know that the prompt is cached from the beginning up to the first thing that changed, so new content should be added at the end rather than inserted in the middle.
Explain the ordering rule — stable instructions and tool definitions first, conversation next, fetched results appended — and why an edit anywhere invalidates everything after it.
Diagnose real misses: volatile fields high in the prompt, non-deterministic serialization, re-sorted context blocks. Weigh a deliberate invalidation to reclaim window space against the recurring cost of carrying dead tokens.
Treat cache hit rate as a budget line with an owner. Decide the prompt-assembly contract every component must honour, and set the policy for when reclaiming window space is worth a cold prefix.
## What prefix caching gives you and what it demands Providers cache the computed state of a prompt prefix so that a follow-up request sharing that prefix skips recomputing it. The reported savings are large — commonly cited around 90% on cost and 85% on latency for the cached portion — which makes cache hit rate a real budget line for any multi-turn agent. The demand is exact-prefix matching. The cache is keyed on the token sequence from the start of the prompt. It survives up to the first position where this request differs from the cached one; from there on, everything is recomputed. That single property drives every rule below. ## Why JIT loading is the natural enemy of a stable prefix Just-in-time loading means the prompt grows unpredictably during a session: a document arrives at turn three, a schema at turn five. If the harness places that content anywhere except the end, it moves the boundary. Three common ways teams break their own cache: - **Injecting fetched content into the system block** — a tempting way to make it feel authoritative, and the worst possible position, because it invalidates literally everything. - **Re-sorting or re-grouping** — collecting all loaded documents into one 'context' section, re-sorted by relevance each turn, so the section's bytes change every turn. - **Editing earlier turns** — trimming old tool results, replacing a payload with a summary, or renumbering references. Each edit resets the cache from that point forward. A fourth, subtler one: **non-deterministic serialization**. If the tool result is a JSON object whose key order or whitespace varies between runs, or carries a timestamp, the bytes differ even when the content does not, and the prefix stops matching for reasons nobody can see in the rendered prompt. ## The discipline: stable prefix, append-only tail Structure the request in layers ordered by volatility, most stable first: 1. **System instructions** — fixed for the deployment. 2. **Tool definitions** — fixed per session. 3. **Preloaded catalog or index** — the tier-one listing, if it is stable for the session. 4. **Conversation history** — grows at the end. 5. **Tool results from JIT fetches** — appended in arrival order, never reordered. With this ordering, every fetch extends the prompt rather than mutating it, so each turn's cache boundary moves *forward* and the previously cached portion stays valid. This is the concrete meaning of the advice to design context append-only. Two corollaries. First, anything volatile — the current timestamp, a per-request user ID, a randomized instruction — must go *after* the stable layers, not in the system block, or it poisons the whole prefix on every request. Second, a catalog that refreshes mid-session should be treated as volatile and placed with the conversation, not with the instructions. ## When breaking the cache is the right call The discipline is not absolute. Window space is finite, and eventually an agent must reclaim it — by dropping stale tool results, summarizing, or reinitializing the session. Each of those invalidates the cache from the edit point, and that is sometimes correct: a window filled with dead payloads produces worse answers, and worse answers cost more than a cold prefix. The way to reason about it is to compare the recurring cost of carrying the dead tokens — paid every turn for the rest of the session — against the one-time cost of reprocessing the prefix. Cleanups that are rare, batched, and large win; cleanups that are frequent and small lose, because you pay a cold prefix repeatedly to reclaim a little space each time. Batching is the lever: do one substantial cleanup rather than trimming continuously. ## What to measure Instrument cache hit rate per turn alongside token counts. A hit rate that collapses at a particular turn number points at a specific mutation — often a harness feature nobody associated with the cache, like re-rendering a header with a timestamp. Also log the *position* of the first mismatch when you can, because that identifies the offending layer directly. ## The interview-shaped summary JIT loading and prefix caching are compatible, but only under an ordering rule: stable things first, never edited; fetched things appended at the end. Where they genuinely conflict — reclaiming space from a full window — treat cache invalidation as a cost to be paid deliberately and in batches, not accidentally on every turn.
- Where should a per-request timestamp or user ID go?After the stable layers, never in the system block. Anything that changes per request sits at the very start of a naively built prompt and invalidates the entire cache on every call — the maximum possible damage from the smallest possible field. Put volatile values in the latest user turn, or omit them when the model does not actually need them. This is one of the most common self-inflicted cache misses.
- If the window fills, is it ever right to delete earlier tool results despite the cache cost?Yes. Dead payloads degrade answers every turn, and a cold prefix is a one-time cost. Compare the recurring cost of carrying the tokens against the single reprocessing charge, and batch the cleanup — one substantial reclamation beats trimming a little on every turn, which pays the cold-prefix penalty repeatedly for small gains.
- Why can a cache miss occur even when the visible prompt text looks unchanged?Because matching is on tokens, not on rendered meaning. Non-deterministic JSON key order, whitespace or trailing-newline differences, a re-rendered header containing a clock value, or a library that reorders tool definitions all change the byte sequence while the prompt looks identical. Serialize deterministically and pin the order of every generated block if you want reliable hits.
saying these in an interview costs you the question
- Believes the cache matches on meaning rather than exact tokens
- Injects retrieved documents into the system prompt
- Re-sorts loaded context by relevance on every turn
- Puts a per-request timestamp at the top of the prompt
- Treats any cache invalidation as automatically unacceptable