Why does a per-request retrieved exemplar block destroy prompt-cache hits?
answer
- caches match a prefix, not a set of blocks
- volatile content poisons everything after it
- order blocks by how often they change
- bundles beat bespoke when reuse matters
- a cache entry never read still costs
basics
~20 sPrompt caching matches an exact token prefix from the start of the request. Retrieved exemplars differ on every call, so everything from that block onward is uncached — placed near the top, the entire prompt is reprocessed every single request.
solid answer
~50 sCaching is prefix-based: providers reuse the processed state of a request only up to the first token that differs from a previous one. A retrieved exemplar block is by construction different every call, so it acts as a cache barrier — every token after it is a miss. If the block sits above the system instructions and tool definitions, you lose all of it. The fix is ordering: put everything stable first (system prompt, tool schemas, static instructions), then the volatile retrieved block as late as possible, so the stable prefix still hits. If the demonstrations must precede other large static content, consider bucketing instead — route each request to one of a small set of precomputed exemplar bundles by intent, so each bundle's prefix is reused across many requests. Note also that where providers charge a premium to *write* a cache entry, a prefix that never repeats costs more than not caching at all.
go deeper
Know that prompt caching depends on the beginning of the prompt staying identical between requests, and that anything chosen per request must therefore go near the end.
Explain the prefix rule and its consequence: a volatile block invalidates every token after it, however static that content is. Be able to reorder a prompt by volatility and say what each block costs.
Do the accounting out loud for a real prompt — stable preamble versus retrieved block, cached-read price versus fresh tokens — and weigh the measured accuracy lift against the multiplied per-call cost. Mention bucketing as the middle path and the write premium as a trap.
Own the frontier: at what lift does per-request retrieval justify losing cache economics across the whole product, and where do you standardise a bundled-exemplar layout so teams cannot silently reintroduce a volatile prefix.
## Why the cost is not obvious The accuracy argument for dynamic exemplars is made on a benchmark, where every request is priced the same. The cost argument is made in production, where it is not. Providers cache the processed prefix of a prompt and price a cached read at a small fraction of a fresh input token, so a long, stable system prompt is nearly free on the second and subsequent calls. A retrieved exemplar block quietly opts you out of that discount, and the accounting change can dwarf the token count of the block itself. This reflects provider prompt-caching practice as of mid-2026: the mechanism is prefix matching, the discount on a cache read is large, and several providers additionally charge a premium on the call that writes a cache entry. ## The prefix rule Caching works on an exact token prefix from position zero. The provider can reuse work only up to the first token that differs from a prior request. This has one consequence that governs all prompt layout: **anything volatile poisons everything after it.** It does not matter that the tool schemas below your exemplar block are byte-identical to yesterday's — they sit behind a changed token, so they are recomputed. A per-request retrieved block is maximally volatile. Its content depends on the incoming input, so no two requests share it except by coincidence. ## Worked example Take a bug-triage assistant with a 6,000-token stable preamble — system instructions, output schema, tool definitions — plus eight retrieved resolved tickets at roughly 250 tokens each, so 2,000 tokens of demonstrations, plus a short query. Ordered *exemplars first*, every request pays full price for all 8,000-plus tokens. Ordered *preamble first, exemplars last*, the 6,000-token preamble is a cache read at a fraction of the price and only the 2,000-token block plus the query is charged fresh. The token count sent is identical in both layouts; the bill is not. That gap — a multiple on effective input cost per call, not a few percent — is the real answer to "what does dynamic retrieval cost?", and it is why the question comes up in senior interviews at all. The same ordering also cuts prefill latency, because the reused prefix does not need reprocessing. ## The layout rule Order prompt blocks by volatility, most stable first: 1. System instructions and persona 2. Tool and output schemas 3. Static, hand-curated demonstrations if you keep any 4. Retrieved exemplars 5. The user's input and per-turn state This is a rule about *cost*, not about pedagogy, and it can conflict with what the model responds to best — some tasks do better with demonstrations immediately before the query, which happily agrees with the rule, and some templates want them early, which does not. When the conflict is real, measure both: the accuracy delta from moving the block is often small enough that the cost delta decides it. ## Bucketing: getting the cache back If the demonstrations must sit high in the prompt, or if the block is large enough that missing on it hurts anyway, trade some per-request fit for reuse. Precompute a modest set of exemplar bundles — one per intent, per queue, per document type — and route each request to a bundle instead of assembling a bespoke block. Each bundle now has a stable prefix that many requests share, so it caches. You give up the last increment of similarity and keep most of the task-fit win, which on many tasks is a good trade. It also restores reproducibility: a bundle id is a thing you can pin in an eval. ## The write premium Where a provider charges more for the tokens on a call that *creates* a cache entry, caching only pays if the entry is read again before it expires. A prefix that is unique per request writes an entry that will never be reused, so you pay the premium for nothing. Two practical implications: do not mark a volatile block as cacheable, and when traffic is sparse enough that entries expire between requests, check whether caching is earning its premium at all. ## What to measure Instrument cache-read versus fresh input tokens per request and watch the ratio, not just the total spend — a regression that reorders one prompt block shows up there long before it shows up in a monthly invoice. Alongside it, track the accuracy lift dynamic retrieval buys over a static or bucketed block. Those two numbers, together, are the entire decision: if the lift is a point or two and the effective input cost multiplies, the honest recommendation is bucketing or a fixed block.
- You move the retrieved block below the tool schemas and cache hits appear, but quality drops slightly. How do you decide?Quantify both sides on the same traffic slice: the accuracy delta from the move, and the change in effective input cost and prefill latency. If the quality loss is within noise, ordering wins outright. If it is real, look for a middle path — a shorter retrieved block, or bundled exemplars high in the prompt — before paying a multiple on every request for a fraction of a point.
- Would putting the retrieved exemplars last always be enough to keep the cache useful?Only if what precedes them is genuinely stable. A timestamp, a request id, a rotating persona line or per-user preamble injected above the block is just as volatile and breaks the prefix the same way. Audit everything before the block for hidden variability — dynamically rendered dates and ids are the usual culprits, and they cost you the cache even without any retrieval.
- When does bucketing exemplars into a few precomputed bundles beat per-request retrieval?When traffic concentrates on a handful of intents, when the block is large, and when latency or cost budgets are tight. A per-intent bundle keeps most of the task-fit gain, caches across many requests, and is reproducible — you can pin a bundle id in an eval. Per-request retrieval wins when the input space is long-tailed enough that no small set of bundles covers it.
saying these in an interview costs you the question
- Thinking the cache matches individual blocks rather than a prefix
- Assuming stable content below a volatile block still caches
- Comparing only token counts and ignoring cached-read pricing
- Marking a per-request block as cacheable and paying the write premium
- Treating the retrieval hop's latency as the only added cost