Why does a timestamp at the top of a system prompt destroy cache hit rate?
answer
- matching is exact, never fuzzy
- one changed byte, whole tail cold
- volatile field sits at position zero
- you still pay to write dead entries
- hoist it out or coarsen it
basics
~20 sPrompt caches match an exact prefix from the first token onward. A line like "Current time: 14:32:07" differs on every request, so the match fails at the very first tokens and the entire prompt behind it is reprocessed and re-stored, forever.
solid answer
~60 sPrompt caching is an **exact prefix match**, not a similarity match. The provider compares your prompt to a stored entry token by token and reuses only the leading run that is identical; the first divergence ends the match and everything after it is processed fresh. A per-second timestamp at position zero diverges on literally every request, so nothing is ever reusable — and worse, each call also pays to *write* a new entry that no future call will ever read, so you get the cost of caching with none of the benefit. The fix is to get the volatile value out of the cacheable region: put the clock in the variable tail of the prompt or in the user turn, or let the model fetch it through a tool. If it genuinely must sit near the top, coarsen it — a date, or an hour bucket, changes the prefix once a day or once an hour instead of once a request, which turns a permanent 0% into a near-perfect hit rate with one boundary miss per bucket.
go deeper
Be able to state that prompt caching matches an exact prefix, so anything that changes per request must not appear near the top of the prompt. Naming the timestamp as the classic offender is enough at this level.
Explain the prefix rule precisely: the match ends at the first differing token and everything after it is reprocessed, so stable content behind a volatile field is unrecoverable. Mention that each call also pays to write an entry nobody reads.
Show how you would find this in a live system — prefix hashing, hash cardinality per route, correlating the step change with a deploy — and then rank the fixes: hoist, expose as a tool, or coarsen into buckets aligned with traffic.
Own the prevention: a prompt-assembly layer where the reusable region is constructed from declared-stable inputs only, with a test that fails when a prefix hash varies across two synthetic requests. Argue why review discipline alone will not hold across a growing team.
## The rule that explains the behaviour A prompt cache stores the provider's processed representation of a *prefix* of your prompt and reuses it when a later request starts with exactly the same tokens. Two properties follow, and together they explain the whole failure: 1. **Matching is exact, not semantic.** "Current time: 14:32:07" and "Current time: 14:32:08" are not similar-enough; they are different tokens and the match stops there. Nothing about prompt caching is fuzzy — that is a different mechanism entirely. 2. **Matching is a prefix, so it is order-sensitive and unrecoverable.** The cache can reuse the leading run of identical tokens and nothing past the first difference. Content that is stable but sits *behind* a volatile field cannot be salvaged, no matter how large or how unchanging it is. A per-second clock at the very top of the system prompt is the worst possible placement: it diverges at token ~4, so the run of reusable tokens is essentially zero and every downstream byte — the 20k of policy, the tool definitions, the few-shot examples — is reprocessed on every single call. ## The second, quieter cost It is tempting to think the outcome is simply "caching does nothing". It is worse than neutral. Storing a prefix is not free: the tokens written into a cache entry are processed and stored, and providers price that write differently from an ordinary read. If every request writes a brand-new entry that no subsequent request can ever match, you are paying a surcharge on every call to populate a cache with garbage. This shows up unmistakably in instrumentation as cache-creation tokens tracking cache-read tokens one-for-one, or as creation with essentially no reads at all. ## Recognising the pattern beyond timestamps A clock is the canonical example because it is so visibly per-request, but the same defect appears in many disguises, and all of them share the shape "a value that varies per request has been placed inside the region you wanted to reuse": - A greeting that interpolates the end user's name or account ID into the system text. - A request or trace ID injected for debugging, often added innocently during an incident and never removed. - An A/B flag or feature-toggle snapshot rendered into the preamble. - A random seed, or few-shot examples shuffled per call to "reduce position bias". - A session counter, remaining-quota figure or token-budget line that decrements as the conversation runs. - Retrieved documents pinned above the instructions in relevance order, where the order changes with the query. Each of these looks harmless in a diff and each converts a large reusable block into dead weight. ## The fixes, in order of preference **Hoist it out.** Move the volatile value below everything you want cached — into the variable tail, or into the user turn. The model reads it just as well from there, and the entire stable region above it becomes reusable. This is almost always the right answer and costs nothing. **Let the model ask for it.** If the value is only needed occasionally, expose it through a tool call rather than pushing it into every prompt. The prefix then never carries it at all. **Coarsen it.** When something time-like genuinely has to sit high in the prompt, reduce its resolution to the coarsest value the task tolerates: a date instead of a second, or an hour bucket. Now the prefix changes once per bucket instead of once per request, which costs you a single cold call per bucket and gives you hits for everything in between. Be honest with yourself about what the task needs — a support agent almost never needs second-level precision, and if it does, a tool is the better answer. **Bucket deliberately, not accidentally.** If you coarsen, align the buckets with traffic. An hourly bucket that rolls over at the top of the hour produces one cold request per hour per prefix; that is a rounding error. A bucket that rolls every 30 seconds under bursty traffic is barely better than no bucketing. ## How to find it in an existing system Do not read the prompt looking for the problem — the whole point is that the offending line is easy to miss. Hash the exact bytes of the cacheable region at the call site and log the hash. If a route that should have a handful of prefix shapes emits thousands of distinct hashes per hour, you have confirmed content variation. Then diff two prefixes with different hashes programmatically and the volatile field falls out immediately. If a deploy caused it, the hash cardinality steps up sharply at that release. ## What not to conclude A longer retention window does not help a prefix that changes on every request — nothing is being evicted, nothing was ever reusable. Nor does retrying, nor does asking the provider to "enable caching harder". The problem is entirely in what you sent, and the fix is entirely in your prompt assembly code.
- The team says the timestamp must stay because the assistant needs the current time. What do you propose?Two options. Either move the time into the variable tail of the prompt or the user turn, where it does not sit inside the reusable region — the model reads it identically from there. Or expose the clock as a tool the model calls when it actually needs it, which keeps the value out of the prompt entirely. If it truly must appear near the top, coarsen it to a date or hour bucket so the prefix changes once per bucket rather than once per request.
- Besides a clock, what other per-request values commonly end up inside a prefix that was meant to be reused?End-user names or account IDs interpolated into the system text, debug trace IDs added during an incident and never removed, A/B or feature-flag snapshots, per-call shuffled few-shot examples, decrementing quota or budget lines, and retrieved documents pinned above the instructions in query-dependent order. All share the same shape: a value that varies per request placed inside the region you wanted stable.
- Would raising the cache retention window fix a 0% hit rate caused by a per-second timestamp?No. Retention governs how long an unused entry survives; here nothing is ever reusable in the first place, because every request produces a prefix that has never been seen before. The entries are not being evicted too early — they are being written and then never matched. The only fix is to stop varying the prefix, by hoisting or coarsening the volatile field.
saying these in an interview costs you the question
- Believing the cache matches semantically similar prompts
- Assuming only the changed line is reprocessed, not the tail
- Thinking a 0% hit rate is merely neutral, not costly
- Expecting a longer retention window to rescue a varying prefix
- Rounding a timestamp to the second and expecting hits