Does a prompt cache's TTL reset on each cache hit, or expire a fixed time after the write?
answer
- think idle timeout, not lease
- what event restarts the clock?
- measured from last use, not write
- every read refreshes the full window
- only a longer-than-window gap kills it
basics
~20 sPrompt caches typically use a sliding window: every read of a cached prefix restarts its expiry at the full TTL. A continuously used prefix stays warm indefinitely, and only an idle gap longer than the window forces a fresh, full-price write.
solid answer
~50 sIn the caching schemes providers expose as of mid-2026, the TTL is measured from the **last use**, not from the write. A prefix written at 12:00 under a five-minute window and read again at 12:04 now lives until 12:09, not 12:05. That changes what you optimise for: the number that decides your hit rate is the **largest gap between consecutive requests sharing that prefix**, not session length or total elapsed time. Only a genuine cache read refreshes the clock — traffic to the same model with a different prefix does nothing, and nothing on your own infrastructure touches the provider's entry. Expiry is also silent rather than an error: the request still succeeds, but the prefix is reprocessed and rewritten, so you lose the discount and the first-token speedup, visible only as write-tokens instead of read-tokens in the usage accounting.
go deeper
Know that a cached prefix has a limited lifetime and that using it again extends that lifetime. Say plainly that an expired cache is not an error — the call still works, it just costs the full price again.
Be ready to explain sliding expiry precisely with a timeline, name a cache read as the only thing that refreshes the clock, and state that hit rate is governed by the gap between requests sharing the prefix, not by session length.
Show that you instrument cache-write versus cache-read tokens per route, size cold-path timeouts for the uncached case, and never let a latency SLO assume a hit. Explain why an expiry mid-conversation costs more than one at the start.
Own the traffic-shape argument: which TTL tier a workload justifies, whether the gap distribution supports caching at all, and what you would give up to make prefixes shared rather than per-user so one warm entry serves many callers.
## What a TTL is bounding here When you mark part of a prompt as cacheable, the provider stores the model's precomputed internal state for that exact token prefix and keeps it available for a bounded time. Writing the entry is the expensive step; reading it back is the cheap, fast one. TTL policy therefore answers one question: how long do I get the cheap step before I have to pay the expensive one again? ## Sliding expiry, not fixed-from-write The common design as of mid-2026 is a **sliding window**. Expiry is measured from the most recent use. Each hit restarts the clock at the full TTL. Concretely, with a five-minute window: write at 12:00, read at 12:04, and the entry now survives until 12:09. Read it again at 12:08 and it survives to 12:13. There is no cumulative budget being spent down — only an idle timer being reset. Two consequences follow, and they are the whole practical content of this topic. First, an actively used prefix can stay resident **indefinitely** without ever being rewritten. A busy support assistant sharing one system prompt across thousands of sessions may write its prefix once at deploy and never again, because natural traffic never leaves a gap longer than the window. Second, the quantity that governs your hit rate is the **distribution of gaps between consecutive requests that share the prefix** — not average traffic volume, not how long a session lasts. A service with high daily volume but per-user prefixes and long think-times can have a near-zero hit rate, while a low-volume service with one shared prefix and steady trickle traffic can sit at nearly 100%. ## What actually refreshes the clock Only a request that genuinely matches the stored prefix refreshes it. Things that do **not** refresh it: traffic to the same model with a different prefix; requests from another account; keep-alive on your own HTTP connections; health checks against your own service. If a request matches only a shorter portion of a longer cached prefix, it refreshes the portion it matched. ## Expiry is silent, not fatal A request whose prefix has expired does not fail. The provider reprocesses those tokens and writes a fresh entry, so the call succeeds — just at full input cost, plus whatever write premium the tier carries, and without the first-token latency benefit. The only signal is in the per-response usage accounting, where tokens land in the cache-write bucket instead of the cache-read bucket. A team that never inspects those fields experiences an expiry regression purely as a cost line that is higher than the prototype implied and a p95 that quietly drifted up. ## Tiers and the multi-turn rhythm Providers commonly offer a short default window measured in minutes plus an optional longer window on the order of an hour, at a higher write premium. Which tier fits is a traffic-shape decision — what fraction of follow-ups arrive inside each candidate window — rather than anything about the prompt itself. In a multi-turn conversation the prefix grows each turn. While turns keep arriving inside the window, each turn reads the existing prefix (refreshing it) and writes only the incremental new portion, which is the cheap steady state. Break the rhythm — the user goes to lunch — and the next turn hits the worst case: a full write of a now much *longer* prefix. The cost of a miss therefore grows over the life of a conversation, which is why long-running sessions with irregular pacing are the pattern that hurts most. ## Rules that follow - Design against the gap distribution, not the mean. - Never make correctness, and preferably not a hard latency SLO, depend on a hit; the window is a ceiling, not a guarantee. - Instrument cache-write versus cache-read tokens per route so a regression is visible the day it lands. - Remember that the first request after any idle period is always a write, so cold-path timeouts must be sized for the uncached case.
- If the window slides on every read, why does a long conversation ever pay a write after the first turn?Because each turn appends tokens, so the prefix grows. The existing portion is read and refreshed, but the newly appended portion has never been cached and must be written. In the steady state you pay only for that increment — which is why an expiry mid-conversation is so costly: the next request rewrites the entire accumulated prefix, not just the new part.
- Does reading a shorter prefix refresh a longer cached one built on top of it?It refreshes what it matched. Cache lookups are prefix matches, so a request that shares only the first portion keeps that portion alive; the longer continuation that the shorter request never touched keeps counting down from its own last use. Shared, stable material at the front of the prompt therefore stays warm far more reliably than any single user's full conversation.
- How would you detect that your effective TTL is shorter than you assume?Compare cache-write and cache-read token counts per route over time. A healthy shared prefix shows writes as a small constant and reads scaling with traffic; if writes track request volume roughly one-to-one, the window is expiring between requests. Pair that with the inter-request gap distribution for the prefix to confirm the cause is timing rather than prefix variation.
saying these in an interview costs you the question
- Says the entry dies a fixed TTL after the write regardless of use
- Thinks an expired cache makes the request fail rather than silently rewrite
- Believes traffic to the same model keeps any prefix warm
- Assumes the clock starts at the first request of the session
- Treats a cache hit as guaranteed once the entry exists