skip to content

Besides TTL expiry, what can drop a cached prompt prefix before its window ends?

level: middleimportance: should knowfreq 38%

answer

  1. a ceiling, not a promise
  2. best-effort optimisation, not a store
  3. capacity pressure and routing both bite
  4. entries are private to your account
  5. engineer the miss path, not the hit

basics

~20 s

Prompt caching is best-effort. Entries can be dropped early under capacity pressure, a request can land on serving capacity that never held the entry, and caches are scoped to your own account, so an identical prompt from another customer never helps you.

solid answer

~50 s

The TTL is an upper bound, not a promise. Providers describe prompt caching as a **best-effort optimisation**, which means an entry can disappear before its window elapses: capacity pressure can evict it, and your request may be routed to serving capacity that never held it, especially across regions or after a provider-side rebalance. Scope is the other surprise — entries are private to the account or organisation that wrote them, so two customers sending byte-identical system prompts get no benefit from each other, and in some setups even separate keys or deployments within one account behave as separate caches. The design consequence is simple and worth saying out loud in an interview: never let correctness depend on a hit, never build a latency SLO that assumes one, and size cold-path timeouts for the full uncached processing time. Treat the hit rate as a cost-reduction metric you monitor, not an invariant you rely on.

go deeper

for a junior

Know that a cache hit is never guaranteed — the same request can be fast one minute and full-price the next — and that the answer itself does not change when the cache misses.

for a middle

Be ready to name the non-TTL causes: eviction under capacity pressure, routing to capacity that never held the entry, and per-account scoping that stops identical prompts from other customers helping you.

for a senior

Show that you engineer the miss path: cold-path timeouts, bounded concurrency for cold bursts, SLOs achievable without a hit, and the ability to read an unexplained write in the usage split without over-reacting to it.

for a principal

Own the isolation argument — why per-account scoping is a side-channel defence and not a missing feature — and the platform tradeoff when sharding across keys or regions multiplies the number of prefixes you must keep warm.

## Best-effort by design A prompt cache is a performance optimisation layered onto inference infrastructure, not a durable store with a contract. Providers state this plainly: an entry *typically* survives its TTL, and the TTL bounds how long it *may* live — it does not guarantee how long it *will*. Reasoning about it as a guarantee is the root of most production surprises here. ## Ways an entry disappears early **Capacity pressure.** The stored state consumes real memory on serving hardware. Under load, or during maintenance and rebalancing, entries can be reclaimed before their window expires. You get no notification; the next request is simply a write. **Routing.** Your request has to reach capacity that holds the entry. Fleet changes, deployments on the provider side, regional failover, or a differently routed request can all land you somewhere that has never seen your prefix. This is invisible from the outside and shows up only as an unexplained write in your usage accounting. **Content divergence.** Not eviction strictly, but the same practical outcome: if anything ahead of or inside the cached region differs from what was stored — a changed model version, an edited instruction, a reordered tool list — the lookup misses even though a perfectly good entry still exists for the old content. ## Scope: caches are private Entries are keyed to the account or organisation that created them. Two unrelated customers sending byte-identical system prompts do not share anything, and popular public prompts are not warmed globally on your behalf. This is a privacy and isolation property first — a shared cache would leak information about what other customers send, and timing differences alone could reveal it — and only incidentally a performance limitation. Within a single organisation, separate keys, projects or regional deployments may also behave as distinct cache domains, which is worth verifying rather than assuming when you fan traffic across several of them. A practical corollary: if you shard the same workload across multiple keys or regions for rate-limit headroom, you have multiplied the number of prefixes that must be independently warmed. That can quietly halve an expected hit rate. ## Designing for the miss Because a miss is always possible, the miss path is the path you must engineer for. **Timeouts.** Size client and gateway timeouts against uncached processing of the full prefix. A timeout tuned to warm-cache first-token latency will fire spuriously the first time an entry is gone, converting a cost event into an availability event — and the retry then pays the write again. **Capacity.** If a burst of traffic arrives cold — after a deploy, a failover, or an eviction — every request reprocesses the same large prefix concurrently. Bound concurrency so a cold start degrades throughput gracefully instead of saturating your rate limits. **SLOs.** Publish latency targets your system can meet without a hit, and treat the hit as headroom. Teams that set a p95 from warm-cache measurements discover on the first eviction that the target was never achievable. **Correctness.** Nothing about the answer changes on a miss — the full prompt is still sent and processed — so caching must never be load-bearing for behaviour. If some content only reaches the model 'because it is cached', the design is already wrong. ## How you would know The only observable is the per-response usage split between cache-write and cache-read tokens. A shared, stable prefix under steady traffic should show writes as a low, roughly flat count with reads scaling with volume. Sporadic unexplained writes that do not line up with deploys or with gaps in traffic are the fingerprint of early eviction or routing changes — worth accepting as background noise rather than chasing, provided the rate stays low. If it does not, the honest conclusion is usually that the workload's shape, not the provider, is the problem.

  • Why is it a security property, not just a performance one, that caches are not shared across accounts?
    A shared cache would let one customer learn what another has sent: a suspiciously fast first token for a guessed prefix reveals that someone else recently sent it. That is a practical side channel over proprietary prompts and confidential documents. Per-account scoping removes the channel, at the cost of every tenant warming its own copy.
  • How should a client timeout be set on a route that usually gets cache hits?
    Against the uncached case. The cached path is faster, but any request can miss, and a timeout tuned to warm first-token latency turns a routine miss into a failed request plus a retry that pays the write a second time. Set the timeout from cold-path measurements and treat the hit as headroom rather than as the budget.
  • You see occasional cache writes that match no deploy and no gap in traffic. Is that a bug?
    Usually not. Early eviction under capacity pressure and routing to capacity that never held the entry both produce exactly this signature, and neither is visible to you. Treat a low, steady background rate as expected. Investigate only when the write rate is high enough to move cost materially, at which point the likely cause is prefix variation rather than the provider.

saying these in an interview costs you the question

  • Treats the TTL as a guarantee that an entry survives that long
  • Builds a p95 latency SLO from warm-cache measurements only
  • Expects an identical public prompt to be warm from other customers' traffic
  • Sets client timeouts from cached first-token latency
  • Assumes sharding across keys or regions preserves one shared cache

context