skip to content

When does enabling prompt caching make a workload more expensive?

level: seniorimportance: should knowfreq 50%

answer

  1. you pay before you profit
  2. a write nobody reads is pure loss
  3. unique inputs, no second call
  4. per-prefix arrival rate, not total volume
  5. loss is bounded by the premium

basics

~20 s

Whenever a written prefix is never reused before it expires. Single-shot calls over unique documents, prompts whose stable portion changes every request, and traffic arriving slower than entries survive all pay the write premium and collect no discount — typically about 25% worse than not caching.

solid answer

~50 s

Caching is net negative exactly when writes outnumber useful reads. Three patterns produce that. **One-shot work**: summarizing a single 8k-token document once means paying roughly a 25% surcharge on those tokens for a reuse that never comes, with no latency gain either since the prefix still has to be read the first time. **High-variance prefixes**: if per-user or per-request content sits inside the supposedly stable block, each request writes a distinct entry and reads nothing. **Sub-lifetime traffic**: a prefix requested less often than entries survive expires unused every time, so the workload pays the premium perpetually. The damage is bounded — you lose the premium, not a multiple — but it is a real, silent regression, and the tell is a cache-creation token count that stays comparable to or larger than the cache-read count. The remedy for the first two is to turn caching off for that request class rather than to tune it.

go deeper

for a junior

Know that caching only pays when the same prompt text is sent again soon. If every request is over a different document, caching adds cost instead of saving it.

for a middle

Explain the three loss patterns — one-shot calls, prefixes that secretly vary, and traffic slower than the entry lifetime — and that the loss equals the write premium on the cached tokens.

for a senior

Show you would scope caching per request class and verify with reported cache-creation versus cache-read counts, broken down rather than averaged. Be able to size the regression and say when the latency gain still justifies it.

for a principal

Own the default policy across a product's request classes: where caching is on by design, what the bounded downside is worth against the potential saving, and what monitoring catches a silent regression before it shows up in a quarterly bill.

## The failure condition in one line Caching costs money up front and refunds it on reuse. So the loss condition is simply: **a write that is never read.** Everything below is a way of arriving at that state. ## Case 1 — the genuine one-shot A document-processing service summarizes each uploaded file exactly once. A user uploads an 8k-token contract; the service sends it with a short instruction and gets a summary back. If caching is switched on for that call, those 8k tokens are billed as a cache write at roughly 1.25x base instead of 1x — about a **25% surcharge** on the input for that request — and nothing ever reads the entry, because the next upload is a different contract. The latency picture is equally unhelpful. The first request must read the prompt regardless; caching only helps a *subsequent* request that reuses it. So a one-shot call gets neither benefit and pays the premium. This is the cleanest example of caching being straightforwardly wrong, and it is common in practice because caching often gets enabled globally in a shared client wrapper rather than per request class. ## Case 2 — a "stable" prefix that is not stable More insidious: the prefix looks fixed but carries something that varies. A tenant name, an account tier, a timestamp, a user's display name, a freshly shuffled example set — any of these embedded in the block makes every request a distinct prefix. From the billing system's perspective this is a firehose of writes with essentially no reads. This is worth flagging precisely because it *looks* like it should be working; the code says caching is on, and the bill goes up rather than down. The diagnosis lives in the reported token counts (below); the correctness question of what belongs in the stable block belongs to prompt structure, not to this cost analysis. ## Case 3 — traffic slower than the entry lives Even with a perfectly stable prefix and genuine repetition, timing can kill it. Cache entries have a finite lifetime measured in minutes for the common tier. A service that touches a given prefix once an hour finds no live entry every time: write, expire unused, write again. Cost settles at the write multiplier — around 25% above baseline forever. The subtlety is that **aggregate volume does not save you**; per-prefix arrival rate does. A system doing a million calls a day spread across 50,000 distinct customer-specific prefixes may have almost no prefix hot enough to hit. Conversely a low-volume system whose traffic is bursty — a batch run, a user's working session — hits constantly, because the reuses cluster in time. Always ask about arrival rate *per distinct prefix*, never total QPS. ## How much can you actually lose? Be honest about the magnitude: the downside is **bounded by the write premium**, roughly 25% of the input cost of the cached portion in the common tier, or up to about 100% on a longer-retention tier. You cannot lose an order of magnitude, and only the cached portion of the input is affected — variable input and output are unchanged. That asymmetry against a potential ~90% saving is why caching is worth trying by default on plausibly-repeated prefixes: the upside dwarfs the downside. But "bounded" is not "free", and on a high-volume misconfigured path it is a meaningful and completely silent regression. ## Detecting it Providers report per-request token counts that distinguish cache-creation from cache-read. The health signal is the **ratio**: healthy caching shows read tokens vastly exceeding creation tokens. Creation persistently at or above reads means you are funding writes that nothing collects. Track it per request class, because a single global average will hide one pathological path behind several healthy ones. ## What to do about it The correct response depends on the case. For the genuine one-shot, **turn caching off** for that request class — there is nothing to tune. For traffic below the entry lifetime, either accept the premium as the price of the latency benefit (if a human is waiting and you value fast first tokens more than the surcharge), or stop caching that path. For the unstable-prefix case, the prompt itself needs restructuring so the fixed part really is fixed. ## The senior framing What an interviewer is listening for is that you treat caching as a **per-request-class decision with a measurable downside**, not a global switch that is always on. The judgement is: enable where a large prefix is plausibly reused within the entry lifetime, disable where inputs are unique, and verify with the reported token counts rather than assuming.

  • Given the downside is bounded at roughly 25%, why not just cache everything and stop thinking about it?
    On plausibly-repeated prefixes that is defensible — the asymmetry between a ~25% loss and a ~90% saving favours trying. But on a high-volume path of genuinely unique inputs it is a permanent, silent 25% surcharge on a large bill for zero benefit, and it also fills entries that nothing will read. Decide per request class and verify with the reported token counts.
  • A team caches per-customer prefixes and sees the bill rise. Total traffic is high. What do you check first?
    Per-prefix arrival rate, not total QPS. High aggregate volume spread across thousands of customer-specific prefixes can mean no individual prefix is touched again before its entry expires, so every request writes and none reads. Compare cache-creation to cache-read tokens broken down by prefix or customer, not in aggregate.
  • Is there ever a reason to keep caching on for a workload where it loses money?
    Yes, when you are buying latency rather than cost. If a human waits on a very large prompt and the reuse pattern is marginal, paying the write premium can still be worth a sharply lower time-to-first-token on the requests that do hit. Make that an explicit trade, sized against the premium, not an accident.

saying these in an interview costs you the question

  • Treating caching as a global switch that can never hurt
  • Assuming an unread cache write costs nothing
  • Judging cacheability by total request volume rather than per-prefix rate
  • Believing caching speeds up the very first call over a document
  • Leaving caching on for one-shot unique-document workloads

context