In prompt caching, how many cache reads repay one cache write?
answer
- premium now, discount every time after
- work in base-token equivalents
- premium above base over discount below base
- break-even is one or two reuses
- then ask whether reuse arrives in time
basics
~20 sBreak-even reads equal the write premium above base divided by the discount below base: with a 1.25x write and a 0.1x read, a single reuse already pays; at a 2x write it takes about two. The real constraint is whether that reuse arrives before the entry expires.
solid answer
~50 sModel it as two totals in base-input-token equivalents. Uncached, N+1 requests over the same prefix cost N+1. Cached, they cost one write at premium w plus N reads at discount r, so `w + N·r`. Setting them equal gives **N = (w − 1) / (1 − r)**. With a typical 1.25x write and 0.1x read that is 0.25/0.9 ≈ 0.28 — the *first* reuse already puts you ahead. With a 2x write premium for a longer-retention tier it is about 1.1, so you need roughly two reads. The startling conclusion is that break-even is one or two reuses, not dozens; caching pays for almost any repeated prefix. The binding constraint is therefore not the arithmetic but *arrival rate*: those reads must land while the entry still exists, or every request pays the premium and never collects the discount.
go deeper
Know the shape: you pay a bit extra to store the prefix once, then much less each time it is reused. Be able to say that only a couple of reuses are needed before caching is cheaper.
Derive it. Write the cached total as one write at the premium plus N reads at the discount, set it against N+1 uncached, and solve for N. Then plug in realistic multipliers and state the answer as a number.
Move past the algebra to the operating condition: request rate against each distinct prefix versus entry lifetime. Show you would size the saving against total spend rather than quoting the prefix-only discount.
Own the portfolio view — which request classes have prefixes large and hot enough to matter, what the realistic invoice-level reduction is, and when the engineering effort is better spent on model routing or output-length control than on caching.
## Why interviewers ask this The break-even calculation is the whole leaf in one question. It is easy to state, easy to get wrong by an order of magnitude, and it separates candidates who have actually reasoned about an LLM bill from those who have only read that "caching saves money". ## Setting up the arithmetic Work in **base-input-token equivalents** — multiples of the model's standard input price — so provider-specific dollar figures drop out. Define: - `w` = the cache-write multiplier (what a stored token costs relative to base), typically about 1.25 for a short-retention entry and around 2 where a provider sells a longer-lived tier. - `r` = the cache-read multiplier, typically about 0.1. - `P` = the size of the shared prefix in tokens. It cancels out, which is why the answer is a *count of reuses*, not a token threshold. One write plus N reads costs `P·(w + N·r)`. The same traffic uncached costs `P·(1 + N)`. Caching is cheaper when `w + N·r < 1 + N`, which rearranges to: **N > (w − 1) / (1 − r)** ## Plugging in real numbers - w = 1.25, r = 0.1 → N > 0.25 / 0.9 ≈ **0.28**. One reuse is already enough; the second request over the prefix is where you go net positive. - w = 2.0, r = 0.1 → N > 1 / 0.9 ≈ **1.11**. Two reuses. - A provider with no write premium at all (w = 1) → N > 0. Any reuse whatsoever wins, and there is nothing to lose by enabling it. This is the number to have in your head: **break-even is one to two reuses.** Candidates routinely guess ten or a hundred, which badly misjudges when caching is worth wiring in. ## The asymptote — what heavy reuse is actually worth As N grows, the cached cost per request tends to `r` and the saving tends to `1 − r`, i.e. around **90%** of the prefix's input cost. Take a nightly classification job over 10,000 support emails, each sent with the same 20k-token taxonomy prompt. Uncached, the prefix alone is 10,000 × 20k = 200M input tokens at base. Cached, it is one write of 20k at 1.25x (25k equivalent) plus 9,999 reads at 0.1x (about 20M equivalent) — roughly a **90% reduction** on that portion of the bill. Note carefully what that figure covers: only the shared prefix. The per-email text and the model's classification output are still billed at full rate, so the *end-to-end* saving is always smaller than 90% and depends on how much of each request is prefix. ## Where the arithmetic stops being the constraint Because break-even is barely above one reuse, the interesting question is not "is this prefix reused enough?" but "**does the reuse arrive in time?**" Cache entries do not live indefinitely. If a service sends one request per hour against a prefix whose entries survive minutes, every single request writes a fresh entry and none of them ever reads one — so the workload pays the write premium perpetually and collects nothing. The effective cost is `w` per request, about 25% *worse* than not caching at all. That converts the break-even into a **traffic-rate condition**: your request rate against a given prefix must comfortably exceed one request per entry lifetime. Bursty traffic — a batch job, a working session, a spike — satisfies this easily even at low daily volume, because the reuses are clustered. Steady low-rate traffic spread thinly across many distinct prefixes does not, even if the daily total looks large. ## Prefix size and the fixed floor One more practical wrinkle: providers usually enforce a **minimum cacheable prefix length**, below which nothing is cached and the machinery is a no-op. Small prefixes therefore fail not on economics but on eligibility. Combined with the break-even result, the decision rule becomes pleasingly simple: *if the prefix is big enough to be cacheable and will be sent again soon, cache it.* ## How to present this in an interview Derive the formula rather than reciting a number, state the plugged-in result (one to two reuses), then immediately move to the condition that actually decides real systems — arrival rate versus entry lifetime — and to the caveat that the ~90% asymptote applies to the prefix only, not to the whole invoice. That progression from arithmetic to operating reality is what a middle-to-senior answer looks like.
- If break-even is barely one reuse, why not cache every prompt by default?Because a prefix that is never reused before it expires costs strictly more — you pay the premium and collect nothing — and workloads with high-variance prompts or very low per-prefix request rates fall exactly there. Default-on is right for large stable prefixes under real traffic; single-shot or highly variable prompts should stay uncached.
- Does the size of the shared prefix change the break-even reuse count?No. Prefix size multiplies both the cached and the uncached total identically, so it cancels out of the inequality — break-even is a count of reuses, not a token threshold. Size does change the *magnitude* of the saving, and providers impose a minimum cacheable length below which caching does not engage at all.
- Your batch job saves 90% on the cached prefix. Why is the invoice down only 40%?The 90% figure applies to the shared prefix's input cost alone. Each request also carries variable input at full price and generated output, which caching never discounts. If prefix, variable input and output are comparable in size, the end-to-end reduction lands well below the headline number — always compute savings against total spend, not against the prefix.
saying these in an interview costs you the question
- Guessing you need dozens or hundreds of reuses to break even
- Applying the 90% read discount to the entire request cost
- Ignoring that reuse must arrive before the entry expires
- Assuming a bigger prefix lowers the break-even reuse count
- Forgetting that unread writes are billed at above-normal price