How do you measure the savings prompt caching actually delivered in production?
answer
- measure, do not assume
- the response tells you the counts
- creation versus read, per class
- discount collected minus premium paid
- report against total spend, not prefix
basics
~20 sUse the per-request token counts providers report, which separate cache-creation from cache-read from ordinary input. Convert each to base-input-token equivalents with its price multiplier, compare against what the same traffic would have cost uncached, and track the result per request class.
solid answer
~50 sDo not infer savings from the invoice — instrument the responses. Providers return per-request usage that distinguishes **cache-creation tokens**, **cache-read tokens** and ordinary uncached input. Convert to base-input-token equivalents: `creation·w + read·r + uncached`, versus the counterfactual `creation + read + uncached` had nothing been cached. The difference, summed over a window, is realized saving; expressing it as a percentage of *total* spend rather than of the prefix keeps the number honest, since output tokens are never discounted. Two derived signals matter more than the headline. The **read-to-creation ratio** tells you whether writes are being repaid at all — healthy caching shows reads dominating. And the numbers must be sliced **per request class or per prefix family**, because one pathological path that writes constantly and never reads is easily hidden inside a healthy aggregate. Pair the cost series with a time-to-first-token series, since the two benefits move independently.
go deeper
Know that responses report how many tokens were written to and read from the cache, and that those counts — not the bill — are how you tell whether caching is working.
Be able to write the saving as discount collected on reads minus premium paid on writes, in base-token equivalents, and explain why the counterfactual is that all those tokens would have been ordinary input.
Show you would slice the counters by request class, alert on the read-to-creation ratio rather than absolute cost, and track time-to-first-token as a separate series segmented by hit.
Own how caching efficiency is reported to the business: savings as a share of total spend, a dimensionless health metric engineering can alert on, and a clear statement that cost and latency are separate outcomes.
## Why the invoice is the wrong instrument A monthly bill is a lagging, aggregated scalar. It moves for many reasons at once — traffic growth, a model change, longer answers, a new feature — so attributing a delta to caching from the invoice alone is guesswork. The measurement has to come from the **per-request usage data the provider returns with every response**, which is both immediate and attributable. ## The three counters Responses report input token counts split into classes: tokens written into the cache (**cache creation**), tokens served from an existing entry (**cache read**), and ordinary uncached input. Output tokens are reported separately and are irrelevant here, because caching never discounts them. ## Turning counters into money Work in **base-input-token equivalents** — multiples of the standard input price — so the arithmetic is provider-agnostic and survives a price change. With write multiplier `w` (about 1.25 for a short-retention tier) and read multiplier `r` (about 0.1): - **Actual cached-path input cost** = `creation·w + read·r + uncached` - **Counterfactual uncached cost** = `creation + read + uncached` (every one of those tokens would simply have been ordinary input) - **Realized saving** = `read·(1 − r) − creation·(w − 1)` That last expression is the one worth memorizing: **the discount collected on reads minus the premium paid on writes.** It goes negative exactly when writes outweigh reads, which is the regression case. Multiply by the model's input price to express in currency, and sum over the window. ## Report it against total spend The most common reporting error is quoting the prefix-level number. Heavy reuse converges on roughly a 90% saving *on the cached prefix*, but a request also carries variable input and generated output at full price. If those are comparable in size, the end-to-end reduction may be a third of the headline. Publish the saving as a share of **total LLM spend for that path**, and keep the prefix-level figure as a diagnostic rather than the KPI. ## The ratio is the health metric The raw saving figure is what finance wants; the **read-to-creation token ratio** is what engineering should alert on. Healthy caching shows read tokens far exceeding creation tokens — each write amortised across many reads. A ratio near or below one means writes are barely being repaid. Because the ratio is dimensionless it is comparable across models, prompt sizes and traffic levels, which makes it a much better alert threshold than absolute cost. ## Slice it, never average it A single global figure is actively misleading. One high-volume path with unique per-request inputs writing constantly and reading nothing can sit inside an otherwise healthy aggregate and look fine. Break the counters down by **request class, prompt template or prefix family** — whatever dimension your traffic actually varies on — and alert per slice. The same discipline catches a regression introduced when someone adds a varying element to a previously stable block: the slice's ratio falls off a cliff while the global number barely twitches. ## Measure the latency benefit separately Cost and latency are independent outcomes with independent conditions, so give latency its own series: **time-to-first-token**, segmented by whether the request reported cache reads. That segmentation is the cleanest possible demonstration of the benefit — same endpoint, same model, TTFT distribution with and without a hit. Do not fold it into an average end-to-end latency, where streaming time swamps the effect on long answers. ## A worked sanity check Suppose a service reports, over a day: 400M cache-read tokens, 8M cache-creation tokens, 100M uncached input. Realized saving = 400M × 0.9 − 8M × 0.25 = 360M − 2M = **358M base-token equivalents**, against a counterfactual input total of 508M — about 70% off the input bill, before accounting for output tokens. The read-to-creation ratio of 50:1 confirms the prefixes are genuinely hot. If the same day instead showed 8M reads against 400M creations, the formula returns roughly −93M: a large, silent regression that no dashboard tracking only "cache enabled: yes" would ever surface. ## What the interviewer is listening for That you measure rather than assume; that you know the counters exist and what they mean; that you convert to a common unit instead of hand-waving; that you report against total spend; and that you slice by request class. Candidates who answer "we saw the bill go down" have not built the observability the question is about.
- What single ratio would you alert on to catch caching turning net negative?Cache-read tokens divided by cache-creation tokens, computed per request class. Healthy caching shows reads dominating writes by a wide margin; a ratio approaching or below one means writes are not being repaid. It is dimensionless, so one threshold works across models, prompt sizes and traffic levels.
- Your dashboard shows a 90% cache saving but finance sees a 30% reduction. Who is wrong?Neither — they are measuring different denominators. The 90% is the discount on the reused prefix; the invoice also carries variable input and output tokens, which caching never discounts. Publish savings as a share of total spend for the path and keep the prefix-level discount as an internal diagnostic.
- How would you demonstrate the latency benefit convincingly rather than anecdotally?Segment time-to-first-token by whether the response reported cache-read tokens, on the same endpoint and model, and compare the distributions rather than the means. Keep it separate from end-to-end latency, where streaming a long answer swamps the effect and hides a real improvement.
saying these in an interview costs you the question
- Judging caching's effect from the monthly invoice alone
- Quoting the prefix-level discount as the overall saving
- Averaging cache metrics across all request classes
- Ignoring cache-creation tokens when computing savings
- Folding time-to-first-token into an end-to-end latency average