skip to content

How do you measure the savings prompt caching actually delivered in production?

level: seniorimportance: nice to knowfreq 38%

answer

  1. measure, do not assume
  2. the response tells you the counts
  3. creation versus read, per class
  4. discount collected minus premium paid
  5. report against total spend, not prefix

basics

~20 s

Use the per-request token counts providers report, which separate cache-creation from cache-read from ordinary input. Convert each to base-input-token equivalents with its price multiplier, compare against what the same traffic would have cost uncached, and track the result per request class.

solid answer

~50 s

Do not infer savings from the invoice — instrument the responses. Providers return per-request usage that distinguishes **cache-creation tokens**, **cache-read tokens** and ordinary uncached input. Convert to base-input-token equivalents: `creation·w + read·r + uncached`, versus the counterfactual `creation + read + uncached` had nothing been cached. The difference, summed over a window, is realized saving; expressing it as a percentage of *total* spend rather than of the prefix keeps the number honest, since output tokens are never discounted. Two derived signals matter more than the headline. The **read-to-creation ratio** tells you whether writes are being repaid at all — healthy caching shows reads dominating. And the numbers must be sliced **per request class or per prefix family**, because one pathological path that writes constantly and never reads is easily hidden inside a healthy aggregate. Pair the cost series with a time-to-first-token series, since the two benefits move independently.

go deeper

for a junior

Know that responses report how many tokens were written to and read from the cache, and that those counts — not the bill — are how you tell whether caching is working.

for a middle

Be able to write the saving as discount collected on reads minus premium paid on writes, in base-token equivalents, and explain why the counterfactual is that all those tokens would have been ordinary input.

for a senior

Show you would slice the counters by request class, alert on the read-to-creation ratio rather than absolute cost, and track time-to-first-token as a separate series segmented by hit.

for a principal

Own how caching efficiency is reported to the business: savings as a share of total spend, a dimensionless health metric engineering can alert on, and a clear statement that cost and latency are separate outcomes.

## Why the invoice is the wrong instrument A monthly bill is a lagging, aggregated scalar. It moves for many reasons at once — traffic growth, a model change, longer answers, a new feature — so attributing a delta to caching from the invoice alone is guesswork. The measurement has to come from the **per-request usage data the provider returns with every response**, which is both immediate and attributable. ## The three counters Responses report input token counts split into classes: tokens written into the cache (**cache creation**), tokens served from an existing entry (**cache read**), and ordinary uncached input. Output tokens are reported separately and are irrelevant here, because caching never discounts them. ## Turning counters into money Work in **base-input-token equivalents** — multiples of the standard input price — so the arithmetic is provider-agnostic and survives a price change. With write multiplier `w` (about 1.25 for a short-retention tier) and read multiplier `r` (about 0.1): - **Actual cached-path input cost** = `creation·w + read·r + uncached` - **Counterfactual uncached cost** = `creation + read + uncached` (every one of those tokens would simply have been ordinary input) - **Realized saving** = `read·(1 − r) − creation·(w − 1)` That last expression is the one worth memorizing: **the discount collected on reads minus the premium paid on writes.** It goes negative exactly when writes outweigh reads, which is the regression case. Multiply by the model's input price to express in currency, and sum over the window. ## Report it against total spend The most common reporting error is quoting the prefix-level number. Heavy reuse converges on roughly a 90% saving *on the cached prefix*, but a request also carries variable input and generated output at full price. If those are comparable in size, the end-to-end reduction may be a third of the headline. Publish the saving as a share of **total LLM spend for that path**, and keep the prefix-level figure as a diagnostic rather than the KPI. ## The ratio is the health metric The raw saving figure is what finance wants; the **read-to-creation token ratio** is what engineering should alert on. Healthy caching shows read tokens far exceeding creation tokens — each write amortised across many reads. A ratio near or below one means writes are barely being repaid. Because the ratio is dimensionless it is comparable across models, prompt sizes and traffic levels, which makes it a much better alert threshold than absolute cost. ## Slice it, never average it A single global figure is actively misleading. One high-volume path with unique per-request inputs writing constantly and reading nothing can sit inside an otherwise healthy aggregate and look fine. Break the counters down by **request class, prompt template or prefix family** — whatever dimension your traffic actually varies on — and alert per slice. The same discipline catches a regression introduced when someone adds a varying element to a previously stable block: the slice's ratio falls off a cliff while the global number barely twitches. ## Measure the latency benefit separately Cost and latency are independent outcomes with independent conditions, so give latency its own series: **time-to-first-token**, segmented by whether the request reported cache reads. That segmentation is the cleanest possible demonstration of the benefit — same endpoint, same model, TTFT distribution with and without a hit. Do not fold it into an average end-to-end latency, where streaming time swamps the effect on long answers. ## A worked sanity check Suppose a service reports, over a day: 400M cache-read tokens, 8M cache-creation tokens, 100M uncached input. Realized saving = 400M × 0.9 − 8M × 0.25 = 360M − 2M = **358M base-token equivalents**, against a counterfactual input total of 508M — about 70% off the input bill, before accounting for output tokens. The read-to-creation ratio of 50:1 confirms the prefixes are genuinely hot. If the same day instead showed 8M reads against 400M creations, the formula returns roughly −93M: a large, silent regression that no dashboard tracking only "cache enabled: yes" would ever surface. ## What the interviewer is listening for That you measure rather than assume; that you know the counters exist and what they mean; that you convert to a common unit instead of hand-waving; that you report against total spend; and that you slice by request class. Candidates who answer "we saw the bill go down" have not built the observability the question is about.

  • What single ratio would you alert on to catch caching turning net negative?
    Cache-read tokens divided by cache-creation tokens, computed per request class. Healthy caching shows reads dominating writes by a wide margin; a ratio approaching or below one means writes are not being repaid. It is dimensionless, so one threshold works across models, prompt sizes and traffic levels.
  • Your dashboard shows a 90% cache saving but finance sees a 30% reduction. Who is wrong?
    Neither — they are measuring different denominators. The 90% is the discount on the reused prefix; the invoice also carries variable input and output tokens, which caching never discounts. Publish savings as a share of total spend for the path and keep the prefix-level discount as an internal diagnostic.
  • How would you demonstrate the latency benefit convincingly rather than anecdotally?
    Segment time-to-first-token by whether the response reported cache-read tokens, on the same endpoint and model, and compare the distributions rather than the means. Keep it separate from end-to-end latency, where streaming a long answer swamps the effect and hides a real improvement.

saying these in an interview costs you the question

  • Judging caching's effect from the monthly invoice alone
  • Quoting the prefix-level discount as the overall saving
  • Averaging cache metrics across all request classes
  • Ignoring cache-creation tokens when computing savings
  • Folding time-to-first-token into an end-to-end latency average

context