How do you measure prompt-cache hit rate in a production LLM service?
answer
- count tokens, not requests
- two counters in the usage block
- read tokens versus creation tokens
- slice by route, version and deploy
- log a hash of the prefix
basics
~20 sMeasure it in tokens, not requests. Providers report per-response counters for tokens read from cache and tokens written to it, so hit rate is cache-read tokens over the reusable prefix. Chart it per route and per deploy.
solid answer
~50 sAs of mid-2026, providers return per-response usage counters that separate **tokens served from cache** from **tokens written into the cache** and from ordinary uncached input. Hit rate is a token ratio, not a request ratio: a request that reuses 2k of a 30k prefix is mostly a miss even though it "hit". So I log all three counters on every call, plus a hash of the cacheable prefix I sent, and I aggregate hit rate as cache-read tokens divided by the prefix tokens I expected to reuse. That gets sliced by route, prompt version and deploy. Two charts do most of the diagnostic work: cache-read versus cache-creation tokens over time (a healthy system reads far more than it writes once traffic is steady), and the count of *distinct* prefix hashes per hour. A route whose creation tokens track its read tokens one-for-one is rewriting the cache on every call, and the hash cardinality tells you whether the cause is prompt variation.
code
python · 9 lines# per-response counters reported in the provider's usage block
cache_read, cache_write, uncached = 28_400, 0, 260
prefix_total = cache_read + cache_write
hit_rate = cache_read / prefix_total if prefix_total else 0.0
coverage = cache_read / (prefix_total + uncached)
print(f"prefix hit rate: {hit_rate:.1%}")
print(f"share of all input served from cache: {coverage:.1%}")go deeper
Know that the provider returns separate token counts for cache reads and cache writes, and that you read those numbers rather than guessing from latency. Be able to say what a high write-to-read ratio implies.
Explain why hit rate must be token-weighted: matching is prefix-based, so a partial match is a partial hit. Be ready to compute the ratio from the three counters and to name the two charts you would build.
Show the diagnosis workflow — chart read versus creation tokens across a day, correlate the inflection with a deploy, then count distinct prefix hashes to separate prompt variation from routing or expiry. Say what each branch rules out.
Own the observability contract itself: which counters and hashes every LLM call site must emit, how they are sliced, and what regression threshold pages someone. Argue where the metric misleads — high hit rate on a prefix so generic that output quality suffers is not a win.
## What a prompt cache actually reports A prompt cache lets the provider skip re-processing a prefix of your prompt that it has already seen. To make that observable, providers report per-response token counters that split input into three buckets, and as of mid-2026 essentially every major provider exposes some version of this in the response's usage block: - **Cache-read tokens** — prefix tokens that were served from an existing cache entry. This is the win. - **Cache-creation (write) tokens** — prefix tokens that were processed fresh *and* stored for reuse. This is the investment. - **Ordinary uncached input tokens** — the variable tail of the prompt, processed normally every time. The key discipline is that these are *token* counts. A request is not a boolean hit or miss. ## Hit rate is a token ratio Define hit rate as cache-read tokens divided by the tokens you *intended* to be reusable — typically cache-read plus cache-creation. Counting requests instead hides the failure mode that matters: a 30k-token system prefix where the cache matches only the first 2k reports a "hit" while 28k tokens are being reprocessed every call. In token terms that is a 7% hit rate, which is roughly worthless. Request-level hit rate would call it 100%. A second useful number is *effective coverage*: cache-read tokens divided by total input tokens. That tells you how much of the whole prompt, tail included, you are getting for free, which is the number that maps to what you are paying and to how much prefill work remains before the first token appears. ## Instrument at the call site The three counters alone tell you *that* you are missing, not *why*. The cheapest addition is to hash the cacheable portion of the prompt before you send it — the system text, tool definitions, any pinned documents — and log the first few hex characters of that hash alongside the counters. Now you can group by hash. If a route that should have exactly one prefix shape is producing thousands of distinct hashes per hour, the prompt is varying and you have located the problem without reading any provider internals. If the cardinality is one and hits are still absent, the cause is elsewhere: traffic too sparse to keep an entry alive, requests landing on different serving replicas, or a model or parameter change that invalidated the entry. Also log the prompt version or template identifier, the route, and the deploy SHA. Cache behaviour is overwhelmingly a *deployment* property — one merged line at the top of a system prompt flips a fleet from 90% to 0% — so the metric is only actionable when it can be correlated with releases. ## The repeatable diagnosis workflow When a service reports near-zero hits, work in this order: 1. **Chart cache-read against cache-creation tokens per request across a day.** If creation tracks read one-for-one, every call is paying to write an entry nobody reads. If both are near zero, the request is not being marked as cacheable at all, or the prefix is below the provider's minimum cacheable length. 2. **Find the inflection point and line it up with deploys.** The flip is almost always a specific release, and the diff is usually small and obvious once you know which release to look at. 3. **Count distinct prefix hashes per route per hour.** High cardinality means content variation; cardinality of one points at timing, routing or invalidation instead. 4. **Diff two prefixes with different hashes byte-for-byte.** Not visually — programmatically. The offending difference is often invisible to a human: reordered keys, a changed float format, a trailing space. 5. **Check request timing.** Sparse traffic on a route can idle an entry out between calls; that shows up as hits during busy hours and misses overnight. ## What the metric cannot tell you Token counters do not tell you whether caching is *worth it* for a route — writes and reads are priced differently, and that economics question is separate. They also do not distinguish "the prefix changed" from "the entry expired", which is why prefix hashing is worth the few lines it costs. And a high hit rate is not automatically a healthy system: you can achieve 100% by caching a prefix so generic that quality suffers. ## Common instrumentation mistakes Inferring hits from latency is unreliable — perceived speed varies with load, output length and network. Reading only aggregate monthly spend hides per-route regressions inside a large bill. And sampling usage on 1% of traffic is usually too sparse to catch a regression on a low-volume but expensive route; these counters are cheap to log in full.
- Your dashboard shows cache-creation tokens roughly equal to cache-read tokens on every request. What is that telling you?That each call is paying to write a cache entry that the next call never reads, so you are getting the write cost with none of the reuse. Either the prefix differs slightly on every request, or requests are spread across serving replicas or accounts that do not share the entry, or traffic is too sparse for an entry to survive until the next call. Check distinct prefix-hash cardinality first — it separates content variation from the other two causes immediately.
- Why is a per-request boolean hit rate a misleading metric for a long system prefix?Because a cache match is a prefix match, not an all-or-nothing lookup. A request whose first 2,000 tokens match a 30,000-token prefix counts as a hit while 28,000 tokens are reprocessed. Boolean hit rate reports 100% for a system that is capturing 7% of the available saving. Token-weighted hit rate exposes the gap and points straight at where in the prompt the divergence begins.
- What would you log alongside the token counters to make misses diagnosable rather than just visible?A short hash of the exact cacheable prefix bytes, plus the prompt-template version, route name and deploy SHA. Grouping by hash turns an opaque "hits are down" into either high cardinality — the prompt is varying, and you can diff two examples — or cardinality of one, which redirects you to timing, routing or invalidation. Without the hash you are guessing at which of those it is.
saying these in an interview costs you the question
- Concluding a cache hit because the response felt faster
- Measuring hit rate per request instead of per token
- Watching only total spend, never per-route usage counters
- Assuming identical-looking prompt text guarantees a hit
- Sampling usage data so sparsely that regressions never surface