Your Anthropic prompt cache never registers a read — how do you debug it?
answer
- it fails silently, no error
- read the two usage counters first
- too short, too variable, too slow
- tools sit at the front of the prefix
- cold parallel starts all miss
basics
~20 sCheck the usage counters first: persistent cache_creation with zero cache_read means the prefix is either below the minimum cacheable size, changing between calls, expiring between calls, or routed to a different model. None of these raises an error.
solid answer
~50 sPrompt caching fails silently, so diagnosis is a checklist against `usage`. If `cache_creation_input_tokens` is always zero, the marked prefix is below the minimum cacheable length — on the order of 1024 tokens for Sonnet- and Opus-class models and 2048 for Haiku-class — and `cache_control` was ignored without an error. If creation is non-zero but `cache_read_input_tokens` never is, the prefix differs between calls or the entry expires first. Hunt for injected variability ahead of the breakpoint: timestamps, user IDs, request IDs, non-deterministic JSON key ordering in tool schemas, a reordered tools array, a system prompt assembled from a set. Then check the gaps between calls against the five-minute sliding TTL, confirm both calls hit the same model, and remember that parallel cold-start requests all miss because the entry is not readable until the first one has processed it.
code
python · 29 linesimport os
from anthropic import Anthropic
client = Anthropic()
MODEL = os.environ["ANTHROPIC_MODEL"]
HANDBOOK = open("handbook.txt").read() # large and stable
def ask(question: str):
resp = client.messages.create(
model=MODEL,
max_tokens=512,
system=[
{"type": "text", "text": "You are a policy assistant."},
{
"type": "text",
"text": HANDBOOK,
"cache_control": {"type": "ephemeral"},
},
],
messages=[{"role": "user", "content": question}],
)
u = resp.usage
print(u.input_tokens, u.cache_creation_input_tokens, u.cache_read_input_tokens)
return resp
ask("What is the notice period?") # expect a write, no read
ask("What about probation?") # expect a read, no writego deeper
Know that caching can silently do nothing, and that the usage counters — not an error message — are how you find out whether a prefix was written or read.
Be able to list the concrete causes: prefix below the minimum size, dynamic text ahead of the breakpoint, unstable tool serialization, expiry, or a different model.
An interviewer wants a method, not a guess: reduce to a two-call reproduction, diff the serialized requests byte for byte, and reason about traffic gaps and cold-start concurrency in the real deployment.
Own the prevention: a single construction point for the cached prefix, a stability hash logged per request, and hit-rate alerting per route so a personalization field added by another team is caught in hours, not in the invoice.
## Start from the counters, not the code Prompt caching never throws. A misplaced or ineffective `cache_control` produces a perfectly normal response, so the only evidence is the `usage` object. Two distinct symptoms point at two different families of cause: - **Creation is always zero.** Nothing is being cached at all. The prefix is too small, or the marker is not where you think it is. - **Creation is non-zero every call, read is always zero.** Caching works, but no two requests share a prefix, or the entry dies before the next request arrives. ## Symptom 1: nothing is cached The most common cause is the **minimum cacheable prefix size**. Prefixes shorter than the threshold are not stored, and the API does not complain. The threshold depends on the model class — around 1024 tokens for Sonnet- and Opus-class models and around 2048 for the smaller Haiku class — so the same prompt can cache on one model and silently not on another. Count the tokens in your prefix before assuming a bug. The second cause is placement: the block you decorated is not actually inside the prefix you think, for example a `cache_control` on a message that sits after the content you wanted covered. Since the marker ends the prefix rather than selecting content, a marker placed too early caches far less than intended. ## Symptom 2: written every time, never read Now you are hunting variability ahead of the breakpoint, because matching is exact. **Injected dynamic text.** "Current date: …", a session or request ID, a user's name, an A/B flag, a random few-shot sample. Anything templated into the system prompt is suspect. **Unstable serialization.** Tool definitions built from a dict or set can serialize with different key or member ordering between processes. So can a JSON Schema assembled by reflection. The bytes must be identical, not merely equivalent. **A changed tool list.** Tools sit at the very front of the prefix, so adding, removing or reordering one invalidates everything behind it — including a system prompt you never touched. **Model or configuration drift.** Cache entries are keyed per model, so a routing layer that spreads traffic across model versions gives each its own entry. Toggling features that change how the prompt is assembled has the same effect. **Expiry.** The default window is about five minutes and refreshes on each hit. Low-QPS services with multi-minute gaps write and never read. Either raise the TTL to the one-hour tier or accept that caching does not pay here. **Cold-start concurrency.** An entry is not available for reads until the first request that writes it has processed the prefix. Deploying ten workers that all fire simultaneously produces ten writes and zero reads. Warm the cache with one request, then fan out. ## Reproducing it cleanly Reduce to a two-call script against a static prefix with no templating at all. Print `usage` from both calls. If call two reads, your production prefix is the variable; diff the serialized requests byte for byte — not the objects, the actual JSON — and the culprit usually appears immediately. If call two still writes, the prefix is too short or the marker is misplaced. ## Making it not recur Build the cached prefix once, in one place, from constants, and assert its stability: hash the serialized prefix and log the hash with the request. A hash that varies across requests that should share a cache is a one-line alarm, and it catches the failure the day someone adds a personalization field, rather than at the end of the billing month. Pair it with a hit-rate metric derived from the two usage counters per route, and both silent failure modes become visible.
- Why can the same prompt cache on one Claude model and not on another?Because the minimum cacheable prefix length differs by model class — larger models cache from about 1024 tokens, the smaller Haiku class from about 2048 — and cache entries are keyed per model anyway. A prompt sitting between the two thresholds caches on one and is silently ignored on the other, with no error and no difference in the response shape.
- How would you prove the prefix is stable rather than guessing?Serialize the exact prefix you intend to cache — tools plus system plus any pre-breakpoint messages — hash it, and log the hash with every request. Requests that should share an entry must show the same hash. This turns an invisible failure into a one-line comparison and catches regressions the moment someone templates a dynamic value into the cached region.
- Hit rate collapsed after a deploy but the prompt text is unchanged. What do you look at?The tool definitions and their serialization order, since tools lead the prefix, plus anything that changed how the request is assembled — a new SDK version, a different JSON encoder, a routing change sending traffic to another model. Also check whether the deploy restarted a fleet of workers simultaneously against a cold entry, which produces a burst of writes and no reads.
saying these in an interview costs you the question
- Expects an error when a prefix is too short to cache
- Blames the API instead of checking prefix stability
- Forgets that tool definitions lead the prefix
- Ignores the five-minute expiry on low-traffic services
- Assumes concurrent first requests share one cache write