skip to content

Which fields in Anthropic's Messages usage object report cache hits and writes?

level: juniorimportance: should knowfreq 44%

answer

  1. two extra counters on usage
  2. creation means miss, read means hit
  3. input_tokens is only the leftovers
  4. streaming reports it in message_start

basics

~10 s

The usage object reports cache_creation_input_tokens for tokens written to the cache on this call and cache_read_input_tokens for tokens served from it. A non-zero read count means a hit; input_tokens counts only the uncached remainder.

solid answer

~40 s

Every Messages response carries a `usage` object, and prompt caching adds two counters to it. `cache_creation_input_tokens` is the number of prompt tokens written into the cache on this request — a miss that populated the entry. `cache_read_input_tokens` is the number served from an existing entry — a hit. The plain `input_tokens` field counts only what was neither read from nor written to the cache, so the three together (plus `output_tokens`) describe the whole request. That split matters for billing, because the three input categories are priced differently: reads at a fraction of the base rate, writes at a premium, ordinary input at list price. In a streaming response these counters arrive in the `message_start` event, since the prompt is fully accounted for before any output is generated.

go deeper

for a junior

Know the two field names and what each means: creation is the write on a miss, read is the hit. Be able to say that a second identical call should show a read.

for a middle

Explain that input_tokens excludes both cached categories, so the total prompt is the sum, and that the split exists because the three categories are priced differently.

for a senior

Show that you instrument these counters as a hit-rate metric per prompt version, and that you know caching fails silently — no error is raised when a prefix is too short or quietly invalidated.

for a principal

Own the cost model: map the counters to per-route spend, decide which surfaces justify caching at all, and set the alerting that catches a prompt change destroying hit rate across a fleet of services.

## The counters An Anthropic Messages response includes a `usage` object. Without caching it holds `input_tokens` and `output_tokens`. With caching enabled, two more fields describe what happened to the prompt: - **`cache_creation_input_tokens`** — prompt tokens processed normally on this request and then written into the cache. This is what a cache *miss* on a marked prefix looks like. - **`cache_read_input_tokens`** — prompt tokens served out of an existing cache entry. This is a *hit*. `input_tokens` is the residual: prompt tokens that were neither read nor written, typically everything after your last breakpoint. It is not the total prompt size, which is the single most common misreading of the object. To recover the total, add all three. ## Reading the pattern The canonical two-call test: - Call 1 — `cache_creation_input_tokens` is large, `cache_read_input_tokens` is 0. The prefix was cold and has now been stored. - Call 2, same prefix, within the TTL — `cache_read_input_tokens` roughly equals the previous creation count, `cache_creation_input_tokens` is 0, and `input_tokens` shrinks to just the new tail. If call 2 still shows creation rather than read, the prefix is not stable, the entry expired, or the prefix is below the minimum cacheable size and was never stored at all. A mixed result — both counters non-zero — is normal and healthy in a multi-breakpoint conversation: the older prefix hit, and the segment added since the last call was written as a new entry. ## Why the split exists The three input categories are billed at different rates. Cache reads cost a small fraction of the base input price; cache writes cost a premium over it; ordinary input is at list price. A single `input_tokens` number could not express that, so the API exposes the breakdown and lets you reconstruct the cost yourself. This is also what makes the counters the right basis for a cache hit-rate metric: emit reads divided by reads-plus-creations per route, and a regression in prompt stability shows up as a falling ratio before it shows up as a surprising invoice. When you request the extended one-hour lifetime, the write counters are also broken out by TTL bucket in a nested `cache_creation` object, so you can tell five-minute writes from one-hour writes — which matters because the two are priced differently. ## Streaming With `stream=true`, the prompt is accounted for before generation starts, so the cache counters appear in the `message_start` event's `usage`. The trailing `message_delta` event carries final output-token totals. If you only read the trailing event, you will see output accounting and conclude caching is not reported at all. ## Operational habits Log the four numbers per request alongside the route or prompt version. Two failure modes are only visible this way. First, silent invalidation: someone adds a dynamic field to a cached system prompt and creation counts replace read counts overnight. Second, pointless caching: short prefixes that never reach the minimum size, so `cache_creation_input_tokens` stays at zero forever while you assume caching is active. Neither raises an error — prompt caching fails quietly by design, and the usage object is the only signal you get.

  • Both cache_creation_input_tokens and cache_read_input_tokens are non-zero on one call. Is that a bug?
    No, it is the expected shape for an incremental cache. An earlier breakpoint matched and was billed as a read, while the content added since the last request was written as a new entry at the later breakpoint. In a growing agent conversation you want exactly this: a large read plus a small write each turn.
  • How would you turn these counters into a cache health metric?
    Emit all four usage numbers per request tagged with the prompt version and route, then track cache_read_input_tokens divided by the sum of reads and creations. A sustained drop means a prompt edit introduced variability ahead of the breakpoint, or traffic slowed enough that entries expire between calls. It surfaces regressions long before the billing does.

saying these in an interview costs you the question

  • Thinks input_tokens is the total prompt size
  • Confuses cache_creation with a cache hit
  • Looks for cache counters only at the end of a stream
  • Assumes zero cache tokens means an API error
  • Adds the counters to input_tokens and double-counts the prompt

context