skip to content

Prompt caching

Marking a stable prefix — system prompt, tool definitions, a large pasted document — with cache_control so later calls reuse it at a fraction of the input price. The questions are about placement: the cache is prefix-based, so anything that changes per request has to sit after the breakpoint.

on this pageshow

questions

5

Where do you place cache_control in an Anthropic Messages request, and why?

level: middleimportance: must knowfreq 72%

answer

  1. prefix, not a fragment
  2. order is tools, system, messages
  3. stable first, volatile last
  4. exact match, one byte breaks it
  5. at most four breakpoints

basics

~20 s

Anthropic's prompt cache is prefix-based: a cache_control marker ends a cacheable prefix rather than caching one block. Put it after everything stable — tools first, then system, then early messages — and keep anything that varies per request after it.

solid answer

~50 s

Anthropic's Messages API caches **prefixes**, not fragments. Adding `"cache_control": {"type": "ephemeral"}` to a content block tells the API to cache everything from the start of the request up to and including that block. The prefix is assembled in a fixed order — `tools`, then `system`, then `messages` — so a breakpoint on the system prompt also covers the tool definitions ahead of it. Lookups are exact matches on that prefix, so a single changed byte anywhere before the breakpoint turns the whole thing into a miss. The design rule follows directly: order content from most stable to most variable, put the breakpoint at the boundary, and keep per-request material (the user's question, retrieved chunks, timestamps, session IDs) strictly after it. A request may carry at most four breakpoints, so spend them on real stability boundaries.

go deeper

for a junior

Know that Anthropic caching is opt-in: you attach cache_control to a block, and it caches everything from the start of the request up to that point. Say plainly that stable text goes before it and the user's question goes after.

for a middle

Be ready to explain the prefix ordering — tools, then system, then messages — and why an exact-match lookup means any edit ahead of the breakpoint is a full miss. Expect to be asked where you would cut a real request.

for a senior

An interviewer expects you to audit a production prompt for accidental variability: injected dates, per-user fields, non-deterministic tool ordering, unstable JSON serialization. Show that you verify placement from the usage counters rather than assuming it works.

for a principal

Own the request layout as an architectural contract: which tiers are stable, how the four breakpoints are budgeted across tools, static context and conversation, and what policy stops teams from slipping volatile data into the cached head as the system grows.

## What "prefix caching" means here When you call Anthropic's Messages API, the request is flattened into one long token sequence in a fixed order: the `tools` array first, then the `system` parameter, then the `messages` array in order. Prompt caching lets the service keep the internal state it computed for a leading run of that sequence and reuse it on a later call, so those tokens do not have to be processed again. The crucial consequence is that the unit of caching is a *prefix* — a contiguous run starting at the very beginning of the request — not an arbitrary block. `cache_control` does not mean "cache this block". It means "the cacheable prefix ends here". ## The mechanics of a breakpoint You mark a breakpoint by attaching `"cache_control": {"type": "ephemeral"}` to a content block: an entry in the `system` array, a tool definition in `tools`, or a content block inside a message. `ephemeral` is the only type; an optional `"ttl": "1h"` extends the default lifetime. On a request, the API hashes the prefix up to each breakpoint and looks it up. A hit means those tokens are billed at a small fraction of the normal input price and skip most of the prefill work; a miss means the prefix is processed normally and then written to the cache for next time. The response reports both outcomes in `usage` via `cache_read_input_tokens` and `cache_creation_input_tokens`. Because matching is exact, *any* difference before the breakpoint is a miss — a changed word in the system prompt, a reordered or edited tool schema, an injected timestamp, a different model. There is no fuzzy or partial matching on content; the only flexibility is that the API will also check for hits on somewhat shorter prefixes near your breakpoint (roughly the preceding twenty content blocks), which is what makes a moving breakpoint in a growing conversation still land on the older cached prefix. ## The layout rule Sort your request from most stable to most variable, then cut: 1. **Tools** — JSON Schema definitions change only when you ship code. They sit at the front of the prefix whether you like it or not, so an unstable tool list poisons everything behind it. 2. **System** — persona, policies, and long static context such as a manual, a codebase excerpt, or a pasted contract. 3. **Early messages** — few-shot examples, an uploaded document turn, or completed conversation turns. 4. **Everything volatile** — the current user question, freshly retrieved RAG chunks, per-user metadata, the current time. Put the breakpoint at the boundary between 3 and 4. The classic beginner failure is interpolating something variable into the system prompt — "Today is {date}", "User: {id}" — which silently makes every request a fresh cache write. That material belongs in the first user message, after the breakpoint. ## Multiple breakpoints A request may carry up to four breakpoints. Use them for genuinely different stability tiers, for example one after the tool definitions, one after the static system context, and one or two rolling markers in a growing conversation. With two breakpoints, a later call can read the matched earlier portion at cache-read price and write only the new segment, so you pay the write premium on the delta rather than the whole prompt. ## Scope and lifetime Cache entries are private to your organization and keyed on the exact prefix plus the model, so different models never share an entry. Entries live five minutes by default, and every hit refreshes that clock, so steady traffic keeps a prefix warm indefinitely. Prefixes below a minimum size — on the order of a thousand tokens, higher for the smaller model tiers — are simply not cached, and no error is raised. ## Verifying it Don't assume; measure. Send the same prefix twice and read `usage`. The first call should show a non-zero `cache_creation_input_tokens`; the second should show a comparable `cache_read_input_tokens` with `input_tokens` reduced to just the tail. If the second call still shows creation rather than read, something before your breakpoint changed between the two requests.

  • If a request has two breakpoints and only the earlier prefix matches, what gets billed?
    You pay cache-read price for the matched leading portion and cache-write price for the new segment between the matched point and the second breakpoint. That is the whole point of a second breakpoint in a growing conversation: the stable head is read cheaply and only the delta is written, instead of re-writing the entire prompt each turn.
  • Does changing max_tokens or temperature invalidate a cached prefix?
    No. Those are request-level sampling parameters, not part of the token sequence that forms the prefix, so they do not affect the lookup. What does invalidate is anything that changes the serialized content ahead of the breakpoint — tool definitions, system text, earlier message blocks — or switching to a different model, since cache entries are model-specific.
  • Where should retrieved RAG chunks go if they change on every query?
    After the breakpoint, in the user turn. Retrieved passages are the most volatile part of a RAG request, so placing them in the system prompt destroys the cache for the stable instructions and tool schemas ahead of them. Cache the persona, policies and tool definitions; pass the per-query evidence as ordinary uncached input.

saying these in an interview costs you the question

  • Thinks cache_control caches only the block it is attached to
  • Interpolates a timestamp or user ID into the cached system prompt
  • Assumes the cache matches similar prompts, not an exact prefix
  • Puts the changing user question ahead of the breakpoint
  • Believes message order in the prefix does not matter

context

open as a page

Which fields in Anthropic's Messages usage object report cache hits and writes?

level: juniorimportance: should knowfreq 44%

basics

~10 s

The usage object reports cache_creation_input_tokens for tokens written to the cache on this call and cache_read_input_tokens for tokens served from it. A non-zero read count means a hit; input_tokens counts only the uncached remainder.

open as a page

What is the default TTL of Anthropic's prompt cache, and what does caching cost?

level: middleimportance: should knowfreq 52%

basics

~20 s

Anthropic cache entries live five minutes by default, and every hit refreshes that window. An optional one-hour TTL is requested with ttl inside cache_control. Writes cost a premium over base input price, reads a small fraction of it.

open as a page

Your Anthropic prompt cache never registers a read — how do you debug it?

level: seniorimportance: should knowfreq 54%

basics

~20 s

Check the usage counters first: persistent cache_creation with zero cache_read means the prefix is either below the minimum cacheable size, changing between calls, expiring between calls, or routed to a different model. None of these raises an error.

open as a page

How would you budget cache breakpoints across a long-running Claude agent loop?

level: principalimportance: should knowfreq 38%

basics

~20 s

Spend the four available breakpoints on stability tiers: one after tool definitions, one after static system context, and one or two rolling markers at the end of settled conversation. Each turn then reads the long stable head and writes only the delta.

open as a page