skip to content

How many cache breakpoints do you set, and at which layers of a prompt?

level: middleimportance: should knowfreq 48%

answer

  1. one per stability tier, not per block
  2. a breakpoint marks the end of a prefix
  3. several give partial hits, not all-or-nothing
  4. most stable layer must lead
  5. dynamic tools belong below the system prompt

basics

~20 s

Set one breakpoint per stability tier, not per block — commonly after tool definitions, after the static documents and examples, and after the last completed conversation turn. Explicit breakpoints are capped at a handful, so spend them on real change boundaries.

solid answer

~50 s

A cache breakpoint marks "the prompt up to here is a reusable prefix", and it is also a candidate hit point: with several of them, a change low in the prompt can still hit at an earlier breakpoint instead of losing everything. So place them where the *rate of change* shifts — after the system prompt and tool definitions, after the large static corpus and few-shot examples, and after the last completed turn in a conversation. One per tier is enough; a breakpoint after every document buys nothing and providers cap explicit markers at a small number anyway. The ordering matters as much as the count: whatever sits first must be the most stable thing you have. Agents that render tool definitions ahead of the system prompt lose the entire prefix the moment a tool is added mid-session.

code

yaml · 16 lines
yaml
tiers:
  - name: static_boot        # changes on deploy
    blocks: [system_prompt]
    breakpoint: true
  - name: tools              # may change mid-session
    blocks: [tool_definitions]
    breakpoint: true
  - name: reference          # changes on content update
    blocks: [policy_corpus, few_shot_examples]
    breakpoint: true
  - name: conversation       # grows every turn
    blocks: [completed_turns]
    breakpoint: true
  - name: request            # never cacheable
    blocks: [retrieved_snippets, user_message]
    breakpoint: false

go deeper

for a junior

Know that a breakpoint marks the end of a reusable region, and that you place it after the parts of the prompt that stay the same rather than sprinkling it between messages.

for a middle

Be ready to name the natural boundaries — after the system prompt and tools, after static documents, after the last completed turn — and to explain that extra breakpoints give partial hits instead of all-or-nothing reuse.

for a senior

Show that you would check how the request is actually assembled in code, not just how the template reads, and catch the case where a dynamic toolset sits above stable content and quietly costs the whole prefix.

for a principal

Treat the layer order and breakpoint policy as an owned contract across teams: a documented tier list, a limited budget of markers spent on real change boundaries, and a review rule that new content is appended to a tier rather than inserted above one.

## What a breakpoint is A cache breakpoint is a marker saying "everything from the start of the request up to this point is a prefix worth storing and reusing". Providers differ in how you express it. Some require you to place explicit markers and cap how many you may use — Anthropic's API allows up to four per request, checking for a hit at each one, as of mid-2026. Others cache automatically once a prefix exceeds a minimum length; OpenAI's automatic caching kicks in above roughly a thousand tokens with no marker at all. The placement reasoning below is provider-independent: even where markers are automatic, you are still choosing *where the stability boundaries fall*, and the same layout decisions apply. ## Why more than one With a single breakpoint you get one all-or-nothing outcome: either the request matches up to that point or it does not. With several, you get graceful degradation. Suppose your prompt is system prompt → tools → policy corpus → history. If the corpus is updated but the system prompt and tools are untouched, a breakpoint after the tools still hits and only the corpus and everything after it is recomputed. Without that earlier breakpoint, the corpus change costs you the whole prefix. That is the entire argument for multiple breakpoints: each one is an insurance policy against changes in the layers below it. ## Which layers Work down the stack and put a breakpoint wherever the answer to "when does this change?" changes: - **After the system prompt and tool definitions.** These move on deploy. If your toolset is dynamic within a session, split them: static system prompt, breakpoint, tools, breakpoint. - **After the large static content** — the policy corpus, the reference document, the few-shot example block. These move on a content update, typically much less often than a session. - **After the last completed conversation turn.** This is the one that moves every turn, and it is the workhorse breakpoint in a chat application. What you do *not* do is place one after every document in a 40-document corpus. Documents in a corpus all change on the same event, so the extra breakpoints protect against nothing, and you will run out of your allowance before reaching the boundaries that matter. ## Order layers by rate of change, not by convention Breakpoint placement only works if the underlying ordering is right, and the ordering rule is: the topmost block must be the most stable thing in the prompt. Convention often violates this. Many agent frameworks assemble the request as tools first, then system prompt, because that is the order the objects were constructed in. That is fine when the toolset is fixed at deploy. It is a trap in an agent that loads tools dynamically — progressive disclosure, per-task toolsets, a plugin activating mid-session. Add one tool at turn six and every token after the tool block changes: the system prompt, the corpus, all six turns of history. The prefix collapses to whatever preceded the tools, which is nothing. Reorder so the immutable system prompt leads and the mutable tool block sits behind it with a breakpoint between them, and the same mid-session tool addition costs you only the tokens below the tools. ## A placement recipe 1. List every block with the event that changes it. 2. Group blocks that change on the same event into a tier. 3. Sort tiers slowest-changing first, and check that this sort did not get overridden by how your code happens to build the request. 4. Put one breakpoint at the end of each tier, from the top, until you run out of breakpoints or tiers. 5. Spend any remaining breakpoint on the boundary you expect to be crossed most often — usually the end of the conversation history. ## When one is enough A single-shot classification or extraction prompt with a fixed system prompt, a fixed instruction block and one variable input has exactly one boundary. Add a second breakpoint and you have added nothing except an extra thing to maintain. Complexity here should be proportional to how many genuinely independent change schedules your prompt contains — most prompts have two or three, not four. ## Common mistakes Treating breakpoints as decorations sprinkled at message boundaries; placing them *before* the block they are meant to protect rather than after it; putting one in the middle of the fastest-changing region, where it can essentially never be hit twice; and forgetting that the block ordering, not the markers, is what actually determines what is reusable. A breakpoint cannot rescue a prompt whose first line contains today's date.

  • What does a second breakpoint buy you that the first one does not?
    Partial reuse. With one breakpoint a change anywhere above it costs the whole prefix; with two, a change below the first breakpoint still hits at the first one. Each breakpoint is a fallback position against churn in the layers beneath it. That only pays when the layers really do change on independent schedules — repeating breakpoints within one tier adds nothing.
  • Where would you put the breakpoint in a single-turn extraction prompt with a fixed instruction and one variable document?
    One breakpoint, at the end of the fixed instruction block, immediately before the document. There is exactly one stability boundary in that prompt, so a second marker has nothing to protect. If the document is large and the instruction small, be honest that caching has little to offer here and look at whether the instruction block can carry more of the weight, such as few-shot examples.
  • Your agent loads extra tools mid-session. Where should the tool block sit?
    Below the immutable system prompt, with a breakpoint between them. Then adding a tool at turn six invalidates only the tokens after the tool block, and the system prompt still hits at the earlier breakpoint. If tools sit above the system prompt, the same addition wipes out the entire prefix including all prior turns — a large loss for a small change.

saying these in an interview costs you the question

  • Puts a breakpoint after every document in the corpus
  • Places dynamic tool definitions above the static system prompt
  • Thinks one breakpoint is always as good as several
  • Sets a breakpoint inside the per-request tail
  • Believes markers matter more than block ordering

context