How would you budget cache breakpoints across a long-running Claude agent loop?
answer
- four breakpoints, spend them on tiers
- read the history, write only the delta
- roll a marker with the conversation
- cadence decides the TTL
- compaction invalidates everything after it
basics
~20 sSpend the four available breakpoints on stability tiers: one after tool definitions, one after static system context, and one or two rolling markers at the end of settled conversation. Each turn then reads the long stable head and writes only the delta.
solid answer
~50 sAn agent re-sends its entire history every turn, so the design goal is that each turn reads almost all of it and writes only what is new. A Messages request allows four `cache_control` breakpoints, and they should map to stability tiers rather than being sprinkled. A typical layout: one after the tool definitions, one after the static system context, and one or two rolling markers placed at the end of conversation content that will not change again — the last completed tool result or assistant turn. On the next turn the earlier breakpoints hit and are billed as reads, while only the segment added since the last marker is written, so the write premium is paid on the delta rather than the whole prompt. Beyond placement, the judgement calls are TTL against turn cadence, whether summarising or truncating history is worth the invalidation it causes, and warming the prefix once before fanning out workers.
code
python · 42 linesimport os
from anthropic import Anthropic
client = Anthropic()
MODEL = os.environ["ANTHROPIC_MODEL"]
TOOLS = [
{
"name": "get_ticket",
"description": "Fetch a support ticket by id.",
"input_schema": {
"type": "object",
"properties": {"id": {"type": "string"}},
"required": ["id"],
},
"cache_control": {"type": "ephemeral"},
}
]
def mark_last_block(messages):
"""Roll a breakpoint onto the newest settled content block."""
for message in messages:
for block in message["content"]:
block.pop("cache_control", None)
messages[-1]["content"][-1]["cache_control"] = {"type": "ephemeral"}
return messages
def turn(messages):
return client.messages.create(
model=MODEL,
max_tokens=1024,
tools=TOOLS,
system=[
{
"type": "text",
"text": "You are a support agent.",
"cache_control": {"type": "ephemeral"},
}
],
messages=mark_last_block(messages),
)go deeper
Know that an agent re-sends its whole history each turn, and that a breakpoint near the end of settled conversation lets that history be re-read cheaply instead of re-billed at full price.
Explain the four-breakpoint budget and the read-plus-small-write pattern per turn, and why a marker must sit on content that will not change again.
Demonstrate operating judgement: matching TTL to inter-turn cadence, warming a prefix before fanning out workers, and recognising that history compaction invalidates everything after the edit point.
Own the whole economics: which agent surfaces justify caching, how compaction policy trades context length against hit rate, tenant isolation as a placement rule, and the metrics that prove the loop still reads more than it writes as it scales.
## The shape of the problem A tool-using agent is a loop: send the whole conversation, get a tool call, execute it, append the result, send everything again. The prompt therefore grows monotonically, and the overwhelming majority of every request is content the model already saw on the previous turn. Without caching you pay full input price for that history on every single iteration, and prefill latency grows with it. With caching placed well, each turn re-reads the history at a fraction of the price and pays the premium only on the newly appended segment. ## Budgeting four breakpoints A request may carry at most four breakpoints, so treat them as a scarce resource allocated to stability tiers, from most stable to least: 1. **After the tool definitions.** They change only on deploy and sit at the very front of the prefix. A breakpoint here means a tool edit costs you this tier and nothing more expensive. 2. **After the static system context.** Persona, policies, and any large fixed corpus — the codebase excerpt, the handbook, the schema. 3. **A rolling marker at the end of settled conversation.** Place it on the last content block that will never change again: the completed tool result or assistant turn from the previous iteration. 4. **Optionally a second rolling marker**, one turn behind the first, so a moving breakpoint always has an older stable entry to land on. With this layout, a turn produces a large `cache_read_input_tokens` and a small `cache_creation_input_tokens` — the shape you should be measuring for. ## Why rolling markers work The API checks for hits not only at your exact breakpoint but on somewhat shorter prefixes near it, covering roughly the preceding twenty content blocks. That tolerance is what lets you move a marker forward each turn and still match the entry written on the previous turn. Without it, advancing the marker would look like a brand-new prefix every time. ## The judgement calls **TTL against cadence.** Interactive agents with a human in the loop go quiet for minutes at a time. If your median inter-turn gap approaches or exceeds five minutes, the default window expires between turns and you pay writes forever; the one-hour tier is the fix, at a higher write premium. Machine-driven loops that iterate in seconds never need it, since every hit refreshes the short window for free. **History compaction versus cache stability.** Summarising or trimming old turns to control context length rewrites the middle of the prefix and invalidates everything after the edit point. That is not a reason to avoid compaction, but it is a reason to do it rarely and in large steps — one big rewrite that re-warms the cache beats trimming a message every turn, which destroys the cache continuously. **What not to cache.** Short-lived or highly variable agents whose prefix is never reused pay the write premium for nothing. Caching is a bet on reuse; when the loop typically terminates after one call, decline the bet. **Warm-up and concurrency.** An entry is only readable after the request that writes it has processed the prefix, so a fleet of workers starting simultaneously against a cold prompt all pay writes. Warm once, then fan out — particularly at deploy time, when every worker restarts at the same moment. **Isolation.** Cache entries are scoped to your organization and matched on the exact prefix, so per-tenant data placed after the breakpoint cannot leak through a shared entry. The corollary is the design rule: shared, non-sensitive context goes ahead of the breakpoint where it is reused across tenants; anything tenant-specific goes after it, which also keeps the shared head reusable. ## Measuring the result Instrument the ratio of cache reads to total prompt tokens per turn and the write volume per turn. A healthy agent loop shows the read share climbing as the conversation grows and the write volume staying roughly flat at the size of one turn's delta. Rising write volume means a marker is not advancing, compaction is thrashing the prefix, or something ahead of the breakpoint has become variable.
- Why does trimming one old message per turn hurt more than a single large compaction?Because any edit inside the prefix invalidates the cache from that point onward. Trimming every turn means the prefix is rewritten every turn, so you pay write premiums continuously and never accumulate reads. One large compaction pays a single invalidation and then lets a long stable stretch build up hits again, which is strictly cheaper for the same reduction in context length.
- Where do per-tenant instructions belong in a multi-tenant agent?After the shared breakpoint. Entries are matched on an exact prefix and scoped to your organization, so tenant-specific text ahead of the marker gives every tenant a private entry and destroys reuse of the common head. Putting shared policies and tool definitions first, tenant data second, keeps one warm shared prefix serving all traffic.
- How do you warm the cache safely at deploy time?Issue one request carrying the full stable prefix, wait for it to return, and only then release traffic to the rest of the fleet. Until the first request has processed the prefix the entry is not readable, so simultaneous cold starts each pay a write. On a large system prompt across many workers that is a real and entirely avoidable cost spike.
saying these in an interview costs you the question
- Placing a breakpoint on every message block
- Summarising history every turn and wondering why hit rate is zero
- Assuming a growing conversation cannot be cached at all
- Putting per-tenant data ahead of the shared breakpoint
- Enabling the long TTL on a loop that iterates every second