Why does compacting an agent conversation wipe out its prompt-cache savings?
answer
- reuse needs an identical prefix
- edits kill everything after them
- append-only is what keeps it cheap
- stable head, volatile tail
- rare and deep beats often and shallow
basics
~20 sCache reuse requires the request prefix to be identical to a previous one. Compaction rewrites the history near the front, so everything after the edit is new and the next turn is billed and computed as an uncached prefill.
solid answer
~50 sPrompt caching pays off in agent loops because their requests are append-only: each turn resends the same long prefix with a little added at the end, so the provider reuses the cached prefix and charges a fraction of the full rate. Compaction breaks that assumption. It removes and rewrites messages in the middle of the context, so from the edit point onward nothing matches the cached entry and the following turn pays full price for the entire rebuilt body. In sessions where prefix caching had been cutting input cost by roughly ninety percent, one compaction resets the session to zero savings until the new prefix has been read enough times to earn them back. The design consequences: keep the truly stable material — system prompt, tool definitions — ahead of any mutable region so it survives untouched, compact rarely and deeply rather than often and shallowly, and batch a tool-result clear into the same rewrite so the invalidation is paid once.
go deeper
Know that repeated agent requests are cheaper because the unchanged beginning of the prompt can be reused, and that changing it removes that saving.
Explain that reuse is a prefix match, so an edit in the middle makes everything after it new, and that this is why agent contexts are designed append-only.
Price a compaction as three costs — the summarizing call, fidelity loss, and the cache reset — and give the design responses: volatility layering, rare and deep compaction, batched edits, plus the telemetry that reveals accidental invalidation.
Own the budget-level tradeoff across a fleet: how compaction frequency, sub-agent isolation and session length interact with cached-token economics, and where the quality floor stops you optimizing purely for cache hits.
## The economics compaction interrupts An agent loop is unusually friendly to prompt caching. Every turn resends the same system prompt, the same tool definitions and the same accumulated history, with one new exchange appended. Because the prefix is identical to last turn's, providers can serve it from cache — reported savings on the order of ninety percent on input cost and a large cut in time-to-first-token, reflecting mid-2026 pricing practice. In a hundred-turn session that discount is not a rounding error; it is most of the bill. The mechanism depends on one property: **the prefix must be unchanged**. Caching is a prefix match, so reuse extends only as far as the request agrees with what was cached, and stops at the first difference. ## What compaction does to that Compaction is precisely a modification of the middle of the context. Dozens of messages are replaced by one summary. The system prompt and tool definitions may still match, but the moment the message list diverges, every token from there to the end is new. The next request therefore reuses only the small stable head and pays full uncached price for the entire rebuilt body — and pays the latency of prefilling it too. So the true cost of a compaction is three-part: the summarizing model call, the fidelity lost in summarizing, and this cache reset. Practitioners often account for the first and ignore the third, which is the larger number in a long session. The same is true of any other edit. Clearing stale tool results, reordering messages, injecting a freshly retrieved document near the top, or stamping the current time into the system prompt each turn all end reuse from the point of change. The rule generalizes: **design the context append-only and treat every edit as a cache event**. ## Design responses **Layer the context by volatility.** Put what never changes first — system prompt, tool definitions, long-lived instructions — then the mutable history. Compaction then rewrites only the volatile tail, and the stable head continues to be served from cache. Putting anything volatile in the header, such as a timestamp or a per-turn retrieved snippet, sacrifices the whole cache every turn and is a classic, expensive mistake. **Compact rarely and deeply.** Since each compaction costs a full reset, ten shallow compactions cost roughly ten resets while one deep compaction costs one. This reinforces the two-threshold trigger policy from a purely economic direction: fire at a moderate mark and reduce far, so the rebuilt prefix has many turns to earn its discount back before the next event. **Batch edits together.** If you are going to clear stale tool results anyway, do it in the same rewrite as the compaction. Trimming a message here and a result there on separate turns pays the invalidation repeatedly for the same benefit. **Let the new prefix stabilize.** Immediately after compaction the context is cold; the first turn pays full price and subsequent turns start reusing again. That is normal, and it is another reason not to schedule compactions close together. ## What this does not mean It does not mean avoiding compaction. Running deep into a degraded window to protect a cache trades quality for cost in the wrong direction, and you overflow eventually anyway. The right conclusion is that compaction is a priced event to be scheduled deliberately — rare, deep, at a clean boundary, batched with any other pending edit — not an invisible housekeeping step that fires whenever a threshold twitches. ## Recognising it in production The signature is a session whose input tokens are largely uncached despite an append-only-looking workload. Providers report cached and uncached input separately; a long agent session that should be around ninety percent cache hits and shows a much lower ratio is usually either compacting too often or mutating something near the front of the prompt every turn. The second cause is worth checking first because it is so easy to introduce accidentally — one dynamic value in a system prompt is enough. ## The interview shape A strong answer states the prefix-match requirement plainly, connects it to why agent loops are cache-friendly in the first place, names the invalidation as a real line item next to the summarizing call, then gives the design responses: volatility layering, rare and deep compaction, batched edits. A weak answer treats caching as an automatic discount that either applies or does not, with no model of why.
- Besides compaction, what commonly invalidates an agent's prompt cache by accident?Anything volatile placed near the front: a current timestamp in the system prompt, a per-turn retrieved snippet injected above the history, a tool list whose order is not stable, or user-specific data interleaved into a shared header. Each changes the prefix every turn, so reuse never extends past it and the session pays full price throughout.
- Does this argue for delaying compaction as long as possible to protect the cache?No. Running deep into the degraded region trades planning quality for cost in the wrong direction, and you overflow eventually anyway. The correct response is to make compactions rare and deep, batched with any other pending edit, so the number of resets stays small — not to postpone them until the window is nearly full.
- How do you verify the effect rather than assume it?Providers report cached and uncached input tokens separately. Compare the ratio for a long agent session against what an append-only workload should achieve, and correlate dips with compaction events in your traces. A persistently low ratio with no compactions points instead at something volatile sitting near the front of the prompt.
A cached prefix is like a book someone has already read up to a bookmark: adding pages at the end is free to resume from, but inserting a rewritten chapter in the middle forces a re-read of everything after it.
saying these in an interview costs you the question
- Believing a summary is still cache-eligible because the tail is unchanged
- Counting only the summarizing call as the cost of compaction
- Trimming a message or two every turn, paying invalidation repeatedly
- Putting a timestamp or per-turn retrieval above the stable prefix
- Assuming caching is an automatic discount independent of request shape