When is clearing old tool results cheaper than compacting an agent conversation?
answer
- diagnose what is filling the window
- payloads versus reasoning
- mechanical beats semantic when possible
- keep the newest results verbatim
- leave a placeholder behind
basics
~20 sWhen the window is full of bulky tool output rather than long reasoning. Dropping old tool results is mechanical, costs no model call, and leaves the reasoning thread intact — compaction is for when the conversation itself, not its payloads, has grown long.
solid answer
~50 sDiagnose what is actually consuming the window. If thirty stale tool results holding a two-hundred-thousand-token log dump are the bulk, clearing them is strictly better: it needs no summarizing call, it is deterministic rather than lossy over the model's own reasoning, and keeping the most recent one or two verbatim preserves whatever the agent is currently acting on. As of mid-2026 providers expose this as a server-side context-editing feature that strips old tool-use and tool-result blocks once a token threshold is crossed, which is why it costs nothing to run. Compaction earns its model call when the length lives in the *conversation* — many turns of planning, decisions and back-and-forth that no mechanical rule can shrink. In practice they compose: clear first, re-measure, and compact only if utilization is still high. Leave a placeholder where a result was cleared so the agent knows the call happened.
go deeper
Know that old tool output can simply be dropped from the history, and that this is different from summarizing the whole conversation.
Be ready to diagnose first — is the window full of payloads or of reasoning? — and explain why the mechanical, no-model-call option should be tried before the summarizing one.
Demonstrate the composed policy: clear stale results, re-measure, compact only if still high; and cover the operational details of placeholders, keep-most-recent counts, and protecting load-bearing output.
Own the tradeoff at fleet level: which workloads should never accumulate large payloads in the main thread at all, versus where clearing and compaction are the pragmatic answer, and what that implies for how tools are designed to return results.
## Two different problems that both look like a full window "Utilization is at 85%" is a symptom, not a diagnosis. Two very different sessions produce it. One is a long-horizon planning session: eighty turns of reasoning, decisions and corrections, each modest in size. The other is a short session that ran a handful of tools which returned enormous payloads — a full log fetch, a large file read, a wide query result. The right remedy differs, and reaching for compaction reflexively wastes a model call on the second case while losing fidelity you did not need to lose. ## Why clearing wins when payloads dominate **It costs nothing to compute.** Removing old tool-use and tool-result blocks is a mechanical edit to the message list. There is no summarizing call, no latency spike, and no chance the summarizer misunderstands what it read. **It is lossless where loss hurts most.** Compaction compresses the model's own reasoning — the decisions and the thread of the plan. Clearing touches none of that; the assistant turns stay exactly as written and only the bulky observations vanish. The reasoning that *interpreted* a log dump usually survives in the assistant turn that followed it, which is the durable part anyway. **It is targeted.** You can keep the most recent one or two results verbatim, which is nearly always what the agent is acting on, and clear only what is stale. Age-based and count-based policies are trivial to express: clear results older than N turns, or keep the newest K. **The reclamation is large.** When a couple of payloads dominate, clearing them recovers most of the window in one stroke. Published results for server-side clearing on tool-heavy sessions report token reductions well past 80% on long search-style runs, reflecting practice as of mid-2026. ## Where clearing is not enough If the window is long because the *conversation* is long, clearing barely helps: there is no fat to trim, only many small turns whose collective length is the problem. That is compaction's job — a semantic compression only a model can do, deciding that forty turns of exploration reduce to three sentences of conclusion. Similarly, if a tool result is still load-bearing — the agent derived a plan from it and will need it again — clearing it silently removes the ground under the plan. The fix there is to summarize that result into the thread, or write it to a file, before it is cleared. ## Sequencing them The cheap, mechanical, lossless move goes first. A reasonable policy: at the trigger threshold, clear tool results older than some age while keeping the newest few; re-measure; compact only if utilization is still above the mark. Many sessions never reach the second step. Sessions that do reach it arrive with a cleaner history to summarize, because the summarizer is no longer wading through log dumps to find the decisions. ## Placeholders and dangling references A cleared result should not vanish without trace. If an assistant turn contains a tool call and the matching result is gone, the model can conclude the call never completed and re-run it — sometimes an expensive or destructive re-run. Leaving a short placeholder in the result's position ("output cleared to save context; re-run the command if needed") preserves the causal structure and tells the model the output is recoverable rather than missing. Provider-side clearing does this for you; a hand-rolled implementation must remember to. ## The cost they share Neither move escapes one consequence: both edit the history, and any edit to the request prefix ends the prompt-cache reuse that made the preceding turns cheap. That argues for doing the clearing in occasional larger batches rather than trimming a message or two every turn, and for pairing a clear with a compaction rather than paying the invalidation twice. It does not change the ordering advice; it changes the frequency. ## The interview shape A good answer starts with diagnosis — what is in the window — rather than jumping to a technique, names both levers, gives the composition order, and mentions the placeholder detail. Candidates who only know compaction reveal that they have never watched an agent session's token breakdown, which in tool-heavy work is usually dominated by observations rather than reasoning.
- A cleared tool result turns out to have been load-bearing. How do you prevent that?Before clearing, have the agent write the durable conclusion somewhere that survives — a line in the conversation thread or a notes file. Clearing policies should also protect anything the agent explicitly marked as still in use, and keeping the newest few results verbatim covers most cases, since load-bearing output is usually recent.
- Does clearing tool results avoid the prompt-cache penalty that compaction pays?No. Any edit to the request prefix ends reuse from the edit point onward, so a clear invalidates just as a rewrite does. The practical implication is frequency, not choice: batch clearing into occasional larger sweeps rather than trimming a message every turn, and pair a clear with a compaction so you pay the invalidation once.
- How would you decide how many recent results to keep verbatim?By what the agent is plausibly still acting on. One or two cover the common case of interpreting the last observation; more if the workflow compares several outputs, such as diffing two query results. Measure it — if the agent starts re-running commands it already ran, the keep count is too low.
saying these in an interview costs you the question
- Reaching for compaction without looking at what actually fills the window
- Believing clearing requires a model call the way summarization does
- Deleting results with no placeholder, so the agent re-runs the tool
- Assuming clearing preserves the prompt cache because it is only a deletion
- Clearing the most recent results along with the stale ones