skip to content

When a coding agent compacts its conversation, what actually happens to the session?

level: juniorimportance: must knowfreq 62%

answer

  1. the window fills, the task does not stop
  2. history out, summary in
  3. recent turns kept verbatim
  4. one-way: nothing is reconstructed
  5. system prompt and tools stay put

basics

~10 s

Compaction replaces the accumulated history with a written summary plus the most recent turns, then continues the same task in a nearly empty window. Work carries on, but the raw earlier messages are gone.

solid answer

~50 s

An agent session grows monotonically: every user message, model turn, tool call and tool result is appended, and the whole list is resent on each request because the model holds no state between calls. Eventually the window fills. Compaction is the reinitialization step: the agent has the model write a structured summary of the session so far — goal, decisions taken, artifacts in flight, position in the plan — then rebuilds the context as system prompt, tool definitions, that summary, and the last few turns kept verbatim. The same task continues; it is not a new chat. The property that matters is that compaction is **lossy and one-way from the model's point of view**: anything not captured in the summary and not recoverable by re-reading a file or re-running a query is gone. What goes into the summary therefore matters far more than the compression ratio.

go deeper

for a junior

Be able to say plainly that compaction swaps a long history for a summary plus the newest turns, and that the same task continues afterwards.

for a middle

Explain the sequence — threshold, summarizing call, rebuilt message list — and why the recent turns are kept raw while the middle of the conversation is condensed.

for a senior

Show that you treat it as lossy and instrument for the loss: which facts existed only in the transcript, and how you detect an agent acting as though a constraint no longer exists.

for a principal

Own the position that compaction is one of three answers alongside file-backed notes and sub-agent isolation, and argue when a session should be redesigned so it never needs compacting at all.

## Why a session needs compacting An agent conversation is append-only by construction. The harness keeps a running list of messages — the user's request, the model's reasoning and tool calls, and the output those tools returned — and resends that entire list on every turn, because the model itself carries no state between calls. A long coding or migration session therefore does not merely get slower and costlier; it eventually exceeds the context window outright. Well before that hard ceiling it enters the degraded region practitioners call context rot, where a very full window measurably hurts recall and planning quality even though nothing has technically overflowed. Compaction is the standard response. Alongside externalizing notes to files and isolating work in sub-agents, it is one of the three accepted answers to a filling window. ## The mechanic, step by step 1. **Trigger.** The harness notices utilization crossing a threshold — a fraction of the usable window, not the advertised number. 2. **Summarize.** It issues one extra model call whose prompt is roughly "write down everything a fresh instance of you would need to continue this task," usually against a fixed template so the output is structured rather than narrative. 3. **Rebuild.** It constructs a new message list: the unchanged system prompt and tool definitions, then the summary as a single message, then the most recent turns copied verbatim — the ones the agent is actively working within. 4. **Continue.** The next turn runs against that new context. Utilization drops from, say, 90% to 30%, and the loop proceeds on the same goal. The verbatim tail matters. A summary alone loses the fine texture of what just happened — the exact error string, the half-finished edit — and the agent's next action usually depends on it. Keeping the last few turns raw is cheap and prevents the most common post-compaction stumble, which is an agent that redoes the step it had just completed. ## What compaction is not - **Not a new session.** The goal, the user's constraints and the plan survive by design. If the agent asks the user to restate the task after compacting, the summary was inadequate. - **Not lossless.** Nothing reconstructs the removed messages. Detail that existed only in the conversation — a reason for a decision, a constraint the user gave in passing — is unrecoverable once dropped. - **Not free.** It costs a model call, and it rewrites the front of the context, which throws away the prompt-cache discount the session had been enjoying. - **Not the only lever.** If the bulk of the window is old tool output rather than reasoning, clearing those results is mechanical, cheaper, and preserves the reasoning thread intact. ## Why the loss is asymmetric Some material is recoverable from outside the window and some is not. A file's contents can be re-read; a query can be re-run; a repository can be re-grepped. Those things do not need summarizing — carrying a path or a query identifier is enough, and re-reading gives the current version rather than a stale copy. But a decision the agent made three hours ago and the reason behind it exist nowhere except in the transcript. Same for a constraint the user stated once. That asymmetry is the whole design principle: summarize what only the conversation knows, and keep pointers to everything the world still holds. ## The failure everyone has seen The canonical bad outcome is a summary that captured the mechanics but dropped a standing instruction. A user says early on "never run this against production"; four hours later the session compacts, that sentence is not in the summary, and the agent's next proposal is exactly the forbidden action. It has not malfunctioned — the instruction simply no longer exists anywhere it can see. This is why mature harnesses treat standing constraints as a required, always-carried section of the summary template rather than leaving their survival to the model's judgement.

  • Why keep the last few turns verbatim instead of summarizing everything?
    Because the agent's immediate next action usually depends on fine detail the summary flattens — the exact error text, the half-applied edit, the last tool result. Keeping the tail raw is cheap and prevents the most common post-compaction stumble, where the agent repeats a step it had already finished or misreads where it left off.
  • How is compaction different from just letting the oldest messages be truncated?
    Truncation drops content blindly, so whichever facts sat oldest — usually the goal and the user's constraints, which arrive first — are exactly what disappears. Compaction spends a model call to decide what to keep, so early, load-bearing material survives in condensed form while bulky middle content goes.
  • What signal tells you a compaction went badly?
    The agent immediately redoes completed work, asks the user for information already given, or proposes something a stated constraint forbids. Those are all one symptom: something load-bearing existed only in the removed history. Instrumenting for them is how you tune the summary template.

saying these in an interview costs you the question

  • Thinks compaction starts a new task or clears the goal
  • Believes the original messages can be restored afterwards
  • Assumes compaction is free rather than costing a model call
  • Thinks a bigger context window removes the need for it entirely

context