Why do agent answers degrade long before the context window fills, and what fixes it?
answer
- degradation before the limit is reached
- long sessions rot, they do not truncate
- attention diluted by stale turns
- summarize and reinitialize the window
- pin goal and constraints outside the summary
basics
~20 sQuality decays as a window fills with stale turns and superseded tool output — context rot — well inside the nominal limit. The fix is context engineering: compact the session into a structured summary and reinitialize the window, keeping only what the remaining work needs.
solid answer
~50 sContext rot is measurable degradation as the window grows: the model starts missing instructions given early, re-runs work it already did, and reasons over superseded intermediate results. It is not a hard truncation error — it appears well inside the advertised limit, because attention is diluted across a lot of low-value tokens and because contradictory states now coexist in the same window. The remedy is to stop treating the window as an append-only transcript. **Compaction** summarizes the session into a structured carry-forward — goal, decisions made, open questions, current state, pinned constraints — and reinitializes the window from it. Alongside compaction you offload bulky payloads behind references and retrieve on demand. As of mid-2026 vendor evaluations of context editing and compaction report large token reductions and double-digit quality gains on long-horizon tasks, which is why compaction is standard in production harnesses rather than an optimization.
go deeper
Know the term context rot and that it means answers get worse as a conversation grows, before any hard limit is reached, and that applications deliberately trim or summarize history rather than sending everything.
Explain the mechanics: attention diluted across low-value tokens, superseded tool results still present, and bulky raw output never leaving. Describe compaction as summarize-and-reinitialize rather than delete-the-oldest.
Demonstrate that you measure it — quality bucketed by turn index or window occupancy — and that you know what a carry-forward summary must preserve. Be ready to describe a long real session and what broke across a compaction boundary.
Own the memory-hierarchy framing: compaction, offloading, retrieval and isolation as one budget policy, with explicit thresholds, an eval that spans compaction boundaries, and a cost model for the tokens you are choosing to carry.
## What context rot actually is Context rot is the observed decline in an LLM's reliability as its context window fills, appearing well before the model's advertised token limit is reached. Symptoms are recognisable: the model forgets a constraint stated at the start of the session, repeats a step it already completed, cites a value that was later corrected, or drifts from the original goal into whatever topic dominates the recent tokens. It is important that this is a *gradual* curve, not a cliff. Truncation — where the oldest messages simply fall out of the request — is a different, sharper failure that your own code causes. Rot happens while everything is still in the window. ## Why it happens Three mechanisms compound. **Attention dilution.** Relevance is competitive. A crucial instruction that occupied 2% of a short window occupies 0.05% of a long one, competing with thousands of tokens of routine tool chatter. **Superseded state.** Long sessions accumulate contradictions by construction. An early tool result said three pods were pending; a later one said zero. Both are in the window, both look authoritative, and nothing marks the first as obsolete. The model may reason over either. **Low-value bulk.** Raw tool output is the biggest offender — full log dumps, entire file contents, verbose API payloads. Most of it was needed for one decision and is dead weight afterwards, but it never leaves. ## How you detect it Do not rely on vibes. Instrument the session: track tokens in the window per turn, and track quality metrics *bucketed by turn index or by window occupancy*. If your task-completion rate at turn 40 is materially below turn 5 on comparable tasks, you have rot, and you now have a threshold to act on. Repeated-action rate and step count are useful proxies too, since re-doing work is one of the earliest symptoms. Complaints that "the model got dumber later in the conversation" are almost always this, and almost always measurable. ## Compaction Compaction is the primary remedy: near a chosen occupancy threshold, summarize the session so far and restart the window from that summary plus the most recent raw turns. What makes it work or fail is *what you preserve*. A free-form prose summary loses precisely the things the next hour depends on. A structured carry-forward is far more reliable: - the original goal and any hard constraints, copied verbatim rather than paraphrased - decisions taken and the reason for each - current state of the world as last observed, with the superseded versions dropped - open questions and the next intended step - handles or identifiers for any offloaded artefacts, so the agent can re-read them - the last few turns kept raw, so immediate continuity is not summarized away Concretely: an on-call copilot riding a six-hour SEV-1 bridge will exceed any window. Compacted twice with that structure, it retains the incident timeline, the mitigations already tried and rejected, and the current hypothesis, while shedding a hundred thousand tokens of raw metric queries. Compacted with a naive "summarize the conversation" prompt, it forgets that a rollback was already attempted and proposes it again — the classic failure. Compaction is lossy and it is a recurring risk: the second compaction summarizes a summary. Guard it by pinning invariant material (goal, constraints, safety rules) outside the summarizable region so it is re-injected verbatim every time, and by keeping the full log durable so anything dropped can be re-read rather than lost. ## The complementary techniques Compaction is one of four levers, and mature harnesses use them together: **Offloading.** Store large payloads externally and put a reference plus a short digest in the window. This is prevention rather than cure; it is what keeps you from needing to compact every ten turns. **Retrieval.** Pull material in on demand instead of pre-loading it, so the window carries what the current step needs rather than what the whole task might need. **Isolation.** Give a bounded piece of work its own window and bring back only the conclusion, so the exploration's token cost never enters the main session. ## Judgement in the interview The strong answer resists two temptations. The first is "use a bigger window" — a larger limit postpones truncation but does nothing about dilution, and it raises cost and latency for every call. The second is "drop the oldest messages" — cheap, but it deletes exactly the goal and constraints stated at the start, which is why sliding-window truncation produces goal drift. Both come up constantly and both are worth naming as the wrong answer before giving the right one. The deeper point is architectural: the context window is a working set you curate every turn, not a transcript you accumulate. Once you accept that, compaction, offloading, retrieval and isolation stop looking like tricks and start looking like the memory hierarchy they are.
- What must survive compaction, and how do you verify it did?The goal, hard constraints, decisions already taken with their reasons, current observed state, open questions, and handles to offloaded artefacts. Verify with an eval that runs long sessions across a compaction boundary and asserts on facts stated before it — for example, that a mitigation already tried is not proposed again. Pin invariant material outside the summarized region so it is re-injected verbatim rather than paraphrased.
- When is compaction the wrong response to a filling window?When the remaining work needs full fidelity of what you would be summarizing — reconciling two long documents line by line, for instance. Then the answer is to restructure: offload the material behind references and have the agent read spans on demand, or hand the bounded comparison to a separate window and bring back only the result. Compaction trades detail for room; do not spend it on the detail the task is about.
- Why is simply moving to a model with a much larger context window not a sufficient fix?Because rot is degradation *within* the limit, not at it. A larger window buys headroom but the same dilution and superseded-state problems recur further along the curve, while every call now pays for more tokens in cost and latency. Larger windows are genuinely useful, but they change when you must curate, not whether.
saying these in an interview costs you the question
- Claims quality is fine until the token limit is hit
- Fixes it by dropping the oldest messages verbatim
- Says a bigger context window removes the need to compact
- Summarizes with free prose and loses goal and constraints
- Treats the window as an append-only transcript