skip to content

At what window utilization should a long agent compact, and how do you prevent thrash?

level: seniorimportance: should knowfreq 52%

answer

  1. two marks, not one
  2. measure against the usable ceiling
  3. reserve room for the next tool result
  4. compact deep so it happens rarely
  5. never mid tool call

basics

~20 s

Trigger on a high-water mark measured against the usable ceiling, not the advertised window, with headroom for the largest plausible next tool result. Then compact deep, to a low-water target, so the session cannot re-trigger within a few turns.

solid answer

~50 s

Use two thresholds, not one. A **high-water mark** — commonly around 70–85% of the *effective* window, since usable context sits well below the advertised figure — fires the compaction, and it must leave headroom for the biggest tool result the session can realistically produce plus the response itself, or the very turn that discovers the problem overflows. A **low-water target** governs how far you compact: reduce to roughly a third, so dozens of turns pass before the next trigger. Thrash is the failure of running one threshold and trimming only just enough: the session compacts every few turns, each pass paying a model call, invalidating the prompt cache, and re-summarizing a previous summary until drift accumulates. Two more rules: compact only at a clean turn boundary, never between a tool call and its result, and clear stale tool results first, since that may drop utilization enough that no compaction is needed.

code

python · 12 lines
python
EFFECTIVE_LIMIT = 400_000   # measured usable ceiling, not the advertised window
HIGH_WATER = 0.75
LOW_WATER = 0.35

def should_compact(used_tokens, largest_result_estimate, max_output, at_turn_boundary):
    if not at_turn_boundary:
        return False
    projected = used_tokens + largest_result_estimate + max_output
    return projected > HIGH_WATER * EFFECTIVE_LIMIT

def compaction_target():
    return int(LOW_WATER * EFFECTIVE_LIMIT)

go deeper

for a junior

Know that compaction is triggered automatically when the context gets close to full, rather than being something the user has to request.

for a middle

Explain the two-threshold shape — a high-water mark to fire and a low-water target to compact down to — and why leaving headroom for the next tool result matters.

for a senior

Show that you price the event: a model call, an invalidated cache, and drift that compounds across passes. Give a concrete policy including the boundary rule and clearing stale results first, and name the telemetry that reveals thrash.

for a principal

Own the sizing of these numbers per model and workload from measured effective ceilings, and argue the wider tradeoff between compacting often, isolating work in sub-agents, and designing sessions that rarely approach the ceiling at all.

## Measure against the usable ceiling, not the label The first mistake is anchoring the threshold to the advertised window. Long-context evaluations agree that effective capacity — where the model still reliably attends to material anywhere in the window — falls well short of the advertised number, and that degradation is cliff-shaped rather than gradual. A policy that only fires at 95% of a very large advertised window means the agent spends most of the session operating in the degraded region, quietly planning worse, while the operator sees nothing but slightly odd behaviour. Set the high-water mark against the ceiling you actually measured for your model and workload. ## Two thresholds, not one **High water** decides *when*. **Low water** decides *how far*. Running only a high-water mark and trimming just enough to get back beneath it is the direct cause of thrash: at 95% you trim to 90%, three turns later you are at 95% again, and you compact perhaps ten times an hour. Each of those passes costs real money and real risk: - a model call to produce the summary, with its latency inserted into the loop; - **cache invalidation** — the rewrite changes the request prefix, so the next turn is a full-price uncached prefill of the whole rebuilt context, and the discount the session had been enjoying is gone; - **compounding drift** — the second compaction summarizes the first summary, the third summarizes that, and every pass is another lossy transform on the same facts. Constraints and decisions blur into vagueness faster than anyone expects. Compacting deep is what buys distance between events. Reducing to roughly a third of the effective ceiling typically means dozens of turns before the next trigger, so a long session pays for a handful of compactions rather than dozens. ## Headroom for the next call The threshold must account for what is about to happen, not only what has happened. If the agent is about to read a large file or fetch a long log, that result lands on top of current utilization, and the output tokens land on top of that. A policy that fires at 90% with a tool capable of returning a very large payload will overflow on the discovering turn. Practical form: trigger when current usage plus a reserve for the largest plausible result plus the maximum response size exceeds the high-water mark. ## Boundaries matter Compact only at a clean turn boundary, where every tool call issued has its matching result present. Rewriting the history between a call and its result leaves an orphaned call — an inconsistent transcript that providers may reject and that models certainly misread, typically by re-issuing the call. "Only at a boundary" is a hard constraint sitting above the numeric thresholds: if the mark is crossed mid-call, finish the call and compact after. ## Compose the cheap lever first At the trigger, clear stale tool results before deciding to compact, then re-measure. In tool-heavy sessions this alone often drops utilization below the mark, and it costs no model call. Age-based clearing (results older than N turns) and count-based clearing (keep the newest K) can run as a continuous background policy, with compaction reserved for the case where the *conversation* rather than its payloads has grown long. ## Sequence the durable write before the clear Whatever the session must not lose — running notes, plan state, facts destined for a store outside the window — is written **before** the history is removed. After the clear the model can no longer see the material it would be writing about, so a harness that clears first and then prompts "record anything important" is asking an unanswerable question, and the failure is silent: you get a plausible, empty note. ## Tuning signals Instrument compactions per session, tokens reclaimed per compaction, and turns between compactions. The last is the thrash detector — if it sits in the low single digits, either the high-water mark is too close to the ceiling or the low-water target is too shallow. Pair those with behavioural signals: repeated work, re-asked questions, violated standing constraints. Cost dashboards show it too, as a session whose uncached input tokens are a large fraction of the total, because every compaction restarted the cache. ## The judgement being tested An interviewer asking this wants to hear that you treat compaction as an event with a price rather than free housekeeping, that you know the price has three parts (call, cache, drift), and that your policy is shaped to make the event rare and clean rather than frequent and shallow.

  • How would you detect compaction thrash from telemetry alone?
    Track turns between compactions and the uncached share of input tokens. A session compacting every few turns, with most input billed uncached because each rewrite restarts the cache, is thrashing. Rising compactions per session without a rising session length is the same signal from another angle.
  • Why not just raise the threshold to 95% and get more work per compaction?
    Because the last stretch of a window is where quality falls off, and the decay is cliff-shaped rather than gradual. You buy a few extra turns of context at the price of planning quality across all of them, and you remove the headroom that keeps a large tool result from overflowing the very turn that discovers the problem.
  • Should the threshold be the same for every workload?
    No. A session whose tools return small structured payloads can run closer to the mark than one that may pull a large log at any moment, because the required headroom differs. Effective ceilings also differ by model. Derive the numbers from measured behaviour per model and workload rather than porting one constant everywhere.

saying these in an interview costs you the question

  • Setting the threshold against the advertised window rather than the usable ceiling
  • Trimming only just below the trigger, causing repeated compactions
  • Ignoring headroom for the next tool result and the response
  • Compacting between a tool call and its result
  • Treating compaction as free housekeeping with no cost to weigh

context