skip to content

In a long agent coding session, what fills the context window, and why does the run degrade rather than stop?

level: middleimportance: should knowfreq 52%

answer

  1. the room per turn is fixed
  2. who spends it, not how big it is
  3. your request stays one paragraph
  4. it narrows rather than halts

basics

~20 s

A long session is filled mostly by its own by-products: files read entire, full command output, failed attempts, its own restatements. The room is finite, so older material goes and the work carries on with less.

solid answer

~50 s

Your request is a paragraph you write once and it stays one; the session's room is consumed by what the work produces. Reading a file to reach one function brings the whole file; a test run brings its whole output, green or not; attempts that did not work stay in the record beside the one that did, as do the run's own plans and restatements. Each turn has to be re-assembled under a fixed ceiling, so once the material exceeds it something has to leave — condensed, dropped or re-read later, depending on the tool. **Nothing fails when that happens**, which is the part worth planning around: you get a run that continues with less rather than an error. The symptoms are re-reading a file it read an hour ago, re-introducing something it removed, and quietly dropping a constraint from the opening request.

code

text · 16 lines
text
WHAT THE SESSION IS CARRYING AT TURN 40  (reporting-migration run)

  your opening request .......... one paragraph, sent once, at turn 1
  files read in full ............ 31, to reach perhaps 40 functions
                                  9 of them read a second time later
  suite output .................. 14 runs, whole output each time,
                                  green runs included
  attempts that did not work .... 6, still in the record next to the one
                                  that did, nothing marking which is live
  the run's own turns ........... 40 plans, restatements and diffs
                                  - the largest category by far

  over the ceiling .............. the oldest material leaves: condensed,
                                  dropped or re-fetched later, depending
                                  on the tool
  what does NOT happen .......... the run stopping, or saying so

go deeper

for a junior

Know that a session's room is finite and that the work itself fills it — files read whole, command output, the run's own turns — far more than your opening request does.

for a middle

Be able to say what happens at the ceiling: older material is condensed or dropped and the work continues, so the failure is silent and shows up as re-reading, re-introducing and dropped constraints.

for a senior

Connect it to the decision you control. Unit size sets the conditions a unit's own final turns run under, and the end of a long unit is both the most degraded part and the least carefully reviewed.

for a principal

The tradeoff worth owning is where a team puts durable facts. Anything a session must not lose belongs in the repository as code, a test or a note, because transcripts are not storage and no window size makes them so.

## The resource nobody budgets Ask what limits a long agent session on a migration and most answers name the model's ability or the size of the codebase. The limit that actually binds is simpler: **the room available to each turn is fixed, and the work spends it.** Every turn is assembled from the material the session is carrying, and that assembly has a ceiling. Under the ceiling, everything is present. Over it, something has to go. What makes this a scoping subject rather than a curiosity is where the spending happens — almost none of it is yours. ## What actually fills it Over a run that moves a ticketing system's reporting module off a date library, the session accumulates roughly this: - **Whole files, read to find one function.** A search hits a file; the file arrives entire. The reporting module's largest file is mostly not about dates and is carried anyway. - **Command output, in full.** A suite run brings its whole output, green included. Fourteen runs bring fourteen copies of much the same text. - **Attempts that did not work.** The branch that failed to compile is still in the record beside the one that did, and nothing marks which is live. - **The run's own turns.** Plans, restatements, descriptions of the change about to be made, summaries of what was done. On a long run this is the largest single category. Your opening request is a paragraph, and it stays one however many turns carry it. Set against forty turns of the above, it is a rounding error — which is why *how much you hand over at the start* is a real question but a different one from this. What you chose is small and fixed; what the work generates is neither. ## Why it degrades instead of stopping When the material exceeds what a turn can hold, something leaves. Whether the tool condenses the oldest turns, drops them, or fetches material again later, one thing is consistent: **the work does not stop.** No check fails, no error is raised about it, and the run answers as well as it can from what it still has. That is the whole difficulty. A limit that stopped the run would be a cheap limit — you would see it, resize the unit and re-run. A limit that quietly narrows what the run is working from produces confident output built on less, and the only evidence is behavioural: | what you see | what it usually means | |---|---| | it re-reads a file it read an hour ago | that content is no longer in front of it | | it re-introduces a helper it removed earlier | the removal left the record | | a constraint from the opening request stops being honoured | the opening turn is no longer carried verbatim | | it describes the plan more and edits less | restatement is filling the space the work used to | What happens inside the model as material accumulates is a separate subject with its own answers. The scoping fact is the one above, and it is enough to act on. ## Why this is a scoping question and not a tooling one If capacity is spent by the work, then **the size of a unit decides the conditions its own last turns run under.** A unit that consumes most of the session's room will finish under worse conditions than it started, and the part of the work that lands in that state is the end — the wiring-up, the cleanup, the last report — which is also the part nobody reviews as carefully. Three consequences follow for how you cut work: 1. **Size a unit to finish comfortably**, not to finish at all. The margin is the point. 2. **Prefer ending a unit to continuing into a degraded session.** A fresh session starts with room and without the record of forty turns, which is a genuine trade rather than a free win: it also starts without what the previous session learned, and re-establishing that costs you. 3. **Leave the durable facts in the repository.** A helper, a test that encodes the rule, a short note in the change — material the next session reads from the code rather than inherits from a transcript. ## The honest limits of this answer None of it says a long session is worthless, and none of it says a short one is safe. A long run over a bounded unit with a real check at the end is fine, and a short run whose result nobody can judge is not. The claim is narrower: capacity is consumed mostly by the work, a run that runs out narrows rather than halts, and unit size is the lever you actually hold. ## What a good answer sounds like Name the spenders first — files read whole, command output, dead attempts, the run's own turns — then say the limit does not announce itself, then connect it to unit size. An answer that stops at "the context window fills up" has named the resource without saying who spends it, which is the half that changes what you do next.

  • If a session is running out of room, why not simply start a fresh one?
    Often you should, at a unit boundary. But it is a trade rather than a free win: the new session starts with room and without forty turns of record, and also without what the previous one established. That is why the durable facts want to be in the repository — a helper, a test, a note in the change — so a fresh session reads them from the code.
  • Does a bigger context window remove the problem?
    It moves the point at which the squeeze starts; it does not change who is spending. Files read whole, command output, dead attempts and the run's own turns all grow with the length of the work, so a longer run refills whatever room it is given. Unit size stays the lever you hold directly.

A long meeting's whiteboard fills with working notes, and when room runs out someone wipes the earliest corner to keep going. The meeting carries on perfectly happily, and what was wiped is exactly what nobody remembers deciding.

saying these in an interview costs you the question

  • The session mostly holds what you gave it at the start
  • A run that hits its limit stops and tells you
  • Only the size of the codebase decides how much room is left
  • Re-reading a file it already read is just wasted effort
  • A bigger window would make unit size stop mattering