Why does one wrong assumption early in an agent session make the turns that follow more confident rather than less?
answer
- The agent reads its own output back
- Nothing marks which lines were checked
- Qualifiers disappear when a claim is restated
- Its edits remove the evidence it would find
- Confidence tracks repetition, not verification
basics
~20 sAn agent's own earlier conclusions sit in the session record, and later turns read them as given rather than as guesses. Each restatement drops the hedge the first one carried, so certainty grows while evidence does not.
solid answer
~40 sA session is a record, and the agent wrote most of it. A turn that says `I searched and did not find other readers of this field` becomes, several turns later, `this field has no other readers` — the hedge rarely survives restatement, and nothing in the record marks which lines were checked and which were guessed. The session has also been editing the code the whole time, so a later search for the old name comes back empty because this run removed every occurrence it could see. Agreement accumulates inside the session while verification does not, which is what rising confidence is made of. The premises worth checking yourself are the ones the run decided rather than the ones you supplied.
code
text · 13 linesturn 3 "I searched the files I could read and did not find
other readers of the status field."
turn 6 "Since status has no other readers, I renamed it in
the model and updated the callers."
turn 11 "Searched for the old name: no occurrences left.
The rename is complete."
turn 15 "No migration is needed here — the field was internal."
# the claim never got stronger; the wording did
# turn 11's search was empty because turn 6 emptied itgo deeper
Recall the shape: the agent reads its own earlier statements back as facts, so an early guess quietly becomes a premise. Name one thing you would verify yourself rather than take from the session.
Explain the mechanism: no provenance in the record, hedges lost in restatement, and a run whose own edits erase the traces a later search would find. Say why that makes confidence rise.
Show the judgment call — correct in place while the premise is cheap, or abandon and restart with the fact stated up front once it is in the code. Say what each choice costs you.
Own the framing: a long run's certainty is an artefact of its record, not a measurement of the codebase, so the premises a change rests on belong in the open where somebody can check them.
## A session is a record, and the agent wrote most of it When an agent works across many turns, what it reasons from next is not your codebase directly. It is the running record of the session: your request, the files it has read, the edits it has made, and **everything it has previously said**. That last category is the trouble. A conclusion the run reached on turn three is, by turn twenty, just a line in the record, sitting beside the lines you wrote and the file contents it read, and **nothing in it is stamped with where it came from.** So `I searched the files I could read and did not find other readers of this field` and `this field has no other readers` end up doing much the same work on a later turn, even though the first is a report about a search and the second is a claim about the world. **Context poisoning** is the usual name for the wrong item entering the session's working record; **context drift** is the usual name for what the session then does with it, as later turns inherit it and build on it. It is unlike an ordinary mistake in a single edit, because the wrong item is not in the output where you would review it — it is in the material that produces all the later output. ## The hedge decays with each restatement Long sessions summarise themselves constantly: a turn refers back to what was decided, a plan is restated, a change is described before it is made. Qualifiers tend to be the first words a summary drops. The claim survives; the uncertainty attached to it does not. | what the early turn actually established | what later turns treat as established | |---|---| | a search over the files it could read found no other readers | the field has no other readers | | the tests that already existed still passed after the rename | nothing outside the code depends on the old name | | no caller inside this repository used the old name | no stored data, configuration or outside consumer uses it | Each right-hand side is a larger claim than its left, and the step between them was taken by restatement rather than evidence. ## The run edits the evidence it would check against Suppose the run renames the field on turn five. On turn eleven, being careful, it searches for the old name to confirm nothing is left behind. The search comes back empty — largely **because the run itself removed every occurrence it could see.** What it still cannot see, such as a column in stored data or a key in a configuration file it never opened, is exactly what it could not see on turn three. The confirmation therefore adds close to no information, and it reads like confirmation. That is the shape of the whole failure: agreement accumulates inside the session while verification does not. ## Why a late correction lands badly Telling the run on turn twenty that the premise was wrong helps less than it should, for reasons worth separating: - The earlier line is still in the record. Your correction sits after it; both are read. - The premise is no longer only in the record. It is in the code the session has written, and the code agrees with it. - A run that has produced a lot of work on a premise has a lot of material agreeing with it, and agreement is what it reasons from. - The correction has to be applied backwards through work built on the wrong footing, which is a bigger request than the one that caused the problem. So a correction lands better the earlier it is made, and once the premise has produced many edits, a fresh session with the corrected fact stated at the top is often cheaper than arguing inside the poisoned one. That trade costs the session's accumulated useful work, which is why it is worth deciding deliberately rather than drifting into. ## What to check, and when 1. **Separate what the run was told from what the run decided.** You supplied some facts; it inferred others. Both kinds can be wrong, and the decided ones are the ones nothing outside the session ever checked. 2. **Check the decided ones yourself, outside the session,** against the thing they are claims about — the stored data, the configuration, the consumers you know of. 3. **Do it early, while the premise is cheap.** A premise checked on turn four costs one question; the same premise checked on turn thirty costs whatever was built on it. 4. **Read a confident restatement as a reason to check, not as evidence.** A confident restatement tells you the claim has been repeated; it does not tell you it has been supported. ## What this is not This is not the context window filling up — how much a session can hold is a separate subject, and a session can be poisoned while well inside its capacity. It is also not an argument that assumptions are avoidable: a run that assumed nothing would ask about everything and get nothing done. The practical question is narrower: which assumptions is this change resting on, and which of those have you confirmed with something other than the session's own say-so.
- How would you tell a poisoned premise from the agent simply being wrong on one turn?A one-turn mistake is local: the edit is wrong and the turns around it are unaffected. A poisoned premise shows up as several unrelated-looking edits that are all consistent with one thing the run believes, including edits that look like tidy-up. If correcting one edit implies correcting the others, you are looking at a premise.
- You spot the wrong premise on turn four rather than turn thirty. Does that change what you do?Yes, and it is the whole reason for looking early. On turn four almost nothing rests on it, so stating the correct fact and continuing is usually enough. By turn thirty the premise is in the code as well as the record, and a fresh session with the fact stated up front is often the cheaper route.
- Can you ask the agent which of its statements were verified?You can ask, and you will get an answer, but the answer is produced from the same record that contains the claim. It is a restatement, not a provenance trail. Treat it as a place to start looking rather than as the check itself, and confirm the premise against the code, the data or the configuration it is a claim about.
A rumour grows more certain with each retelling, not because anyone learned anything new, but because each teller drops the hedge the last one used. A long session does the same thing to its own early guesses.
saying these in an interview costs you the question
- Telling the agent it was wrong removes the wrong assumption from the session
- How confidently the agent states something tracks how well supported it is
- A bad assumption will surface as a failing build before it spreads
- Asking the agent to double-check its own earlier conclusion settles it
- A bigger context window would prevent this
- Drift is just the model being unreliable, so nothing you do about the session matters