skip to content

Why do agents write the plan to an external todo file instead of keeping it in context?

level: juniorimportance: must knowfreq 62%

answer

  1. the conversation is not durable storage
  2. addressed, not remembered
  3. one artifact, three jobs
  4. pointers to results, never payloads
  5. stale is worse than absent

basics

~20 s

A written plan survives what the conversation does not. Context gets truncated, summarized or crowded out on long runs, so an external plan file keeps the goal and per-step status stable and re-readable, and doubles as an audit trail of what the agent actually did.

solid answer

~50 s

On a long run, the conversation is not durable storage: the window fills, earlier turns get compacted or cleared, and the original goal is exactly the material most likely to be paraphrased away. Writing the plan to a file — a `plan.md` with one line per subtask and a status marker, rewritten as each item completes — makes it re-readable at any point without depending on what is still in the window. That gives you three things at once. It is working memory: the agent re-reads the file to answer "what is done and what is next". It is a handoff surface: a fresh session, or a person taking over, can reconstruct state from the file alone. And it is an audit trail: the sequence of edits shows what the agent believed and when. The discipline that matters is rewriting it as work completes, not writing it once — a stale plan file is worse than none, because it is confidently wrong.

code

markdown · 9 lines
markdown
# Goal
Produce a diligence memo on the target's supplier obligations.

## Plan
- [x] Collect data-room documents -> raw/ (12,041 files)
- [x] Dedupe near-identical versions -> clean/ (9,318 files)
- [ ] Privilege-screen clean/ -> flags/privileged.csv
- [ ] Summarize contract families -> summaries/
- [ ] Assemble memo -> memo.md

go deeper

for a junior

Be ready to say the conversation window is finite and lossy on long runs, and that a written plan with per-item status can be re-read at any point regardless of what is still in context.

for a middle

Explain the three jobs the file does — working memory, handoff state, audit trail — and the discipline that makes it work: update on completion, store pointers to artifacts rather than the artifacts themselves.

for a senior

Show operational judgment: name the stale-plan failure mode, explain why a plan file proves intent rather than correctness, and describe what you would put in it so an incident reviewer can reconstruct the run.

for a principal

Own it as a policy: an externalized plan is what makes agent work auditable and resumable across sessions, so its format and retention are organizational decisions, not per-agent improvisation.

## The problem the file solves An agent's plan has to remain available for as long as the run lasts. On a short run, the conversation itself is fine — the plan was stated a few turns ago and is still visible. On a long run it stops being fine, for a mundane reason: the context window is finite and the run is not. As turns accumulate, older material is compacted into summaries, cleared, or simply pushed past the point where the model attends to it reliably. The plan stated at turn 3 is precisely the kind of content that gets summarized into a sentence by turn 80 — and a paraphrase of a plan is not a plan, because it loses the per-item status that made it useful. Writing the plan to a file removes the dependency on what is still in the window. The file is addressed, not remembered: the agent re-reads it when it needs it, and the cost is one small read rather than carrying the whole thing forward in every turn. ## Three jobs, one artifact **Working memory.** The file answers "what has been done, what is in progress, what is next" without the agent reconstructing it from the transcript. Reconstruction from a long transcript is exactly where agents invent completed steps or repeat finished ones. **Handoff surface.** Because the state lives outside the conversation, a new session — or a human — can pick the work up by reading the file. This is what makes long-horizon work resumable at all: the run's progress is not trapped inside one conversation's history. **Audit surface.** The file's successive versions record what the agent believed the remaining work was at each point. When something goes wrong, that history is usually more informative than the raw transcript, because it is already at the level of tasks rather than tokens. For work with any review or compliance angle — a diligence review, a change to production systems — being able to show the plan as it evolved is often a requirement rather than a nicety. ## What goes in it Keep it small and status-bearing. One line per subtask, each with an unambiguous scope, a status marker, and — where it matters — a pointer to the artifact the step produced. Pointers rather than payloads is the key discipline: the plan says `findings written to findings/auto-renewal.csv`, not the contents of that file. The plan is an index of work, not a store of results; letting tool output leak into it defeats the whole point by making it too large to re-read cheaply. The goal statement belongs at the top, verbatim. It is the single most valuable line in the file, because it is the thing an agent is most likely to lose on a long horizon and the thing everything else is judged against. ## Rewrite discipline The file only works if it tracks reality. That means updating it at completion boundaries — mark the item done, note the artifact, and only then move on — rather than writing it once at the start and treating it as a record of intent. A plan file that says three items remain when five actually do is worse than no file, because the agent trusts it: it is now confidently wrong, and it will act on that. The converse failure is over-writing: rewriting the whole file every turn burns tokens and invites the model to quietly reword items, which drifts scope. Update on completion, not on every thought. ## Limits An external plan is not a correctness mechanism. It does not make the decomposition good, it does not verify that a step marked done actually succeeded — the agent writes both the work and the checkmark — and it does not stop the agent from doing something the plan never mentioned. It buys durability and legibility, which are prerequisites for the other mechanisms rather than substitutes for them. Interviewers reward candidates who say that plainly: the file makes state inspectable, and inspectability is what everything else builds on. ## Interview framing The expected answer names the finite, lossy nature of the conversation as the driver, and then names more than one benefit — memory, handoff and audit — rather than only "the model forgets". Mentioning the pointers-not-payloads discipline and the stale-file failure mode is what distinguishes someone who has run long agent sessions from someone who has read about them. This reflects practice as of mid-2026, where externalized notes and plans are a standard part of context engineering rather than an optional flourish.

  • What belongs in the plan file, and what should stay out of it?
    In: the goal verbatim, one line per subtask with an unambiguous scope, a status marker, and a pointer to whatever artifact the step produced. Out: tool output, retrieved documents, long reasoning. The file is re-read often, so it has to stay cheap to read; the moment results are pasted into it rather than referenced, it grows past the point where re-reading it is worthwhile and it stops being used.
  • How do you keep the plan file from drifting out of sync with what actually happened?
    Tie the write to the completion boundary: a step is not finished until its status and artifact pointer are written, so the update is part of the step rather than an afterthought. Avoid rewriting the whole file every turn — that costs tokens and lets the model quietly reword items, which drifts scope. Where possible, have the item's status set from the verification result rather than from the model's own assertion.
  • Does an external plan file make the agent's work more reliable on its own?
    No. It makes state durable and inspectable, which is a prerequisite for reliability rather than reliability itself. The agent still writes both the work and the checkmark, so an item marked done proves intent, not success. Pair the file with an independent check — a test, a schema, a reviewer — if you want the status column to mean anything stronger than 'the agent believes this is finished'.

saying these in an interview costs you the question

  • Says the context window reliably retains the original plan
  • Writes the plan once and never updates item status
  • Pastes full tool output into the plan file
  • Treats the file as documentation rather than the agent's live state
  • Claims a plan file guarantees the steps were actually done correctly

context