skip to content

Why is an AI agent's context window not the same thing as its memory?

level: middleimportance: must knowfreq 58%

answer

  1. stateless call, reassembled every time
  2. bounded, volatile, and degrading
  3. the doorway, not the room
  4. working memory can live in a file
  5. a bigger desk is not a filing cabinet

basics

~20 s

The context window is the input assembled for one model call — bounded, rebuilt by your code every time, and gone afterwards. Memory is durable state stored outside the model that the agent deliberately writes and reads back. Continuity comes from the store, not the window.

solid answer

~50 s

A model call is stateless: the application sends a complete set of messages and gets a completion back, and the model retains nothing between calls. Anything that looks like the agent "remembering" the previous turn is really the harness putting that content back into the next request. So the context window is the working set for one call, with a hard token ceiling, reassembled every time — not a place where anything is stored. That disqualifies it as memory on three counts: it is bounded, so long sessions must drop or compress content; it is volatile, so a crash, a new session or a fresh subagent starts empty; and its quality degrades as it fills, since models attend less reliably to material buried in the middle of a very long context. Working memory is a role that the context window usually plays, but a scratchpad file the agent writes and re-reads plays the same role while surviving a context rebuild.

code

python · 17 lines
python
def build_context(system_prompt, store, session_id, user_msg):
    facts = store["semantic"]
    recent = store["episodes"][session_id]
    return [
        {"role": "system", "content": system_prompt + "\n" + "\n".join(facts)},
        *recent,
        {"role": "user", "content": user_msg},
    ]


store = {
    "semantic": ["prefers metric units"],
    "episodes": {"s1": [{"role": "assistant", "content": "Ready."}]},
}

for message in build_context("You are a helper.", store, "s1", "Convert 5 miles."):
    print(message["role"], "->", message["content"])

go deeper

for a junior

Be able to say that each model call carries its whole input and the model keeps nothing between calls, so the application is what makes a conversation feel continuous.

for a middle

Explain the three disqualifying properties — bounded, volatile, quality-degrading — and separate the working-memory role from the context-window mechanism. Expect to be asked whether a bigger window solves it.

for a senior

Describe a real incident shape: a constraint dropped by truncation, or a subagent re-deriving a decision it never saw. Show how you keep the window small and high-signal by holding references and writing notes out to durable files.

for a principal

Own the cost and reliability argument. Every token is paid on every call, so window size is an operating-expense decision as much as a design one, and the durability story has to survive restarts, crashes and hand-offs between agents.

## A stateless call, restaged every turn Every request to a language model carries its entire input. The model does not hold a session; it reads the messages you send, produces a completion, and keeps nothing. The illusion of a continuing conversation is manufactured by the application, which stores the transcript and resends it — usually with the system prompt, tool definitions, retrieved documents and any notes — on the next call. Once you internalise that the harness rebuilds the input each time, the framing question answers itself: the context window is a per-call buffer, and a buffer that your code refills from somewhere else is not the place the information lives. ## Three properties that disqualify it as storage **It is bounded.** There is a hard token ceiling. A long session, a verbose tool result, or a large file read will hit it, and something has to give — truncation, dropping old turns, or summarizing them into a shorter form. Storage that silently evicts your data under load is not storage. **It is volatile.** Close the session, crash the process, hand work to a fresh subagent, or start tomorrow, and the window begins empty. Nothing carries over unless it was written somewhere durable first. **It degrades before it fills.** Quality is not flat across the window. Long contexts suffer what practitioners call context rot: material in the middle is attended to less reliably than material at the edges, and irrelevant or contradictory content actively hurts, because everything in the window is competing for attention. A fact that is technically present but ignored is functionally forgotten. ## What memory is instead Memory is durable state that lives outside the model call: files on disk, rows in a database, an append-only event log, an index the agent can query. Two things make it memory rather than incidental data. First, the agent or harness writes to it deliberately. Second, it is read back deliberately — something decides that this stored item is relevant now and lifts it into the next call's context. Memory only ever influences the model by passing through the context window; the window is the doorway, not the room. ## Working memory versus the context window These are not synonyms, and the distinction is where interviewers probe. Working memory is a role — the state of the task in progress. The context window is one mechanism that carries that role. A long-running agent frequently keeps working memory outside the window: it writes a plan and running notes to a scratchpad file, calls a tool, then re-reads the file to reorient. If a compaction or a restart wipes the conversation, the scratchpad survives and the agent picks up where it left off. That file is still working memory — same task, same session, discarded when the job ends — it simply is not in the window at every moment. ## The desk and the filing cabinet The context window is your desk: everything you are actively working with, fast to reach, and finite. Memory is the filing cabinet behind you: far larger, slower, and useless unless you know which drawer to open. You cannot solve a cluttered desk by buying a bigger desk forever, and a document that is on the desk but under a pile might as well be filed. ## What goes wrong when the two are conflated The common bug is a constraint stated early in a long conversation — "never touch the production database" — that gets dropped when the transcript is truncated, after which the agent cheerfully violates it. Another is a multi-session assistant that appears to know a user until the process restarts, then greets them as a stranger, because everything it "knew" only ever existed in a resent transcript. A third is a subagent launched with a fresh window that lacks a decision the parent made, and re-derives it differently. ## Does a bigger window fix it? No, and this is the follow-up worth being ready for. Larger windows help, but cost scales with tokens on every single call, retrieval quality inside a very full window degrades, and volatility is untouched — a million-token window still starts empty tomorrow. The 2026 practice runs the other way: keep the window small and high-signal, hold references such as file paths and record IDs rather than payloads and fetch just in time, write notes to durable files, and isolate expensive exploration in a subagent that returns a short summary. Summarize-and-restart techniques for a window that does fill are a subject of their own; the point here is that they are damage control for a buffer, not a substitute for a store.

  • If context windows keep growing, does external memory eventually become unnecessary?
    No. Size does not address volatility — a huge window still starts empty after a restart or in a fresh subagent. It also does not address cost, since every token is paid for on every call, or retrieval quality, which degrades as the window fills with lower-signal material. Bigger windows raise the ceiling on a single task; they do not give the agent continuity.
  • Where does a scratchpad file the agent writes mid-task sit in this picture?
    It is working memory that happens to live outside the window. Same task, same session, discarded when the job ends — but because it is on disk it survives truncation, compaction and restarts. That is precisely why long-horizon agents externalize their plan and notes: the role is unchanged, the durability is not.
  • How would you prove to a sceptical reviewer that the model itself remembers nothing?
    Send the same request twice from two independent processes and show that identical inputs give equivalent behaviour with no carry-over, then drop an earlier turn from the message list and watch the agent lose that information entirely. Continuity tracks exactly what the application resends, which is the observable signature of statelessness.

The context window is the desk you are working at — fast, finite, cleared at the end of the day. Memory is the filing cabinet behind it: much larger, slower, and only useful if you know which drawer to open.

saying these in an interview costs you the question

  • Says the model remembers earlier turns even when they are not resent
  • Claims a large enough context window removes the need for memory
  • Equates working memory with the context window as if they were synonyms
  • Assumes content present in a long context is reliably used by the model
  • Treats resending the transcript as persistence across sessions

context