skip to content

Why does an LLM assistant forget facts between sessions, and how do you fix it?

level: juniorimportance: must knowfreq 70%

answer

  1. the model keeps nothing between calls
  2. in-session recall is replay
  3. new session, new window
  4. facts must be written outside the window
  5. write, store, then re-inject

basics

~20 s

Everything the model can use in a turn is text in its context window, and that window is rebuilt from scratch for each new session. A fact survives only if the application writes it to a store outside the window and puts it back in later.

solid answer

~40 s

The model itself is stateless: each request sends a fresh token sequence and the model retains nothing afterwards. Inside a session it *appears* to remember because the application replays earlier turns on every request — that is session history, and it dies when the session ends or when old turns are dropped to fit the budget. Durable memory is a separate thing you build: an explicit **write** of a fact to a store outside the window (commonly a file-backed store the model can edit, or a database), plus a **read** that re-injects it into a later session's context. A personal-finance assistant only stops asking "do you file jointly?" every month if "files jointly" was written to a user-scoped store and re-injected. Without that write step, a bigger context window changes nothing.

go deeper

for a junior

Be able to say plainly that the model keeps nothing between calls, that in-session recall is the application resending earlier turns, and that cross-session memory needs an explicit store outside the window.

for a middle

Explain the three moving parts — write, store, re-inject — and why a missing read path silently kills the feature. Be ready to say why a larger context window is not memory.

for a senior

Show judgement about what deserves persisting: stable, decision-relevant facts, not every utterance. Talk about how you would detect a write-only store in production and what a stale persisted fact costs on later turns.

for a principal

Own the policy: what the product persists by default, who it is scoped to, how a user deletes it, and why per-user fine-tuning is the wrong instrument for facts that change and must be erasable on request.

## The model is stateless A large language model does not accumulate state across calls. Every request is a complete token sequence — system instructions, prior turns, retrieved documents, tool results — and the model produces the next tokens conditioned only on that sequence. When the response ends, nothing about it is retained inside the model. The weights are the same before and after the conversation. So when people say an assistant "remembers", they are describing something the *application* does, not something the model does. ## Session history is replay, not memory Inside one chat session, the harness keeps a list of turns and resends them on every request. That is why turn 12 can refer to something said in turn 3: turn 3 is literally still in the prompt. This is often called working or short-term memory, and it has two hard properties worth stating in an interview: - **It is bounded.** The window has a token limit, and long before that limit is reached, quality degrades — the widely used term is *context rot*, where effective usable context is well below the advertised size. Something has to be dropped or condensed eventually. - **It is scoped to the session.** When the user starts a new conversation, the harness builds a new list. Nothing carries over unless something else carried it. ## Durable memory is a store outside the window Durable memory is an explicit engineering construct with three parts: 1. **A write.** Some component decides that a fact is worth keeping and persists it — to a file the agent can edit, to a memory tool backed by a directory, or to a database. Providers now ship file-backed memory tools for exactly this, but the concept predates and outlives any one product. 2. **A store.** The persisted facts live outside the context window. They cost nothing while they sit there. They can be inspected, edited and deleted, which matters for privacy and correctness in a way that weights never can. 3. **A read.** On a later turn or session, something puts the relevant subset back into the window. Nothing in the store affects the model until it is re-injected as tokens. Missing any one of the three gives you a system that looks like it has memory and does not. The most common broken version is a store that is written diligently and never read back. ## A worked example A personal-finance assistant is asked in January to estimate a tax refund. The user says "we file jointly." In February the user opens a new session and asks about a deduction. Filing status is a stable, decision-relevant fact about the user, so at the end of the January session it should have been written to a user-scoped store; in February the assistant loads it and answers without re-asking. Contrast that with "the user asked about deductions on 3 February" — an event record, useful for a history log but usually not worth spending tokens on in every future turn. Practitioners split durable memory into tiers — episodic event records, semantic facts, and procedural know-how — with different write and read policies for each. The important point at this level is simply that the tier decision is a design choice about what deserves to be persisted and re-injected, not a property the model provides. ## What durable memory is not - **Not fine-tuning.** Training on conversations changes general behaviour, not per-user facts. It is far too slow and expensive per fact, and you cannot delete one fact from a set of weights on request. - **Not a bigger window.** Million-token windows help you fit more into *one* request. They do not carry anything into the *next* request, and pushing utilization high enough is exactly where quality degrades. - **Not automatic.** A chat product that appears to remember has a memory feature built into it. A raw model API has none. ## Failure modes to name - **Write-everything.** Persisting every utterance produces a store full of low-value and contradictory entries that later poison the context. - **Write-only.** Facts are saved but never re-injected, so behaviour never changes and nobody notices the feature is dead. - **Session junk promoted to permanent.** Transient state ("the file I'm editing right now") gets written as if it were a durable preference, and months later the agent still believes it. The crisp interview answer: session history is transient replay of the current conversation; durable memory is an explicit store outside the window with a deliberate write path, a read path, and a lifecycle.

  • If the whole conversation is already visible to the model, why write anything to a durable store mid-session?
    Because the window is finite and the session is not permanent. Long sessions get trimmed or condensed to fit the budget, so a fact stated early can be evicted before it is needed. Writing it out at the moment it appears makes it survivable independently of what happens to the transcript, and makes it available to a future session that has no access to this transcript at all.
  • Could you just fine-tune the model on each user's conversations instead of building a memory store?
    No, for three reasons. Weights encode general behaviour, not reliable per-fact recall, so a single stated preference may not survive training at all. The turnaround and cost per user make it impractical for facts that change weekly. And you cannot delete one fact from a weight set on request, which breaks correction and privacy obligations that a store handles with a single row deletion.
  • How would you tell, in production, that the memory feature is not actually working?
    Instrument both halves. Count writes and count how often stored content actually reaches a prompt; a store growing while injections stay near zero means the read path is broken. Then look at behaviour: if the assistant re-asks a question whose answer is in the store, that is a concrete, reproducible failure you can turn into a regression case.

The model is like a consultant who reads a briefing packet before every meeting and remembers nothing once they leave. Durable memory is the filing cabinet the packet is assembled from.

saying these in an interview costs you the question

  • Thinks the conversation updates the model's weights
  • Assumes a million-token window removes the need for a store
  • Believes memory is automatic behaviour of any model API
  • Writes facts to a store but never re-injects them
  • Cannot distinguish session history from persisted memory

context