skip to content

Agent Memory

What an agent remembers between steps and between sessions: the in-context scratchpad, external stores it retrieves from, and the compression that keeps both from overflowing. Interviewers ask what you write, when you read it back, and how you stop stale memories from poisoning later decisions.

part ofAI agentsoverview, primer and where to startread it →
on this pageshow

questions

15

In an AI agent, what are working, episodic, semantic and procedural memory?

level: juniorimportance: must knowfreq 70%

answer

  1. four labels borrowed from cognitive science
  2. now, what happened, what is true, how to
  3. events versus facts is the key split
  4. procedural usually means code, not prose
  5. working memory dies with the session

basics

~20 s

Working memory holds what the agent is using right now for the current task. Episodic memory records what happened in past sessions. Semantic memory stores durable facts. Procedural memory holds learned how-to — reusable skills the agent can apply again.

solid answer

~50 s

The four labels are borrowed from cognitive science and each maps to a different engineering problem. **Working memory** is the current task's state — the latest message, the plan so far, tool results, scratch notes — scoped to this session and thrown away when it ends. **Episodic memory** is an ordered, timestamped record of what happened: past sessions, actions taken, outcomes. **Semantic memory** is the distilled durable facts, statements that are true independent of when they were learned. **Procedural memory** is reusable know-how, usually stored as code or skill files rather than prose. A tutoring agent shows all four: this session's turns are working memory, the log of the previous fourteen sessions is episodic, "struggles with fraction division" is semantic, and "how to run a Socratic hint sequence" is procedural. They differ in lifetime, access pattern and failure mode, which is why real systems rarely put them all in one undifferentiated store.

go deeper

for a junior

Be able to name all four types and give one concrete example of each without hesitating. Say plainly that working memory is this task, episodic is what happened, semantic is what is true, procedural is how to do it.

for a middle

Explain what makes the types different beyond the labels: lifetime, access pattern, and whether the data is a sequence or a set. Be ready to classify an ambiguous item, such as a session summary, and justify the call.

for a senior

Show you would not build all four unless the product needs them, and describe the failure each type introduces in production — overflow, unbounded logs, stale facts, unsafe skills. Interviewers listen for whether you have actually had to debug a wrong memory.

for a principal

Own the position that the taxonomy is vocabulary, not architecture. Be ready to argue when one substrate serving all four is the right call versus four purpose-built subsystems, and what that choice costs in review burden, migration risk and vendor lock-in.

## Where the vocabulary comes from Agent memory terminology is borrowed from cognitive psychology, which separates the small amount of information a person is holding in mind right now from the much larger store they can recall later, and then splits that long-term store into memory for events, memory for facts, and memory for skills. Agent frameworks adopted the same four labels because they line up with four distinct engineering problems: what to put into the next model call, what history to keep, what durable knowledge to accumulate, and what capability to reuse. The biological analogy is loose — nothing in a language model works like a hippocampus — but the vocabulary is now standard, and an interviewer expects you to use it precisely. ## Working memory Working memory is the state of the task in progress: the user's latest message, the plan, results returned by tools, intermediate notes. Most of it sits in the context window assembled for the current model call, but not all of it needs to. An agent that writes a scratchpad file partway through a long job and re-reads it after each tool result is also using working memory — the file is simply working memory that survives a context rebuild. The defining property is scope, not location: working memory belongs to one task or session and is discarded when that ends. Nothing in it persists unless something deliberately writes it somewhere durable. ## Episodic memory Episodic memory is the record of what happened, in order, with timestamps: on this date the user asked for X, the agent called tool Y, the call failed, the agent retried with different arguments and succeeded. It answers "what did we do last time?" Its natural shape is append-only and time-ordered — you replay or summarize episodes rather than edit them. Raw episodes are verbose, so production systems usually keep a condensed summary per session alongside, or instead of, the full transcript. ## Semantic memory Semantic memory holds durable facts, either distilled from episodes or supplied directly: this customer is on the enterprise plan, the deploy script lives at a particular path, this student struggles with fraction division. The shape is a set of statements rather than a sequence of events. Facts can still expire or be superseded, but they are not inherently tied to the moment they were observed. This is the tier that gets updated and deduplicated, because two different episodes can yield the same fact — or contradictory ones. ## Procedural memory Procedural memory is know-how: a routine the agent has learned or been given and can invoke again. In practice it is stored as executable code, scripts, or structured skill documents rather than as free prose, because a procedure can be tested and a paragraph cannot. The Voyager Minecraft agent is the canonical illustration — it wrote programs, verified them against the environment, kept the ones that worked in a growing skill library, and composed them into harder tasks later. ## Why the split matters The four types have different lifetimes (one task, one session, indefinite, versioned), different access patterns (rebuild every call, replay in order, look up by similarity or key, fetch by name), and different failure modes (overflow, unbounded growth, staleness and contradiction, broken or unsafe skills). Collapsing them into a single vector store of text chunks is the most common design mistake: you end up running similarity search over an event log and retrieving a fragment of an old transcript when what the agent needed was one current fact. ## What the taxonomy does not give you It is a naming scheme, not an architecture. It says nothing about when to write, what to retrieve, or how to decide something has gone stale — those are separate policy questions with their own tradeoffs. Some systems deliberately implement all four in one substrate, such as a directory of files, and let the folder layout carry the distinction. Naming the four types correctly is a warm-up; the interview usually moves quickly to which types your system actually needs and what happens when one of them is wrong. ## Common confusions worth pre-empting Retrieval over a static document corpus is not agent memory in this sense — it is external knowledge the agent reads, not experience it formed and can update. A rolling conversation summary is compressed episodic memory, not semantic memory, until the facts inside it are extracted as standalone statements. And working memory is not "the small memory": it is often the largest thing in play by token count, just the shortest-lived.

  • Is retrieval over a company document corpus a form of semantic memory?
    Usually no. Retrieval-augmented generation over a static corpus is external knowledge the agent reads; agent memory is state the agent formed from its own experience and can update. The line blurs when the agent writes back into that corpus, but the useful distinction is authorship and mutability: memory is agent-written and revisable, a knowledge base is curated for it.
  • Where does a rolling conversation summary sit in this taxonomy?
    It is compressed episodic memory — a shorter account of what happened, still tied to a sequence of turns. It becomes semantic memory only when standalone facts are lifted out of it and stored as statements in their own right. Treating a summary as if it were a fact store is a common source of stale, un-updatable claims.
  • Which of the four can a simple single-session assistant skip entirely?
    All but working memory. If nothing needs to survive the session, the conversation itself is the memory and adding episodic, semantic or procedural stores buys complexity with no payoff. The types earn their keep the moment a user returns and expects continuity, or the agent repeats a procedure often enough that re-deriving it is wasteful.

Working memory is what is spread across your desk for today's job; the other three are the filing cabinet behind you — the diary of past jobs, the folder of facts, and the binder of procedures.

saying these in an interview costs you the question

  • Calls every long-term store "the vector database" regardless of memory type
  • Treats a conversation summary as semantic memory rather than compressed episodes
  • Believes working memory persists between sessions on its own
  • Says procedural memory just means storing prompt templates
  • Describes the taxonomy as an architecture rather than a naming scheme

context

open as a page

When is agentic memory recall better than pre-injecting top-k memories in an agent?

level: middleimportance: must knowfreq 62%

basics

~20 s

Agentic recall — the agent calling a memory tool when it notices it needs a fact — wins when most turns need no memory and precision matters. Pre-injected top-k wins for a small always-relevant profile and tight latency budgets.

open as a page

Why is an AI agent's context window not the same thing as its memory?

level: middleimportance: must knowfreq 58%

basics

~20 s

The context window is the input assembled for one model call — bounded, rebuilt by your code every time, and gone afterwards. Memory is durable state stored outside the model that the agent deliberately writes and reads back. Continuity comes from the store, not the window.

open as a page

When should an agent write to long-term memory: per turn, at session end, or on an explicit remember tool?

level: middleimportance: must knowfreq 62%

basics

~20 s

Write triggers trade freshness for cost. Per-turn extraction catches everything but adds a model call to every turn. End-of-session extraction is cheaper and better informed. An explicit remember tool is precise but captures only what someone thought to flag.

open as a page

How do scope and metadata filters on agent memory reads prevent cross-user leakage?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Every memory read must be constrained by a scope key — user, tenant, session or project — derived from the authenticated session, applied inside the query before ranking. Semantic search has no notion of ownership, so nothing but an explicit filter keeps one user's memories out of another's context.

open as a page

When a new fact contradicts a stored agent memory, how should the write path resolve it?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Detect the conflict before inserting, then choose deliberately: overwrite (last-writer-wins) is simple but destroys history and breaks on out-of-order input. Superseding keeps the old record, marks it replaced, and records both when the fact was stated and when it was written.

open as a page

Why do agent memory retrievers score recency and importance alongside relevance?

level: middleimportance: should knowfreq 54%

basics

~20 s

Semantic similarity alone surfaces old, trivial memories that merely resemble the query. Adding a recency term favours what the agent learned lately, and an importance term favours consequential facts over small talk, so the top few slots go to memories that change the answer.

open as a page

Which backing stores fit working, episodic, semantic and procedural agent memory?

level: middleimportance: should knowfreq 44%

basics

~20 s

Access pattern picks the store. Working memory suits a fast session-keyed cache or a scratchpad file; episodic memory suits an append-only, time-ordered event log; semantic memory suits a row store with a vector index for fuzzy recall; procedural memory suits versioned skill files in source control.

open as a page

How do you stop an agent's memory store filling with near-duplicate facts?

level: middleimportance: should knowfreq 45%

basics

~20 s

Look before you insert. Normalize each candidate fact into a canonical key, search the store for that key and for semantically similar records, then merge instead of adding: one record, one claim, with a link to every source that corroborated it.

open as a page

How should agent memory retrieval handle facts that were only true for a period?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Store a validity period on each fact and filter at read time against the moment being reasoned about, so superseded facts never enter context. Ranking cannot fix this: an expired fact can be the most semantically relevant one in the store.

open as a page

Why is agent procedural memory usually stored as code or skill files, not prose?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Procedures need to be verified, repeated exactly, reviewed and versioned. Code and structured skill files give all four — you can test them, diff them and roll them back. Prose "lessons learned" cannot be tested, accumulate contradictions, and are recalled fuzzily rather than executed.

open as a page

When should an agent memory be scoped per-session, per-user, or globally?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Default to the narrowest scope that makes the memory useful: session for task state, user for anything personal or account-specific, global only for knowledge that is true for everyone. Scope is decided at write time and enforced as a hard partition, because a wrong global memory affects every user.

open as a page

In agent memory, how do you decide which facts from a session are worth storing?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Store only facts that are durable, specific to this user or account, and likely to change a future decision. Skip general world knowledge the model already has, and skip anything provisional. A wrong stored fact costs far more than a missed one.

open as a page

How would you measure whether an agent's memory retrieval is actually helping?

level: principalimportance: should knowfreq 38%

basics

~20 s

Measure usefulness, not recall. Track how often a retrieved memory is actually used in the reply, whether turns that used memory produced better outcomes than the same turns with memory disabled, and how often a retrieved memory made the answer wrong.

open as a page

Why design deliberate forgetting into agent memory, and how would you choose TTLs?

level: principalimportance: should knowfreq 40%

basics

~20 s

Unbounded memory degrades: stale facts get recalled as current and crowd out live ones. Assign an expiry at write time from how volatile the fact class is — months for a situational claim, none for a structural one — and prefer archiving over erasing so decisions stay auditable.

open as a page