skip to content

Working, Episodic and Semantic Memory

The borrowed-from-cognitive-science split that agent frameworks actually implement: working memory (this turn's context), episodic memory (what happened in past sessions), semantic memory (durable facts) and procedural memory (learned how-to). Naming them correctly is a common warm-up.

part ofAI agentsoverview, primer and where to startread it →
on this pageshow

questions

5

In an AI agent, what are working, episodic, semantic and procedural memory?

level: juniorimportance: must knowfreq 70%

answer

  1. four labels borrowed from cognitive science
  2. now, what happened, what is true, how to
  3. events versus facts is the key split
  4. procedural usually means code, not prose
  5. working memory dies with the session

basics

~20 s

Working memory holds what the agent is using right now for the current task. Episodic memory records what happened in past sessions. Semantic memory stores durable facts. Procedural memory holds learned how-to — reusable skills the agent can apply again.

solid answer

~50 s

The four labels are borrowed from cognitive science and each maps to a different engineering problem. **Working memory** is the current task's state — the latest message, the plan so far, tool results, scratch notes — scoped to this session and thrown away when it ends. **Episodic memory** is an ordered, timestamped record of what happened: past sessions, actions taken, outcomes. **Semantic memory** is the distilled durable facts, statements that are true independent of when they were learned. **Procedural memory** is reusable know-how, usually stored as code or skill files rather than prose. A tutoring agent shows all four: this session's turns are working memory, the log of the previous fourteen sessions is episodic, "struggles with fraction division" is semantic, and "how to run a Socratic hint sequence" is procedural. They differ in lifetime, access pattern and failure mode, which is why real systems rarely put them all in one undifferentiated store.

go deeper

for a junior

Be able to name all four types and give one concrete example of each without hesitating. Say plainly that working memory is this task, episodic is what happened, semantic is what is true, procedural is how to do it.

for a middle

Explain what makes the types different beyond the labels: lifetime, access pattern, and whether the data is a sequence or a set. Be ready to classify an ambiguous item, such as a session summary, and justify the call.

for a senior

Show you would not build all four unless the product needs them, and describe the failure each type introduces in production — overflow, unbounded logs, stale facts, unsafe skills. Interviewers listen for whether you have actually had to debug a wrong memory.

for a principal

Own the position that the taxonomy is vocabulary, not architecture. Be ready to argue when one substrate serving all four is the right call versus four purpose-built subsystems, and what that choice costs in review burden, migration risk and vendor lock-in.

## Where the vocabulary comes from Agent memory terminology is borrowed from cognitive psychology, which separates the small amount of information a person is holding in mind right now from the much larger store they can recall later, and then splits that long-term store into memory for events, memory for facts, and memory for skills. Agent frameworks adopted the same four labels because they line up with four distinct engineering problems: what to put into the next model call, what history to keep, what durable knowledge to accumulate, and what capability to reuse. The biological analogy is loose — nothing in a language model works like a hippocampus — but the vocabulary is now standard, and an interviewer expects you to use it precisely. ## Working memory Working memory is the state of the task in progress: the user's latest message, the plan, results returned by tools, intermediate notes. Most of it sits in the context window assembled for the current model call, but not all of it needs to. An agent that writes a scratchpad file partway through a long job and re-reads it after each tool result is also using working memory — the file is simply working memory that survives a context rebuild. The defining property is scope, not location: working memory belongs to one task or session and is discarded when that ends. Nothing in it persists unless something deliberately writes it somewhere durable. ## Episodic memory Episodic memory is the record of what happened, in order, with timestamps: on this date the user asked for X, the agent called tool Y, the call failed, the agent retried with different arguments and succeeded. It answers "what did we do last time?" Its natural shape is append-only and time-ordered — you replay or summarize episodes rather than edit them. Raw episodes are verbose, so production systems usually keep a condensed summary per session alongside, or instead of, the full transcript. ## Semantic memory Semantic memory holds durable facts, either distilled from episodes or supplied directly: this customer is on the enterprise plan, the deploy script lives at a particular path, this student struggles with fraction division. The shape is a set of statements rather than a sequence of events. Facts can still expire or be superseded, but they are not inherently tied to the moment they were observed. This is the tier that gets updated and deduplicated, because two different episodes can yield the same fact — or contradictory ones. ## Procedural memory Procedural memory is know-how: a routine the agent has learned or been given and can invoke again. In practice it is stored as executable code, scripts, or structured skill documents rather than as free prose, because a procedure can be tested and a paragraph cannot. The Voyager Minecraft agent is the canonical illustration — it wrote programs, verified them against the environment, kept the ones that worked in a growing skill library, and composed them into harder tasks later. ## Why the split matters The four types have different lifetimes (one task, one session, indefinite, versioned), different access patterns (rebuild every call, replay in order, look up by similarity or key, fetch by name), and different failure modes (overflow, unbounded growth, staleness and contradiction, broken or unsafe skills). Collapsing them into a single vector store of text chunks is the most common design mistake: you end up running similarity search over an event log and retrieving a fragment of an old transcript when what the agent needed was one current fact. ## What the taxonomy does not give you It is a naming scheme, not an architecture. It says nothing about when to write, what to retrieve, or how to decide something has gone stale — those are separate policy questions with their own tradeoffs. Some systems deliberately implement all four in one substrate, such as a directory of files, and let the folder layout carry the distinction. Naming the four types correctly is a warm-up; the interview usually moves quickly to which types your system actually needs and what happens when one of them is wrong. ## Common confusions worth pre-empting Retrieval over a static document corpus is not agent memory in this sense — it is external knowledge the agent reads, not experience it formed and can update. A rolling conversation summary is compressed episodic memory, not semantic memory, until the facts inside it are extracted as standalone statements. And working memory is not "the small memory": it is often the largest thing in play by token count, just the shortest-lived.

  • Is retrieval over a company document corpus a form of semantic memory?
    Usually no. Retrieval-augmented generation over a static corpus is external knowledge the agent reads; agent memory is state the agent formed from its own experience and can update. The line blurs when the agent writes back into that corpus, but the useful distinction is authorship and mutability: memory is agent-written and revisable, a knowledge base is curated for it.
  • Where does a rolling conversation summary sit in this taxonomy?
    It is compressed episodic memory — a shorter account of what happened, still tied to a sequence of turns. It becomes semantic memory only when standalone facts are lifted out of it and stored as statements in their own right. Treating a summary as if it were a fact store is a common source of stale, un-updatable claims.
  • Which of the four can a simple single-session assistant skip entirely?
    All but working memory. If nothing needs to survive the session, the conversation itself is the memory and adding episodic, semantic or procedural stores buys complexity with no payoff. The types earn their keep the moment a user returns and expects continuity, or the agent repeats a procedure often enough that re-deriving it is wasteful.

Working memory is what is spread across your desk for today's job; the other three are the filing cabinet behind you — the diary of past jobs, the folder of facts, and the binder of procedures.

saying these in an interview costs you the question

  • Calls every long-term store "the vector database" regardless of memory type
  • Treats a conversation summary as semantic memory rather than compressed episodes
  • Believes working memory persists between sessions on its own
  • Says procedural memory just means storing prompt templates
  • Describes the taxonomy as an architecture rather than a naming scheme

context

open as a page

Why is an AI agent's context window not the same thing as its memory?

level: middleimportance: must knowfreq 58%

basics

~20 s

The context window is the input assembled for one model call — bounded, rebuilt by your code every time, and gone afterwards. Memory is durable state stored outside the model that the agent deliberately writes and reads back. Continuity comes from the store, not the window.

open as a page

Which backing stores fit working, episodic, semantic and procedural agent memory?

level: middleimportance: should knowfreq 44%

basics

~20 s

Access pattern picks the store. Working memory suits a fast session-keyed cache or a scratchpad file; episodic memory suits an append-only, time-ordered event log; semantic memory suits a row store with a vector index for fuzzy recall; procedural memory suits versioned skill files in source control.

open as a page

Why is agent procedural memory usually stored as code or skill files, not prose?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Procedures need to be verified, repeated exactly, reviewed and versioned. Code and structured skill files give all four — you can test them, diff them and roll them back. Prose "lessons learned" cannot be tested, accumulate contradictions, and are recalled fuzzily rather than executed.

open as a page

When should an agent memory be scoped per-session, per-user, or globally?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Default to the narrowest scope that makes the memory useful: session for task state, user for anything personal or account-specific, global only for knowledge that is true for everyone. Scope is decided at write time and enforced as a hard partition, because a wrong global memory affects every user.

open as a page