Which backing stores fit working, episodic, semantic and procedural agent memory?
answer
- let the read pattern pick the store
- order matters for one tier only
- facts need identity, scope and timestamps
- skills belong in version control
- one vector store for everything is the trap
basics
~20 sAccess pattern picks the store. Working memory suits a fast session-keyed cache or a scratchpad file; episodic memory suits an append-only, time-ordered event log; semantic memory suits a row store with a vector index for fuzzy recall; procedural memory suits versioned skill files in source control.
solid answer
~60 sEach type is read differently, and the read pattern should drive the storage choice. **Working memory** is written and re-read many times within one session and never after it, so a fast key-value store keyed by session — or simply a scratchpad file next to the task — is the right shape. **Episodic memory** is a time-ordered sequence you replay or page through, so an append-only event table with timestamps fits; you rarely update an episode, you add to the log. **Semantic memory** is a set of facts you look up by similarity or by attribute, and that must be updated as facts change, so a row store carrying identity, scope and timestamps, paired with a vector index for fuzzy recall, is the standard pairing. **Procedural memory** is fetched by name and must be reviewable and testable, so a versioned directory of skill or script files in source control beats any embedding index. As of mid-2026 the consensus is hybrid and routed rather than one store for everything — and file-directory memory has proved surprisingly competitive with purpose-built vector systems on long-conversation benchmarks.
go deeper
Know that different memory types go in different places, and be able to say that a vector index is for fuzzy fact lookup rather than for everything an agent stores.
Derive each mapping from its read pattern — ordered replay, similarity plus filters, exact fetch by name — instead of reciting a table. Explain why facts need identity, scope and timestamps rather than just an embedding.
Talk about the operational consequences: unbounded episodic growth, per-tier retention and deletion for privacy, and how the split makes it possible to say which tier produced a bad answer during an incident.
Own the build-versus-simplify call. Argue when a single file-based substrate is sufficient and when a tier earns its own system, and be explicit about the migration and lock-in cost of committing to a specialised memory product.
## The rule: read pattern picks the store The mistake this question exists to catch is choosing one store — almost always a vector index — and pushing all four memory types into it. The types are read in genuinely different ways, and a store optimised for one read pattern serves the others badly. Ask of each type: how is it fetched, how often does it change, and does order matter? ## Working memory Read pattern: written and re-read repeatedly inside one session, never afterwards; latency matters; correctness is exact, not fuzzy. That points at a fast key-value store keyed by session id, with an expiry, or at a plain scratchpad file living alongside the task. A cache such as Redis is a natural fit for a service handling many concurrent sessions; a file is a natural fit for an agent working inside a sandbox or workspace, and has the pleasant property that the agent can read it with the same tools it uses for everything else. What you do not want is a similarity index — you know exactly which key you need, and approximate retrieval of your own in-progress notes is a strictly worse experience. ## Episodic memory Read pattern: ordered replay, or a filtered scan such as "the last five sessions for this user". Episodes are immutable once written and their sequence carries meaning. An append-only table or event log with a timestamp, a session id and an actor is the canonical shape. Ordering and cheap range scans are what you are buying. Two practical notes: raw episodes grow without bound, so most systems store a condensed per-session summary alongside the full log and read the summaries by default; and putting episodes in a vector index alone is a recognised anti-pattern, because similarity search returns a fragment from an arbitrary point in history with no sense of when it happened or what came next. ## Semantic memory Read pattern: look up facts relevant to the current situation, by meaning as well as by attribute, and update them when they change. This is the one tier where a vector index genuinely earns its place — but on its own it is not enough. Facts need identity so they can be updated rather than duplicated, they need scope so they reach the right user, and they need timestamps so staleness is visible. So the standard arrangement is a row store carrying the fact, its scope, its provenance and its validity, with a vector index over the fact text for fuzzy recall, and metadata filters applied alongside the similarity search. A knowledge-graph store is the alternative when relationships between entities and their changes over time are the point; graph-backed temporal memory leads the bi-temporal reasoning benchmarks as of mid-2026. ## Procedural memory Read pattern: fetch a named capability, in full, and execute or follow it. Nothing about that wants approximate retrieval. A versioned directory of skill files or scripts in source control gives you exactly the properties that matter: diffs, code review, tests, rollback, and a stable name to fetch by. Selection among many skills is usually done on short metadata — a name and a one-line description — with the full instructions and any bundled scripts loaded only when the skill is activated. ## What the split buys you Different retention and deletion policies per tier, which matters for privacy and for cost; different growth curves, so unbounded episodic logs do not degrade fact lookup; and different review requirements, since a bad skill file is a code-safety problem while a stale fact is a correctness problem. It also makes debugging tractable — when an agent behaves oddly you can ask which tier lied. ## The honest caveat This mapping is a default, not a law, and 2026 practice is less dogmatic than it looks. Filesystem-based memory, where a directory of markdown files carries all four types and the agent navigates it with ordinary file tools, performs competitively with purpose-built memory frameworks on long-conversation benchmarks, and providers now ship file-directory memory as a first-class agent capability. The reason it works is that the agent does the routing itself instead of relying on blind similarity search. The defensible position in an interview is: start with the simplest substrate that supports the access patterns you actually have, split a tier out when its read pattern or its retention policy diverges from the rest, and be able to say which pattern justified the split.
- What specifically breaks if you keep episodes only in a vector index?You lose ordering and recency semantics. Similarity search returns a fragment from an arbitrary point in history with no sense of what preceded or followed it, so the agent can act on a superseded state as if it were current. You also cannot cheaply answer "what happened in the last session", which is the most common episodic query there is.
- When is a knowledge graph a better fit than a row store plus vector index for facts?When the relationships between entities, and how they change over time, are what you actually query — who reports to whom, which service depends on which, what was true during a given window. Graph-backed temporal memory handles multi-hop and bi-temporal questions that flat fact rows answer poorly, at the cost of a heavier write path and more schema thinking.
- Would you ever deliberately put all four types in one store?Yes, and it is a reasonable starting point. A single directory of files, navigated by the agent with ordinary file tools, covers all four and performs well because the agent routes its own lookups instead of trusting blind similarity search. Split a tier out when its retention policy, growth rate or review requirement genuinely diverges from the rest.
saying these in an interview costs you the question
- Puts every memory type into one vector index by default
- Stores episodes without timestamps or ordering guarantees
- Keeps facts with no identity, so updates create duplicates instead
- Uses similarity search to fetch a skill that has an exact name
- Assumes a vector database is required before any memory can work