How does ListMemory change what an AutoGen AssistantAgent sends to the model?
answer
- a protocol with several implementations
- the hook runs before each model call
- it does not actually search
- injected block, not replaced instructions
- an event tells you what was inserted
basics
~20 sAn AssistantAgent given memory=[ListMemory(...)] calls each memory's update_context before every model call, and ListMemory appends all of its stored MemoryContent items to that call's context. It ignores the query, so the injected block grows with everything you have added.
solid answer
~40 s`ListMemory` from `autogen_core.memory` is the simplest implementation of AutoGen's `Memory` protocol: `add`, `query`, `update_context`, `clear`, `close`. You pass instances via `AssistantAgent(..., memory=[my_memory])`. Before each model call the agent invokes `update_context` on each attached memory, and ListMemory appends everything it holds — in insertion order — into that call's model context as an extra system-level block. It is a list, not a retriever: `query` returns the stored items regardless of what you asked, so the injected text grows linearly with `add` calls and silently inflates every prompt. You add content as `MemoryContent(content=..., mime_type=MemoryMimeType.TEXT)`. The injection is observable — a `MemoryQueryEvent` appears in the run stream showing exactly what was inserted. When the store outgrows a prompt, swap in a vector-backed `Memory` implementation; the agent code does not change because the protocol is the same.
code
python · 27 linesimport asyncio
from autogen_agentchat.agents import AssistantAgent
from autogen_agentchat.ui import Console
from autogen_core.memory import ListMemory, MemoryContent, MemoryMimeType
from autogen_ext.models.openai import OpenAIChatCompletionClient
async def main() -> None:
prefs = ListMemory(name="user_prefs")
await prefs.add(
MemoryContent(content="The user prefers metric units.", mime_type=MemoryMimeType.TEXT)
)
await prefs.add(
MemoryContent(content="Answer in British English.", mime_type=MemoryMimeType.TEXT)
)
client = OpenAIChatCompletionClient(model="gpt-4o")
agent = AssistantAgent("helper", model_client=client, memory=[prefs])
# A MemoryQueryEvent in the stream shows exactly what was injected.
await Console(agent.run_stream(task="How far is a marathon?"))
await client.close()
asyncio.run(main())go deeper
Know that memory is passed as a list to AssistantAgent, that you add items as MemoryContent, and that ListMemory simply keeps everything you put in it.
Explain the update_context hook: it runs before each model call and ListMemory appends all stored content, ignoring the query, so prompts grow linearly and every extra turn re-pays for the block.
Show the cost and debugging angle — count injections per run rather than per session, read MemoryQueryEvent to see what was inserted, and move to a vector-backed Memory implementation behind the same protocol when the set outgrows the prompt.
Decide what belongs in prompt-injected memory at all versus in retrieval or in the system message, and set the boundary so per-call token overhead stays predictable as the number of agents and turns scales.
## The protocol, then the implementation AutoGen separates the *contract* for memory from any particular store. `autogen_core.memory.Memory` defines a handful of async methods: - `add(content)` — put a `MemoryContent` in. - `query(query, ...)` — retrieve relevant content, returning a `MemoryQueryResult`. - `update_context(model_context)` — the hook the agent calls; the memory mutates the context that is about to be sent. - `clear()` / `close()` — empty it, release resources. `ListMemory` is the reference implementation and the one most examples use. It stores `MemoryContent` objects in insertion order and, crucially, **its `query` ignores the query**: everything it holds is considered relevant. That makes it perfect for a handful of durable facts ("the user prefers metric units", "always answer in Spanish") and wrong for a knowledge base. ## Where injection happens The interesting mechanic is `update_context`, not `query`. An `AssistantAgent` constructed with `memory=[m1, m2]` runs each memory's `update_context` against the model context *immediately before* each model call, in list order. ListMemory appends its contents as an additional system-level message ahead of the conversation. Three consequences fall out of this: 1. **Every model call pays.** Injection is per call, not per session. In a multi-agent team where an agent may be invoked many times, and inside a tool-calling loop where each iteration is another model call, the same memory block is re-sent each time. Tokens multiply by turns. 2. **It happens after your system message.** Memory content does not replace the agent's instructions; it stacks with them. Contradictions between a system message and injected memory are resolved by the model, unpredictably. 3. **It is per agent.** Memory attaches to agents, not teams. Two participants that both need a fact either share the same memory instance or each carry their own copy — sharing an instance is fine and is the usual answer. ## Observability Because injection is invisible in the final answer, AgentChat emits a `MemoryQueryEvent` into the message stream when a memory updates the context. That event names the memory and carries the content it inserted. This is the debugging tool for "why did the agent suddenly claim the user lives in Berlin" — you can see the exact block that was prepended. In a UI you usually hide these events from end users and keep them in your trace. ## Adding content `MemoryContent` carries `content`, a `mime_type` (`MemoryMimeType.TEXT`, `MARKDOWN`, `JSON`, and a binary option), and optional `metadata`. Text is by far the common case. You add outside the agent loop — from your application when a user states a preference, or from a callback after a run — because `ListMemory` has no self-updating behaviour; nothing writes to it unless you do. ## When it stops fitting ListMemory has no eviction, no scoring, no relevance. The failure curve is predictable: prompts grow, cost grows linearly with the number of stored facts, and eventually you exceed the context window or the model starts ignoring the middle of an over-long block. The fix is to move to a `Memory` implementation backed by a vector store, which performs an actual similarity query in `query` and injects only the top matches. Because the protocol is identical, the agent construction line is the only thing that changes. The boundary worth stating in an interview: AutoGen's memory surface is *where and how* content is injected into a model call. Which chunks are worth retrieving, how they are embedded, and how retrieval quality is measured are retrieval-design questions that live outside the framework. ## Memory versus state Do not conflate memory with persistence. Memory shapes the prompt; saved state records the conversation. A `ListMemory` populated at startup from your database is a normal pattern precisely because the memory object itself is not what gets snapshotted for you.
- Two agents in the same team need the same user preferences. How do you wire that?Construct one memory instance and pass it in the `memory` list of both agents. Memory attaches per agent, so the alternative — a separate instance each — duplicates content and doubles the maintenance. Sharing one instance also means an addition made during a run is visible to both agents on their next model call.
- When would you stop using ListMemory?As soon as the stored set stops fitting comfortably in every prompt. ListMemory injects all of its content on every model call, so cost grows linearly with the number of facts and eventually the block crowds out the conversation. At that point you move to a vector-backed implementation of the same `Memory` protocol so only relevant items are injected.
- How do you verify what memory actually put into a prompt?Watch the run stream for `MemoryQueryEvent` items. AgentChat emits one when a memory updates the model context, naming the memory and carrying the injected content, so you can see the exact text rather than inferring it from the answer. This is the first thing to check when an agent asserts a fact nobody typed.
saying these in an interview costs you the question
- Thinking ListMemory searches or ranks by relevance
- Assuming memory is injected once per session, not per model call
- Attaching memory to a team instead of to agents
- Expecting memory content to override the system message
- Believing memory writes itself from the conversation