How do you stop an AutoGen AssistantAgent's model context from growing unbounded?
answer
- the agent is stateful between calls
- the default keeps everything
- a window is a constructor argument
- instructions survive truncation
- truncation is silent forgetting
basics
~20 sPass a bounded model_context when constructing the agent — BufferedChatCompletionContext keeps only the last N messages, HeadAndTailChatCompletionContext keeps the first and last slices. The default is UnboundedChatCompletionContext, which re-sends the entire history on every model call.
solid answer
~40 sAn `AssistantAgent` is stateful: every incoming message, its own replies and its tool events accumulate in its model context, and each turn re-sends that whole context to the model. The default is `UnboundedChatCompletionContext`, so a long-lived agent gets steadily slower and more expensive and eventually fails on the model's context limit. The fix is to inject a bounded context at construction — `model_context=BufferedChatCompletionContext(buffer_size=n)` for a sliding window of the last n messages, or `HeadAndTailChatCompletionContext` when the opening of the conversation must survive truncation. The `system_message` is not part of that window: it is prepended at call time, so truncation never drops the agent's instructions. Complementary levers are calling `on_reset()` between unrelated tasks and creating a fresh agent per session rather than reusing one process-wide instance.
code
python · 10 linesfrom autogen_agentchat.agents import AssistantAgent
from autogen_core.model_context import BufferedChatCompletionContext
from autogen_ext.models.openai import OpenAIChatCompletionClient
agent = AssistantAgent(
name="assistant",
model_client=OpenAIChatCompletionClient(model="gpt-4o"),
model_context=BufferedChatCompletionContext(buffer_size=10),
system_message="You are a helpful assistant.",
)go deeper
Know that the agent remembers previous turns and that the history is sent to the model again each time, so conversations get more expensive as they get longer.
Name the default unbounded context and the bounded alternatives, and explain that the context is a constructor parameter rather than something you trim after the fact.
Show the diagnosis: rising per-turn latency and cost, then a hard context-limit failure deep into a session. Discuss window sizing against tool payloads and the tool-call pairing hazard.
Frame it as a memory-architecture decision: what must never be evicted, whether that belongs in the conversation at all, and how agent lifetime per session bounds both cost and cross-user exposure.
## The failure mode People treat `agent.run(task=...)` as a request/response call. It is not. `AssistantAgent` in `autogen-agentchat` 0.7.x holds a model context object, and every turn appends to it: the user message, the assistant reply, tool-call requests and tool results. On the next turn the agent sends the system message plus *everything in that context* to the model. Under the default `UnboundedChatCompletionContext` nothing is ever dropped. The symptom curve is characteristic. Latency creeps up turn by turn. Cost per turn climbs roughly linearly with conversation length, which means total cost for a session climbs quadratically. Then, at some point, a request fails outright because the assembled prompt exceeds the model's context window — and it fails on turn 40 of a session that worked fine in every test, which is why this usually reaches production. ## The lever: model_context The context is a constructor parameter, so bounding it is a wiring decision made once: - **`UnboundedChatCompletionContext`** — the default. Everything, forever. Fine for short, single-shot tasks and for evaluation runs where you want the full transcript. - **`BufferedChatCompletionContext(buffer_size=n)`** — a sliding window of the most recent n messages. Simple, predictable, and the right default for a chat-shaped agent. The cost is amnesia: facts established early silently fall out of the window, and the agent will confidently contradict what it agreed to twenty messages ago. - **`HeadAndTailChatCompletionContext`** — keeps a slice from the start and a slice from the end. This is the pragmatic choice when the opening of a conversation carries durable framing (the user's goal, constraints, an identifier) that must not be evicted while recent turns stay in view. One detail that reassures people: the `system_message` is **not** stored in the window. The agent assembles each request as its system messages followed by the context's messages, so truncation can never delete the agent's instructions. Only conversational history is at risk. ## Truncation is not free A sliding window is a lossy compression with no notification. Two consequences deserve airtime in an interview: 1. **Tool-call pairing.** A tool-call request and its result are separate entries in the history. A naive window boundary can leave a dangling half of that pair, and some providers reject a request whose tool-call and tool-result messages do not line up. Choose a window large enough to comfortably contain a full tool round-trip, and test with a tool-heavy conversation rather than a chatty one. 2. **Behavioural cliffs.** Quality does not decay smoothly; it falls off when the fact the agent needed happened to be the message that just aged out. If early context is load-bearing, a head-and-tail context or an explicit summary you re-inject as part of the task is more honest than a bigger window. ## Other levers around the same problem - **`on_reset()`** clears the agent's history. Call it between unrelated tasks on a reused instance. It is the difference between "this agent is a session" and "this agent is a service handling many sessions". - **Agent lifetime.** The cheapest fix is often architectural: build the agent per session or per request instead of holding one module-level instance shared by every user. A shared instance is not just a cost problem, it is a data-leak problem — one user's conversation becomes another user's context. - **Tool output size.** Tool results go into the context verbatim. A tool that returns a 50 KB document blows the window in one call. Truncate or summarize inside the tool; that is the cheapest token you will ever save. ## How to size the window Size it from the workload, not from a round number. Estimate the tokens per turn (user message + reply + typical tool payloads), decide how many turns of genuine memory the task needs, add headroom for the largest tool result you allow, and check the total against the model's context limit with the system prompt and tool schemas included — tool schemas are sent on every request and are not small when an agent has a dozen tools. ## What to say in an interview Name the default as unbounded, name the bounded alternatives, and then move immediately to the consequence: bounding the context trades correctness for predictability, so the real question is which facts must never be evicted and how you keep those alive — a summary re-injected into the task, a head-and-tail split, or state carried outside the conversation entirely.
- If you use a sliding window, does the agent lose its system_message once the window fills?No. The system message is not stored in the model context — the agent prepends its system messages to whatever the context returns when it builds each request. Truncation therefore only ever drops conversational history: user turns, assistant replies and tool events. This is why a windowed agent still behaves in character while forgetting facts, which can be more confusing to diagnose than an agent that visibly falls apart.
- What breaks if the window boundary cuts between a tool call and its result?You can end up sending a tool-result entry whose originating tool-call request has been evicted, or vice versa. Some providers reject such a request outright rather than tolerating the dangling half, so the failure looks like an intermittent API error rather than a memory problem. Size the window to comfortably hold a full tool round-trip, and exercise it with a tool-heavy conversation during testing.
- Is reusing one AssistantAgent instance across users ever acceptable?Only if you reset it between sessions and accept the serialization that implies. The context is per-instance, so a shared agent means one user's messages become part of another user's prompt — a correctness and privacy defect, not just a cost one. The usual design is an agent (and its context) constructed per session, with the model client shared, since the client is the expensive, stateless part.
saying these in an interview costs you the question
- Assumes each run() call sends only the latest message to the model
- Thinks the system_message is dropped when the window truncates
- Claims the framework automatically summarizes old messages
- Reuses one process-wide agent instance for every user session
- Ignores large tool outputs as a context-growth source