How do you persist and resume an AutoGen team across restarts with save_state and load_state?
answer
- both calls are coroutines
- a JSON-serializable mapping
- conversation state, not configuration
- participants are matched by name
- grows with the transcript
basics
~20 sCall await team.save_state() to get a JSON-serializable mapping of the conversation state, store it, then rebuild an identically configured team in the new process and call await team.load_state(state). Configuration — model clients, tools, system messages — is not in the state and must be reconstructed in code.
solid answer
~50 s`save_state()` and `load_state()` are async methods on both agents and teams. `await team.save_state()` returns a `Mapping[str, Any]` that is JSON-serializable, so you write it to a database row or a file keyed by your session id. It contains the runtime conversation state: each participant's model context (the message history that agent will send next time) and the group chat manager's internal bookkeeping, such as whose turn it is and the current message count. It does **not** contain the team's configuration — model clients, API keys, tools, system messages, termination conditions and participant wiring all live in your code. Resuming therefore means constructing the same team again with the same participant names, then `await team.load_state(state)`. Loading into a differently-shaped team fails or silently misassigns history, because participants are matched by name. Agent-level `save_state()` gives you the same facility at single-agent granularity.
code
python · 32 linesimport json
from pathlib import Path
from autogen_agentchat.agents import AssistantAgent
from autogen_agentchat.conditions import MaxMessageTermination
from autogen_agentchat.teams import RoundRobinGroupChat
from autogen_ext.models.openai import OpenAIChatCompletionClient
STATE_FILE = Path("session-42.json")
def build_team() -> RoundRobinGroupChat:
"""Configuration lives here and must be identical on resume."""
client = OpenAIChatCompletionClient(model="gpt-4o")
return RoundRobinGroupChat(
[
AssistantAgent("researcher", model_client=client),
AssistantAgent("editor", model_client=client),
],
termination_condition=MaxMessageTermination(8),
)
async def run_turn(task: str) -> None:
team = build_team()
if STATE_FILE.exists():
await team.load_state(json.loads(STATE_FILE.read_text()))
await team.run(task=task)
state = await team.save_state()
STATE_FILE.write_text(json.dumps(state))go deeper
Know that both save_state and load_state are awaited, that save_state returns a plain JSON-serializable mapping you can store, and that you must build the team again before loading into it.
Explain the split: state is conversation history plus the manager's bookkeeping, configuration is agents, tools, clients and termination. Say why participants are matched by name and what breaks when they are renamed.
Demonstrate the operational picture — snapshot at run boundaries, version the stored blob against your code, bound model contexts so state does not grow forever, and fail closed to a fresh session on schema mismatch.
Own the persistence contract across the fleet: what a session is, retention and deletion of transcripts that contain user data, how state versioning survives deploys, and which agents are worth persisting at all given the storage cost.
## The problem it solves An AgentChat team is an in-memory object. Its participants accumulate a model context as the conversation runs, and the group chat manager tracks whose turn it is. Kill the process and all of that is gone. `save_state` / `load_state` is the mechanism for taking a snapshot of that runtime state and rehydrating it later — a different process, a different pod, tomorrow. ``` state = await team.save_state() # Mapping[str, Any] ... await team.load_state(state) ``` Both are coroutines; forgetting the `await` gives you a coroutine object that serializes to nonsense, which is a classic first-attempt bug. ## State versus configuration — the key distinction This is the point interviewers are probing. AutoGen deliberately splits two things: - **State** — what happened: message histories, the manager's turn pointer, per-agent bookkeeping. Captured by `save_state`. - **Configuration** — what the system *is*: which agents exist, their names and system messages, which model client each uses, which tools are registered, the termination condition. Captured by your code, or declaratively by the component-config surface (`dump_component()` / `load_component()`). Secrets are a consequence of that split: because model clients are configuration, API keys are not in the saved state. That is a feature — you can store state in a normal application database without it becoming a credential store. The practical rule: **rebuild, then load**. Construct the team exactly as you did the first time, with the same participant names, then apply the state. Participants are keyed by name, so renaming an agent between save and load orphans its history. ## What is inside For a team you get a nested structure: the team's own state plus a map of participant name to that agent's state. An `AssistantAgent`'s state is essentially its `llm_context` — the serialized messages it will replay on its next model call. `autogen_agentchat.state` defines the typed models behind this, but you should treat the mapping as opaque and version-sensitive rather than reaching into it. Do not hand-edit it to "fix" a conversation; construct the team differently instead. ## The failure modes that matter in production **Unbounded growth.** State size is proportional to the transcript. A long-running session accumulates every message the agents exchanged, and a team is chattier than a single assistant by construction. Persisting raw state per user per session gets expensive quickly; capping the model context so agents forget old turns is the mechanism that bounds it, and that cap must be part of the agent configuration you rebuild. **Version skew.** The state shape follows the library version and your team's topology. A stored blob from last month's deployment may not load into today's code. Store a schema/app version alongside the state and decide explicitly what happens on mismatch — usually "start a fresh session", which is far safer than a half-loaded team. **Mid-run snapshots.** Save between runs, at a natural boundary, not while a `run_stream` is in flight. Snapshotting concurrently with execution gives you state whose consistency you cannot reason about. **Confusing reset with resume.** Clearing a team's state and loading a saved state are opposite operations; a resume path that clears first and then loads nothing leaves you with a team that has silently forgotten the user. ## Granularity choices You can persist per agent instead of per team when only one participant carries knowledge worth keeping — for example a research agent whose gathered context you want across sessions while the reviewers start clean each time. That is a deliberate design decision, and it is usually cheaper than storing whole-team transcripts. Finally, note that state persistence is not observability. The saved state tells you where a conversation *is*; it does not tell you why it got there, what each model call cost, or how long it took. Those come from the message stream and from tracing.
- What happens if you load a saved state into a team whose participants were renamed?The state is keyed by participant name, so a renamed agent finds no matching entry and starts empty while the old entry is orphaned. You get a team that looks resumed but has silently lost one agent's history. Treat participant names as part of your persisted contract and version them alongside the state.
- Are API keys or tool definitions included in the saved state?No. `save_state` captures runtime conversation state only; model clients, tools, system messages and termination conditions are configuration you reconstruct in code, or serialize separately through the component-config surface. That is why saved state can live in an ordinary application database without becoming a secrets store.
- How do you stop persisted state from growing without bound in a long-lived session?Bound the source, not the snapshot. State size tracks the agents' model contexts, so configure those contexts to keep only a recent window or a summary, and the persisted blob shrinks with them. Alternatively persist only the agents whose history matters and let the rest start fresh each session.
- Where do you snapshot state in a service that streams runs?At a run boundary — after the stream has yielded its final result, before you accept the next task. Snapshotting mid-run captures a team halfway through an orchestration step, and you cannot reason about whether the manager's turn pointer and the participants' contexts agree.
saying these in an interview costs you the question
- Calling save_state without await and storing a coroutine
- Expecting load_state to recreate agents, tools or model clients
- Assuming saved state includes API keys or system messages
- Loading state into a team with different participant names
- Treating state size as constant regardless of transcript length