In a multi-user AutoGen service, how do you decide what to persist and what to discard per session?
answer
- three artefacts, three lifecycles
- configuration is never per user
- bound the source, not the snapshot
- version the blob, fail closed
- deletion must reach every copy
basics
~20 sSeparate three stores: configuration rebuilt from code or component definitions, resume state from save_state kept per session and bounded, and the transcript kept for audit under its own retention. Persist only what a resume genuinely needs; everything else is observability data.
solid answer
~50 sTreat the three artefacts a run produces as separate stores with separate lifecycles. **Configuration** — agents, tools, model clients, termination — is code, or a serialized component definition via `dump_component()` / `load_component()`; it is versioned with your deploy, never with a user. **Resume state** — what `save_state()` returns — is the minimum needed to continue a conversation; it grows with the transcript, so bound the agents' model contexts rather than hoping the blob stays small, and store one row per session with a schema version so a mismatched blob fails closed to a fresh session. **Transcript and telemetry** — the streamed messages, token usage and spans — answer *why* a run behaved as it did and belong in your observability store with its own, usually shorter, retention. The judgment call is which agents deserve resume state at all: often only one participant carries knowledge worth continuing, and letting reviewers start clean each session cuts storage and prompt cost together.
code
python · 24 linesimport json
from dataclasses import dataclass
STATE_SCHEMA = 3
@dataclass
class StoredSession:
session_id: str
schema: int
state_json: str
async def load_session(team, row: StoredSession | None) -> bool:
"""Return True if prior state was applied; fail closed on version skew."""
if row is None or row.schema != STATE_SCHEMA:
return False
await team.load_state(json.loads(row.state_json))
return True
async def store_session(team, session_id: str) -> StoredSession:
state = await team.save_state()
return StoredSession(session_id, STATE_SCHEMA, json.dumps(state))go deeper
Know that saved state and team configuration are different things, and that a resumed session needs the team rebuilt in code before any stored state is applied.
Be able to explain why state grows with the transcript, why participants are matched by name, and why the transcript you keep for debugging is separate data from the state you keep for resuming.
Demonstrate the operational build: bounded contexts, one versioned row per session, fail-closed on schema mismatch, and metrics that separate completed runs from limit-terminated ones.
Own the whole data model — which artefacts exist, where each lives, how long it is kept, how deletion reaches every copy, and the cost argument for persisting only the agents whose continuity a user would actually notice.
## Three artefacts, three lifecycles Every AutoGen run produces three things that people routinely conflate: | artefact | source | lifecycle | |---|---|---| | configuration | your code, or a serialized component definition | versioned with the deploy | | resume state | `save_state()` | per session, bounded, deletable | | transcript + telemetry | the run stream, usage fields, spans | retained for audit and debugging | The design failure this question probes is a single blob that mixes them: storing a per-user record that also encodes which model and which tools the agents had. That record becomes un-migratable the day you change a system message, and it quietly turns your user database into a place where model configuration drifts per user. AutoGen makes the split easy because `save_state` deliberately omits configuration. The declarative side has its own surface: components expose `dump_component()` and `load_component()` so a team's *shape* can be serialized and reloaded independently of any conversation — which is also how a visual builder hands a team definition to code. ## Bounding resume state Saved state is proportional to the agents' accumulated model contexts. In a team, that multiplies by participants: five agents that each retain a full transcript store roughly five copies of the conversation. Left alone this grows without bound in a long-lived session. The lever is not the snapshot, it is the source. Configure agents' contexts to keep a bounded window or a running summary, and the persisted blob shrinks with them. This is the same decision as prompt cost — a context you bound for storage reasons is also a context you stop re-sending on every call — so it usually pays twice. The second lever is *which* agents you persist. Ask per participant: if this agent started fresh next session, would the user notice? A research agent that accumulated findings: yes. A formatting critic: no. Persisting per agent rather than per team is supported and is often the cheaper design. ## Versioning and failing closed Stored state couples to two things: the library version and your team's topology, since participants are matched by name. A deploy that renames an agent orphans its history; a library upgrade may change the state shape. Store an explicit version tag alongside every blob and decide the mismatch policy up front. The safe default is to discard and start a fresh session, telling the user, rather than attempting a partial load that yields a team with half its memory. ## What the transcript is for Resume state answers "where are we". It does not answer "why did this cost two hundred thousand tokens". That answer lives in the streamed messages — per-message `models_usage`, the `stop_reason` on the result, the sequence of tool calls — and in spans. Keep those, keep them queryable by session id, and count runs terminated by a limit as failures rather than successes. A team you cannot explain after the fact is a team you can only rerun. ## Privacy is the constraint that decides ties Agent transcripts contain whatever users typed and whatever tools fetched. Three copies of it now exist: the resume state, the transcript store, and possibly the tracing backend. Deletion must reach all three, and only a deliberate design makes that possible. This is frequently what tips the decision toward persisting less: a session that resumes from a short summary rather than a full transcript is cheaper, smaller, and easier to erase on request. ## Prototyping versus production A visual builder is excellent for exploring team shapes and terrible as a production runtime: its storage model is its own, and its convenience is exactly the coupling you do not want in a service. The healthy pattern is to prototype interactively, export the team definition, and let your service own configuration, state and telemetry explicitly. ## What a strong answer sounds like Name the three artefacts, give each a store and a retention, bound the state at its source, version the blob, and justify per-agent persistence with a cost argument. Candidates who answer "just call save_state and put it in Postgres" have not yet operated one of these at scale.
- What role does dump_component / load_component play next to save_state?They serialize a team's *configuration* — which agents exist, their models and tools — as a declarative definition, while `save_state` captures only the running conversation. Keeping them separate lets configuration version with your deploy and state version with the user's session, so changing a system message does not invalidate every stored session.
- Where does AutoGen Studio fit in this picture?It is a prototyping surface: you assemble and try teams interactively, then export the team definition rather than running production traffic through it. Its storage and session model are its own. Treating it as the runtime couples your service to a tool designed for exploration, and gives you none of the state, retention or tracing control a service needs.
- A user asks you to delete their data. What has to be touched?Every copy: the resume state row, the stored transcript, and anything the tracing backend captured if spans or model-call events carried prompt text. If deletion was not designed for up front, the third one is the copy people forget. This is a strong argument for capturing less payload in telemetry and resuming from summaries rather than full transcripts.
saying these in an interview costs you the question
- Storing agent configuration inside the per-user session record
- Assuming saved state stays small in a long-lived session
- Persisting every participant's context when only one matters
- No version tag, so a deploy silently half-loads old state
- Treating a prototyping UI as the production runtime