skip to content

State, Streaming & Observability

You will learn how to persist a team's conversation, stream it to a console or UI, and trace what every agent did. Interviewers ask because a multi-agent transcript is the only way to explain why a team burned two hundred thousand tokens and produced nothing.

part ofAI agent & RAG frameworksoverview, primer and where to startread it →
on this pageshow

questions

6

In AutoGen, what does Console(team.run_stream(...)) give you over awaiting team.run()?

level: middleimportance: must knowfreq 68%

answer

  1. one is blocking, one is incremental
  2. generator yields messages, then the result
  3. a UI helper that prints and returns
  4. message-granular, not token-granular
  5. final yielded item is TaskResult

basics

~20 s

run_stream yields each agent message and event as it is produced; Console consumes that async generator and prints items as they arrive. run() blocks and returns only the final TaskResult. Console returns that same TaskResult at the end.

solid answer

~40 s

`await team.run(task=...)` is the blocking form: it executes the whole multi-agent conversation and hands back a single `TaskResult` at the end, so you see nothing while it works. `team.run_stream(task=...)` returns an async generator that yields each item as the team produces it — agent replies, tool-call request and execution events, handoffs — and yields the `TaskResult` last. `Console` from `autogen_agentchat.ui` is a helper that consumes such a stream, pretty-prints every item to stdout, and returns that final item, so `result = await Console(team.run_stream(task=...))` gives you live output *and* the result object. Passing `output_stats=True` makes Console print a usage and timing summary as well. For a web UI you skip Console and iterate the generator yourself, forwarding each message over your own transport.

code

python · 29 lines
python
import asyncio

from autogen_agentchat.agents import AssistantAgent
from autogen_agentchat.conditions import MaxMessageTermination
from autogen_agentchat.teams import RoundRobinGroupChat
from autogen_agentchat.ui import Console
from autogen_ext.models.openai import OpenAIChatCompletionClient


async def main() -> None:
    client = OpenAIChatCompletionClient(model="gpt-4o")
    writer = AssistantAgent("writer", model_client=client)
    critic = AssistantAgent("critic", model_client=client)
    team = RoundRobinGroupChat(
        [writer, critic],
        termination_condition=MaxMessageTermination(6),
    )

    # Console prints every item and returns the final TaskResult.
    result = await Console(
        team.run_stream(task="Write a haiku about rain."),
        output_stats=True,
    )
    print(result.stop_reason)

    await client.close()


asyncio.run(main())

go deeper

for a junior

Know the two calls: team.run() waits and returns a result, team.run_stream() yields items as they happen, and Console prints a stream in the terminal. Be able to write the one-liner that wraps a stream in Console.

for a middle

Explain that the stream yields the same message and event objects that end up in TaskResult.messages, that the last item is the TaskResult, and that streaming here is message-granular rather than token-granular.

for a senior

Show that you consume the generator yourself in a service, branch on item type to decide what reaches the user, and pass a CancellationToken so a looping team can be killed. Treat Console as a dev tool.

for a principal

Frame it as an operability decision: the stream is the only place a multi-agent run explains itself, so what your platform captures from it — per-message provenance, cost, latency — determines whether teams can be debugged or only rerun.

## Two entry points, one execution An AutoGen AgentChat team exposes two ways to run a task: - `await team.run(task="...")` returns a `TaskResult` once the conversation has terminated. - `team.run_stream(task="...")` returns an `AsyncGenerator` that yields items during the run and yields the `TaskResult` as its final item. These are not different execution engines. `run()` is the convenience wrapper; the only difference is whether intermediate items reach you. Individual agents have the analogous pair (`run` / `run_stream`), which matters when you drive one agent rather than a team. ## What the stream actually yields The items are the same objects that end up in `TaskResult.messages`. In `autogen-agentchat` 0.7.x you will typically see `TextMessage` for an agent's reply, `ToolCallRequestEvent` when a model asks for a tool, `ToolCallExecutionEvent` with the tool's return value, `ToolCallSummaryMessage`, `HandoffMessage` when control moves between agents, `MemoryQueryEvent` when a memory injected content, and `UserInputRequestedEvent` when a user proxy is waiting on a human. The distinction the type names encode: *messages* are conversation content that other agents see; *events* are internal happenings surfaced for observability. Everything is a Pydantic model, so it serializes cleanly for a UI. One thing the stream does **not** give you by default is token-by-token output. Streaming at this level is message-granular: an agent's reply appears once the model call completes. Chunk-level items only appear when the agent itself has been configured to stream from its model client, which is a property of the agent rather than of `run_stream`. ## What Console is and is not `Console` is deliberately small: an async function that takes a stream, prints each item with the source agent's name, and returns the last item it saw. Because it returns that item, the idiomatic line is: `result = await Console(team.run_stream(task="..."))` and `result` is a `TaskResult` with `.messages` and `.stop_reason`. If you call it on a single agent's stream, it returns that agent's `Response` instead — same contract, different last item. Useful arguments: `output_stats=True` prints a summary line per message and a total (token usage and elapsed time), which is the cheapest way to see what a run cost. `user_input_manager` wires Console into human-in-the-loop flows so a `UserInputRequestedEvent` produces a prompt on the terminal. Console is a **development tool**. It writes to stdout, it has no backpressure story, and it cannot be rendered anywhere but a terminal. In a service you consume the generator directly: ``` async for item in team.run_stream(task="..."): ... ``` (pseudocode; see the code example) and decide per item type what to forward to the browser, what to log, and what to drop. ## Why interviewers ask The honest reason is cost forensics. A multi-agent team can loop for dozens of turns; if you only ever call `run()`, a failed run gives you a final message and no explanation. The stream is where you see that the critic kept asking for revisions, or that a tool returned an error string that the model then ignored. Candidates who reach for `run()` in production and only add streaming after their first runaway bill usually say so; candidates who have operated these systems reach for `run_stream` from the start and treat `run()` as sugar for batch jobs. ## Cancellation and cleanup Both forms accept a `CancellationToken`, which is how you abort a runaway team from outside — you cannot interrupt a team by simply stopping iteration of the generator without leaving the underlying work running. Remember also that model clients hold connections; `await model_client.close()` at the end of a script is the usual cleanup step.

  • How do you tell the final TaskResult apart from the messages in the same stream?
    By type. Every intermediate item is a chat message or an agent event; the last item is a `TaskResult` from `autogen_agentchat.base`. In a consumer loop you branch on `isinstance(item, TaskResult)` rather than counting items, because the number of messages varies per run. `Console` does this for you and returns that object.
  • Console shows nothing for several seconds while an agent thinks. Is that a bug?
    No — streaming here is message-granular. An agent's `TextMessage` is emitted after its model call returns, so a slow model looks like silence. Token-level chunks require enabling streaming on the agent's side; that is an agent configuration concern, not something `run_stream` can add on its own.
  • How do you stop a team mid-stream from outside the loop?
    Pass a `CancellationToken` into `run_stream` and call `cancel()` on it from another task, typically on a timeout or a user action. Just breaking out of the `async for` loop stops your consumption but does not reliably tear down the in-flight agent work, so you leak model calls and cost.

saying these in an interview costs you the question

  • Thinking run_stream streams tokens rather than whole messages
  • Assuming Console returns None so discarding the TaskResult
  • Believing run() and run_stream() execute different logic
  • Using Console in a production web service to render output
  • Breaking out of the async for loop as a way to cancel a run

context

open as a page

What does AutoGen's TaskResult contain, and where do you read a run's token usage?

level: middleimportance: must knowfreq 62%

basics

~20 s

TaskResult holds messages, the full ordered list of the run's messages and events, and stop_reason, a string saying why the run ended. Token usage is not a top-level field: each message carries models_usage with prompt_tokens and completion_tokens, and you sum those.

open as a page

How do you persist and resume an AutoGen team across restarts with save_state and load_state?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Call await team.save_state() to get a JSON-serializable mapping of the conversation state, store it, then rebuild an identically configured team in the new process and call await team.load_state(state). Configuration — model clients, tools, system messages — is not in the state and must be reconstructed in code.

open as a page

How does ListMemory change what an AutoGen AssistantAgent sends to the model?

level: middleimportance: should knowfreq 45%

basics

~20 s

An AssistantAgent given memory=[ListMemory(...)] calls each memory's update_context before every model call, and ListMemory appends all of its stored MemoryContent items to that call's context. It ignores the query, so the injected block grows with everything you have added.

open as a page

How do you wire OpenTelemetry tracing into an AutoGen team, and what do the spans show?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Build an OpenTelemetry TracerProvider with an exporter, pass it as tracer_provider to a SingleThreadedAgentRuntime, and hand that runtime to the team constructor. The runtime then emits spans for message dispatch and agent processing, so one multi-agent run becomes a single nested trace.

open as a page

In a multi-user AutoGen service, how do you decide what to persist and what to discard per session?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Separate three stores: configuration rebuilt from code or component definitions, resume state from save_state kept per session and bounded, and the transcript kept for audit under its own retention. Persist only what a resume genuinely needs; everything else is observability data.

open as a page