skip to content

In AutoGen, what does Console(team.run_stream(...)) give you over awaiting team.run()?

level: middleimportance: must knowfreq 68%

answer

  1. one is blocking, one is incremental
  2. generator yields messages, then the result
  3. a UI helper that prints and returns
  4. message-granular, not token-granular
  5. final yielded item is TaskResult

basics

~20 s

run_stream yields each agent message and event as it is produced; Console consumes that async generator and prints items as they arrive. run() blocks and returns only the final TaskResult. Console returns that same TaskResult at the end.

solid answer

~40 s

`await team.run(task=...)` is the blocking form: it executes the whole multi-agent conversation and hands back a single `TaskResult` at the end, so you see nothing while it works. `team.run_stream(task=...)` returns an async generator that yields each item as the team produces it — agent replies, tool-call request and execution events, handoffs — and yields the `TaskResult` last. `Console` from `autogen_agentchat.ui` is a helper that consumes such a stream, pretty-prints every item to stdout, and returns that final item, so `result = await Console(team.run_stream(task=...))` gives you live output *and* the result object. Passing `output_stats=True` makes Console print a usage and timing summary as well. For a web UI you skip Console and iterate the generator yourself, forwarding each message over your own transport.

code

python · 29 lines
python
import asyncio

from autogen_agentchat.agents import AssistantAgent
from autogen_agentchat.conditions import MaxMessageTermination
from autogen_agentchat.teams import RoundRobinGroupChat
from autogen_agentchat.ui import Console
from autogen_ext.models.openai import OpenAIChatCompletionClient


async def main() -> None:
    client = OpenAIChatCompletionClient(model="gpt-4o")
    writer = AssistantAgent("writer", model_client=client)
    critic = AssistantAgent("critic", model_client=client)
    team = RoundRobinGroupChat(
        [writer, critic],
        termination_condition=MaxMessageTermination(6),
    )

    # Console prints every item and returns the final TaskResult.
    result = await Console(
        team.run_stream(task="Write a haiku about rain."),
        output_stats=True,
    )
    print(result.stop_reason)

    await client.close()


asyncio.run(main())

go deeper

for a junior

Know the two calls: team.run() waits and returns a result, team.run_stream() yields items as they happen, and Console prints a stream in the terminal. Be able to write the one-liner that wraps a stream in Console.

for a middle

Explain that the stream yields the same message and event objects that end up in TaskResult.messages, that the last item is the TaskResult, and that streaming here is message-granular rather than token-granular.

for a senior

Show that you consume the generator yourself in a service, branch on item type to decide what reaches the user, and pass a CancellationToken so a looping team can be killed. Treat Console as a dev tool.

for a principal

Frame it as an operability decision: the stream is the only place a multi-agent run explains itself, so what your platform captures from it — per-message provenance, cost, latency — determines whether teams can be debugged or only rerun.

## Two entry points, one execution An AutoGen AgentChat team exposes two ways to run a task: - `await team.run(task="...")` returns a `TaskResult` once the conversation has terminated. - `team.run_stream(task="...")` returns an `AsyncGenerator` that yields items during the run and yields the `TaskResult` as its final item. These are not different execution engines. `run()` is the convenience wrapper; the only difference is whether intermediate items reach you. Individual agents have the analogous pair (`run` / `run_stream`), which matters when you drive one agent rather than a team. ## What the stream actually yields The items are the same objects that end up in `TaskResult.messages`. In `autogen-agentchat` 0.7.x you will typically see `TextMessage` for an agent's reply, `ToolCallRequestEvent` when a model asks for a tool, `ToolCallExecutionEvent` with the tool's return value, `ToolCallSummaryMessage`, `HandoffMessage` when control moves between agents, `MemoryQueryEvent` when a memory injected content, and `UserInputRequestedEvent` when a user proxy is waiting on a human. The distinction the type names encode: *messages* are conversation content that other agents see; *events* are internal happenings surfaced for observability. Everything is a Pydantic model, so it serializes cleanly for a UI. One thing the stream does **not** give you by default is token-by-token output. Streaming at this level is message-granular: an agent's reply appears once the model call completes. Chunk-level items only appear when the agent itself has been configured to stream from its model client, which is a property of the agent rather than of `run_stream`. ## What Console is and is not `Console` is deliberately small: an async function that takes a stream, prints each item with the source agent's name, and returns the last item it saw. Because it returns that item, the idiomatic line is: `result = await Console(team.run_stream(task="..."))` and `result` is a `TaskResult` with `.messages` and `.stop_reason`. If you call it on a single agent's stream, it returns that agent's `Response` instead — same contract, different last item. Useful arguments: `output_stats=True` prints a summary line per message and a total (token usage and elapsed time), which is the cheapest way to see what a run cost. `user_input_manager` wires Console into human-in-the-loop flows so a `UserInputRequestedEvent` produces a prompt on the terminal. Console is a **development tool**. It writes to stdout, it has no backpressure story, and it cannot be rendered anywhere but a terminal. In a service you consume the generator directly: ``` async for item in team.run_stream(task="..."): ... ``` (pseudocode; see the code example) and decide per item type what to forward to the browser, what to log, and what to drop. ## Why interviewers ask The honest reason is cost forensics. A multi-agent team can loop for dozens of turns; if you only ever call `run()`, a failed run gives you a final message and no explanation. The stream is where you see that the critic kept asking for revisions, or that a tool returned an error string that the model then ignored. Candidates who reach for `run()` in production and only add streaming after their first runaway bill usually say so; candidates who have operated these systems reach for `run_stream` from the start and treat `run()` as sugar for batch jobs. ## Cancellation and cleanup Both forms accept a `CancellationToken`, which is how you abort a runaway team from outside — you cannot interrupt a team by simply stopping iteration of the generator without leaving the underlying work running. Remember also that model clients hold connections; `await model_client.close()` at the end of a script is the usual cleanup step.

  • How do you tell the final TaskResult apart from the messages in the same stream?
    By type. Every intermediate item is a chat message or an agent event; the last item is a `TaskResult` from `autogen_agentchat.base`. In a consumer loop you branch on `isinstance(item, TaskResult)` rather than counting items, because the number of messages varies per run. `Console` does this for you and returns that object.
  • Console shows nothing for several seconds while an agent thinks. Is that a bug?
    No — streaming here is message-granular. An agent's `TextMessage` is emitted after its model call returns, so a slow model looks like silence. Token-level chunks require enabling streaming on the agent's side; that is an agent configuration concern, not something `run_stream` can add on its own.
  • How do you stop a team mid-stream from outside the loop?
    Pass a `CancellationToken` into `run_stream` and call `cancel()` on it from another task, typically on a timeout or a user action. Just breaking out of the `async for` loop stops your consumption but does not reliably tear down the in-flight agent work, so you leak model calls and cost.

saying these in an interview costs you the question

  • Thinking run_stream streams tokens rather than whole messages
  • Assuming Console returns None so discarding the TaskResult
  • Believing run() and run_stream() execute different logic
  • Using Console in a production web service to render output
  • Breaking out of the async for loop as a way to cancel a run

context