skip to content

In AutoGen, how does AssistantAgent.on_messages() differ from run()?

level: middleimportance: should knowfreq 52%

answer

  1. two layers, same agent
  2. one returns a reply, one a transcript
  3. teams share the outer signature
  4. inner events versus conversation messages
  5. neither clears the agent's history

basics

~20 s

on_messages() is the low-level agent protocol: pass a sequence of messages plus a CancellationToken and get back a Response holding chat_message and inner_messages. run() is the task-runner wrapper: pass a task string and get a TaskResult with the full message list and a stop_reason.

solid answer

~40 s

Both are async and both advance the same agent, but they sit at different layers of `autogen-agentchat`. `on_messages(messages, cancellation_token)` is the agent contract — it takes a sequence of chat messages and returns a `Response` whose `chat_message` is the agent's reply and whose `inner_messages` carries the tool-call events produced along the way. That is the method a custom agent implements and the one the team runtime calls. `run(task=...)` comes from the task-runner interface: it wraps a plain string task into a message, drives the agent, and returns a `TaskResult` with `messages` and `stop_reason`. Crucially, teams expose the *same* `run()` signature, so `run()` is the surface that lets you swap a single agent for a whole team without changing the caller. Neither resets state — both append to the agent's model context.

code

python · 25 lines
python
import asyncio

from autogen_agentchat.agents import AssistantAgent
from autogen_agentchat.messages import TextMessage
from autogen_core import CancellationToken
from autogen_ext.models.openai import OpenAIChatCompletionClient


async def main() -> None:
    client = OpenAIChatCompletionClient(model="gpt-4o")
    agent = AssistantAgent(name="assistant", model_client=client)

    response = await agent.on_messages(
        [TextMessage(content="Say hi.", source="user")],
        CancellationToken(),
    )
    print(response.chat_message.to_text(), response.inner_messages)

    result = await agent.run(task="Now say goodbye.")
    print(len(result.messages), result.stop_reason)

    await client.close()


asyncio.run(main())

go deeper

for a junior

Recall that both are async and that run(task="...") is the easy entry point which returns a result object containing the messages. Do not claim either one resets the agent.

for a middle

Explain the layering: the message-handling contract returns one Response with chat_message and inner_messages, while the task-runner contract returns a TaskResult with messages and stop_reason.

for a senior

Talk about why the shared run() signature between an agent and a team matters for substitutability, and about wiring a per-request CancellationToken so an abandoned client stops burning model tokens.

for a principal

Own the interface argument: a narrow message-handling contract is what lets orchestration, termination and observability be implemented once, above every agent type, instead of being reinvented per agent.

## Two layers, one agent `autogen-agentchat` separates *being an agent* from *being something you hand a task to*. `AssistantAgent` implements both interfaces, which is why it has two entry points that look redundant until you see who calls each one. ## on_messages: the agent contract ``` async def on_messages(messages, cancellation_token) -> Response ``` It takes a sequence of chat messages (for example a `TextMessage(content=..., source="user")`) and a `CancellationToken`, and returns a single `Response`. The `Response` has two fields that matter: - `chat_message` — the one message the agent contributes to the conversation. For a text reply that is a `TextMessage`; when tools ran without reflection it is a `ToolCallSummaryMessage`. - `inner_messages` — the events produced on the way there, such as the tool-call request and the tool-call execution events. These are *not* part of the conversation other agents see; they are the audit trail of the turn. This is the method you override when you write your own agent type, and it is what a team's runtime invokes on each participant when it is that participant's turn. The `CancellationToken` is a required positional argument precisely because an agent turn can block on a model call, a tool, or human input, and the caller must be able to abort it. There is a streaming counterpart, `on_messages_stream`, which yields inner events as they happen and finishes by yielding the `Response`. ## run: the task-runner contract ``` await agent.run(task="...") -> TaskResult ``` `run()` is convenience plus polymorphism. It accepts a task as a plain string (or as a message / list of messages), wraps a string into a user message for you, invokes the agent, and returns a `TaskResult` containing `messages` — the whole sequence, including the task message you supplied — and `stop_reason`, a human-readable string explaining why it finished. For a single agent the stop reason is trivial; for a team it names the termination condition that fired. The reason to care is substitutability. A team object exposes the identical `run()` / `run_stream()` signature. Code written against `run()` can be handed one agent today and a group of agents tomorrow with no change at the call site. Code written against `on_messages()` is bound to the single-agent shape. ## What neither of them does **Neither clears state.** `AssistantAgent` keeps a model context across calls, so a second `run()` on the same instance sees the first exchange. If you want a clean slate you call `on_reset()` (or reset the team). People routinely mistake `run()` for a stateless request/response call and then wonder why costs and context climb across an interactive session. **Neither is a loop over turns for a single agent.** Handing a hard task to `run()` does not make the agent keep working until success; it performs its bounded turn and returns. ## Choosing between them - Application code driving a conversation, or code that might later be pointed at a team: use `run()` / `run_stream()`. - Writing a custom agent type, or building your own orchestration that decides message-by-message who speaks next: implement and call `on_messages()` / `on_messages_stream()`. - Needing the tool-call events for logging or a UI: they are `Response.inner_messages` from `on_messages()`, and they also appear inline in the `TaskResult.messages` sequence from `run()`. ## Cancellation in practice A `CancellationToken` passed into either call is what stops an in-flight turn. Cancelling it aborts the underlying model request rather than leaving a coroutine hanging until the HTTP timeout. In a service, create one token per request, hold it, and cancel it when the client disconnects — otherwise an abandoned request keeps burning tokens. ## The mental model to state in an interview `on_messages()` is "here are messages, produce your one reply". `run()` is "here is a task, work it and tell me the transcript and why you stopped". The first is the plumbing that teams are built on; the second is the uniform handle that makes an agent and a team interchangeable to a caller.

  • Which of the two do you implement when writing your own agent type?
    `on_messages()` (and optionally `on_messages_stream()`), plus `on_reset()` and the produced-message-types declaration. That is the agent contract the runtime calls. `run()` comes from the task-runner side and is provided for you — you do not reimplement it, and building a custom agent around it would leave the agent unusable inside a team, since teams drive participants through the message-handling contract.
  • What exactly lands in Response.inner_messages?
    The events generated while producing the turn — principally the tool-call request and tool-call execution events, and streaming chunk events when model_client_stream is on. They are the turn's audit trail rather than conversational content: other participants react to `chat_message`, while `inner_messages` is what you log or render in a debug panel to show why the agent answered as it did.
  • Does calling run() twice on the same agent start a fresh conversation?
    No. `AssistantAgent` retains its model context between calls, so the second `run()` sees the first exchange and re-sends it to the model. To start clean, call `on_reset()` on the agent (or reset the team) before the next task. Assuming statelessness is a common source of both surprising answers and quietly growing per-call cost in long-lived services.

saying these in an interview costs you the question

  • Thinks run() is stateless and starts a fresh conversation each call
  • Calls on_messages() a synchronous convenience wrapper around run()
  • Believes inner_messages are seen by the other agents in a team
  • Says run() keeps working until the task actually succeeds
  • Ignores the CancellationToken as optional boilerplate

context