In AutoGen, what happens in one AssistantAgent turn when the model calls tools?
answer
- one model call, not a loop
- tool results may skip the model
- a flag decides whether it reflects
- default summary is the raw result
- iterations are capped at one
basics
~20 sAssistantAgent makes one model call with its system message plus its context, executes every requested tool, then by default returns a ToolCallSummaryMessage holding the raw tool results. It only sends those results back to the model when reflect_on_tool_use is True.
solid answer
~40 sA turn in `autogen-agentchat` 0.7.x is deliberately short. The agent prepends its `system_message` to the messages in its model context, calls the `model_client` once, and emits a `ToolCallRequestEvent` if the model asked for tools. It then runs those tool functions and emits a `ToolCallExecutionEvent`. What it returns next depends on configuration: by default (`reflect_on_tool_use=False`) it returns a `ToolCallSummaryMessage` whose content is the tool results rendered with `tool_call_summary_format`, which defaults to `"{result}"` — the model never sees the tool output. With `reflect_on_tool_use=True` it makes a second model call including the results and returns a natural-language `TextMessage`. It also does not loop by default: `max_tool_iterations` is 1, so one round of tool calls is all you get unless you raise it or put the agent in a team that keeps handing it turns.
code
python · 26 linesimport asyncio
from autogen_agentchat.agents import AssistantAgent
from autogen_ext.models.openai import OpenAIChatCompletionClient
def get_weather(city: str) -> str:
"""Return the current weather for a city."""
return f"{city}: 23C, sunny"
async def main() -> None:
client = OpenAIChatCompletionClient(model="gpt-4o")
agent = AssistantAgent(
name="weather_agent",
model_client=client,
tools=[get_weather],
system_message="You are a concise weather assistant.",
reflect_on_tool_use=True,
)
result = await agent.run(task="What is the weather in Paris?")
print(result.messages[-1].to_text())
await client.close()
asyncio.run(main())go deeper
Know that you pass plain Python functions as tools and that the agent runs them for you. Be able to say that the reply you get back may be the tool's raw output rather than a sentence written by the model.
Explain the turn mechanically: system message plus context, one model call, tool execution, then either a tool-call summary or a second reflective model call. Name reflect_on_tool_use and say what each setting costs.
Show you debug from the event stream: point at the tool-call request and execution events to prove where a turn ended, and connect a doubled bill to reflection adding a model call per turn on a context that also grew.
Own the argument that the loop belongs to the orchestrator rather than inside the agent, and that bounded turns are what make termination, budget caps and agent substitution enforceable at all in a multi-agent system.
## What an agent turn is In AutoGen's AgentChat layer (`autogen-agentchat` 0.7.x), an agent is a message handler: you hand it messages, it hands back one `Response`. `AssistantAgent` is the model-backed implementation of that contract. Everything it does in a turn is bounded — it is not an open-ended "keep working until done" loop, and this is the single most common surprise for people arriving from tutorials. ## Step by step 1. **Context assembly.** The incoming messages are appended to the agent's model context. At call time the agent builds the request as its `system_message` (a plain string set in the constructor, defaulting to a generic helpful-assistant line) followed by the messages currently in that context. Any tools passed as `tools=[...]` are converted into tool schemas from the function signature and docstring and sent along. 2. **One model call.** The `model_client` is invoked once. If the model returns plain text, the turn is essentially over: the agent returns a `TextMessage`. 3. **Tool requests.** If the model returns tool calls, the agent emits a `ToolCallRequestEvent` as an inner event and then executes the requested tools. Multiple tool calls requested in the same model response are executed together rather than as separate turns. 4. **Tool results.** Results are emitted as a `ToolCallExecutionEvent`. A tool that raises does not crash the run — the error text becomes the result string the model or the summary sees. 5. **How the turn ends.** This is the branch that matters. ## The reflect_on_tool_use branch - `reflect_on_tool_use=False` (the default when no structured output type is configured): the agent returns a `ToolCallSummaryMessage`. Its content is built from `tool_call_summary_format`, which defaults to `"{result}"` — that is, the *raw* tool return value, verbatim. No second model call is made. This is cheap and fast, and it is why an agent wired to a tool that returns JSON will happily emit JSON as its final answer instead of prose. - `reflect_on_tool_use=True`: the tool results are appended to the context and the model is called a second time to write the reply. You get a `TextMessage` in natural language, at the cost of an extra model call and extra latency. Setting `output_content_type` for structured output flips the default to reflection, because a structured object has to come from the model, not from a tool's raw string. ## The looping branch Historically `AssistantAgent` performed exactly one model call plus one round of tools, and the way to get a multi-step ReAct-style loop was to place the agent in a team that repeatedly gave it turns until a termination condition fired. Recent 0.7.x versions expose `max_tool_iterations` (default `1`): raise it and the agent will keep alternating model call → tool execution inside a single turn until the model stops requesting tools or the limit is reached. Leaving it at the default is a correctness trap when a task genuinely needs two dependent tool calls (look up an id, then use the id) — the agent will stop after the first. ## Statefulness The turn is not stateless. Everything — incoming messages, the assistant reply, tool events — is retained in the agent's model context, so the next turn on the same instance sees all of it. Two consecutive `run()` calls on one agent are a conversation, not two independent requests; clearing that history is an explicit act (`on_reset()`), not a side effect of the turn ending. ## What to check when debugging - The final message is raw tool output → `reflect_on_tool_use` is off. - The agent "gave up" after one tool → tool iterations are capped, or the agent is running standalone rather than inside a team that would give it another turn. - The model never calls the tool at all → the tool's docstring and parameter annotations are the schema; a bare `def f(x)` with no types or docstring gives the model almost nothing to go on. - Costs doubled after a config change → reflection is a second model call per turn, on top of a context that now also carries the tool results. ## Why the design is like this AutoGen's position is that the loop belongs to the orchestrator, not the agent. Keeping the agent's turn bounded makes it composable: a team can interleave agents, apply a termination condition, and stop a runaway conversation, which is impossible if every agent internally loops until it decides it is finished. The knobs on the agent are there for the cases where a bounded internal loop is genuinely simpler than a team.
- Your agent keeps returning raw JSON from a tool as its final answer. What is the fix?That is the default `ToolCallSummaryMessage` path: with `reflect_on_tool_use=False`, `tool_call_summary_format` (`"{result}"`) puts the tool's return value straight into the reply. Either set `reflect_on_tool_use=True` so the model writes prose from the results, or change `tool_call_summary_format` to a template that frames the result. Reflection costs an extra model call per turn, so pick it only when you actually need natural language.
- The task needs two dependent tool calls, but the agent stops after the first. Why?`max_tool_iterations` defaults to 1, so `AssistantAgent` performs one model call plus one round of tool execution and then returns. Raise `max_tool_iterations` so it keeps alternating model call and tool execution within the turn, or place the agent in a team that hands it another turn under a termination condition. Both bound the loop deliberately — an unbounded agent loop is how runs become unbounded in cost.
- How does the model know what a tool does when you pass a plain Python function?AutoGen builds the tool schema from the function itself: the name, the parameter type annotations, and the docstring become the description the model sees. An unannotated, undocumented function produces a nearly empty schema, and the usual symptom is a model that never selects the tool or calls it with nonsense arguments. Treat the signature and docstring as prompt text, not as internal documentation.
saying these in an interview costs you the question
- Says AssistantAgent loops until the task is complete by default
- Assumes the model always sees tool results before replying
- Believes reflect_on_tool_use is enabled by default
- Thinks a tool exception aborts the whole run
- Treats each run() call as a fresh, stateless request