What does AutoGen's TaskResult contain, and where do you read a run's token usage?
answer
- two fields, and neither is a token total
- usage hangs off each message
- None where no model was called
- source tells you which agent spent it
- the string that says why it stopped
basics
~20 sTaskResult holds messages, the full ordered list of the run's messages and events, and stop_reason, a string saying why the run ended. Token usage is not a top-level field: each message carries models_usage with prompt_tokens and completion_tokens, and you sum those.
solid answer
~40 s`TaskResult` from `autogen_agentchat.base` has two fields: `messages`, the ordered list of every chat message and agent event the run produced, and `stop_reason`, a string that is `None` while a run is live and is set to a human-readable reason when the run terminates. There is no `total_tokens` on it. Usage lives per message: each message or event has `models_usage`, either `None` or a `RequestUsage` with `prompt_tokens` and `completion_tokens`. So a run's cost is a sum over `result.messages` of the non-`None` `models_usage` values, and messages that never involved a model call — the initial task, tool-execution events — legitimately report `None`. Per-agent cost falls out of the same loop by grouping on `message.source`. `Console(..., output_stats=True)` prints the same arithmetic for interactive use.
code
python · 18 linesfrom collections import defaultdict
from autogen_agentchat.base import TaskResult
def usage_by_agent(result: TaskResult) -> dict[str, int]:
totals: dict[str, int] = defaultdict(int)
for message in result.messages:
usage = getattr(message, "models_usage", None)
if usage is None:
continue # task input, tool events: no model call
totals[message.source] += usage.prompt_tokens + usage.completion_tokens
return dict(totals)
def ended_cleanly(result: TaskResult) -> bool:
reason = result.stop_reason or ""
return "maximum number of messages" not in reason.lower()go deeper
Know that a run returns a TaskResult with a list of messages and a stop_reason string, and that the answer is normally found among the last messages rather than in a dedicated field.
Explain that token usage is per message via models_usage.prompt_tokens and completion_tokens, that it is None for non-model items, and that a run total is a sum you compute yourself.
Show the operational read: group usage by message.source for per-agent attribution, bucket stop_reason into completed versus limit-hit, and accumulate usage during streaming so a runaway run can be cancelled before it finishes.
Own the measurement contract: decide which per-run facts every team must emit, ensure limit-terminated runs are counted as failures in dashboards, and keep cost attribution at agent granularity so a team's budget can be argued about with data.
## The shape of a result `TaskResult` is intentionally thin. It is a Pydantic model in `autogen_agentchat.base` with: - `messages` — a sequence of every chat message and agent event produced during the run, in order, including the task you submitted as the first item. - `stop_reason` — an optional string describing why the run ended. That is all. Everything an interviewer wants you to extract — the answer, the cost, the agent that spoke last, the tool calls attempted — comes out of `messages`. ## Reading the answer The conventional read is `result.messages[-1]`, which is the last thing the team produced. This is idiomatic but fragile: the last item may be an event rather than a content message, and in a team the last speaker is whichever agent the orchestration happened to land on. Robust code filters by type and by `source` (the agent name string every message carries) rather than trusting the tail. `stop_reason` is your machine-readable-ish record of *why* it ended, and it is the difference between "the team finished the job" and "the team hit a ceiling and was cut off". A run stopped by a message limit and a run stopped because an agent said the magic word both return a `TaskResult` with the same shape; only `stop_reason` distinguishes them. Treat a run whose `stop_reason` mentions a limit as a failure in your metrics, not a success — this is the single most common instrumentation mistake in multi-agent deployments. ## Where usage actually lives Every message and event type in AgentChat carries an optional `models_usage` field. When set it is a `RequestUsage` (from `autogen_core.models`) with two integer fields, `prompt_tokens` and `completion_tokens`. Note what is absent: there is no `total_tokens`, no cost in currency, and no model name on the usage object itself. You add the two numbers yourself, and you map to price yourself from the model you configured. Why per message rather than per run? Because a team run is many model calls by different agents against potentially different model clients. Attributing a single total to the run would hide the thing you actually need, which is *which agent* burned the budget. Grouping by `message.source` gives you a per-agent breakdown from the same loop, and that breakdown is usually the artefact that ends an argument about why a crew cost twenty dollars for one document. `models_usage` is `None` for anything that did not come from a model call: the initial user task, tool execution events, handoff bookkeeping. Code that assumes it is always present crashes on the first run; code that silently skips `None` is correct. ## Streaming and results are the same data The items yielded by `run_stream` are exactly the objects that end up in `messages`, so you can accumulate usage live rather than waiting for the run to finish — which is what you want if you intend to cancel a run that has already blown its budget. Reading usage only from the final `TaskResult` means you learn the cost after you have paid it. ## Practical instrumentation A minimal but sufficient set of metrics per run: total prompt and completion tokens, tokens by agent name, message count, wall-clock duration, and `stop_reason` bucketed into "completed" versus "limit hit". Those five make the difference between a team you can operate and one you can only rerun. Everything richer — spans, per-tool latency, prompt payload capture — belongs in tracing rather than in this object. ## Version note This describes `autogen-agentchat` 0.7.x, the redesigned API. Older 0.2-era code returned chat histories from `initiate_chat` and tracked cost with separate helper utilities; that surface no longer applies.
- Why is models_usage None on some of the messages in a completed run?Because not every item came from a model call. The initial task you submitted, tool execution events, and handoff bookkeeping are recorded in the message list for provenance but never hit an LLM, so they carry no `RequestUsage`. Summation code must skip `None` rather than assume the field is populated.
- You want to abort a run once it passes a token budget. Where do you enforce that?Accumulate usage while consuming `run_stream` rather than reading the final `TaskResult`, and cancel through a `CancellationToken` when you cross the threshold. Reading usage from the result tells you the cost after you have already paid it. A token-usage termination condition covers the in-framework case, but external budget enforcement needs the live stream.
- Is result.messages[-1] a safe way to get the team's answer?Only loosely. The tail may be an event rather than content, and in a group chat the final speaker depends on orchestration. Filter by message type and by `source` for the agent you expect to produce the deliverable, or have that agent write to a known output channel. Relying on the tail breaks the first time you add a participant.
saying these in an interview costs you the question
- Looking for a total_tokens field on TaskResult
- Assuming models_usage is set on every message
- Counting a run stopped by a message limit as a success
- Reporting one run-level cost instead of per-agent attribution
- Believing RequestUsage carries the model name or price