How do you bound cost and runtime for an AutoGen team running unattended in production?
answer
- one stop means success, the rest mean failure
- or the fuses together
- boundaries are between messages
- broadcast makes cost superlinear
- the stop reason is a metric
basics
~20 sLayer the stops: one semantic condition that marks real success, ORed with hard fuses on messages, turns, tokens and wall clock, plus limits outside the framework at the model client. Then alert on TaskResult.stop_reason, because a fuse firing routinely means the design is not converging.
solid answer
~50 sTreat termination as policy, not plumbing. Compose an explicit success signal — a `TextMentionTermination` on a token the agents are instructed to emit — with `MaxMessageTermination`, `TimeoutTermination` and `TokenUsageTermination` using `|`, and set `max_turns` on the team as the last-resort fuse. Add an `ExternalTermination` you can `set()` from a supervisor, and pass a cancellation token to `run()` for hard aborts. Recognise the framework-specific cost drivers: teams broadcast, so prompt size grows every turn; `SelectorGroupChat` adds a model call per turn; `MagenticOneGroupChat`'s orchestrator spends extra calls on its ledgers and re-plans when progress stalls, bounded by `max_stalls`. Then instrument: record `stop_reason` on every run and alert when safety nets, rather than the success signal, are what ends most runs. Anything the team cannot bound — per-call token limits, tenant budgets — belongs at the model client or gateway.
go deeper
Know that you must always give a team a termination condition and a turn cap, and that run() reports why it stopped.
Explain how to compose a success signal with message, time and token fuses using |, and why checks happen between messages rather than inside a turn.
Design the whole envelope: layered fuses, external stop and cancellation paths, per-call limits at the model client, and logging stop_reason so a cut-off run is never returned as a finished one.
Own the policy across a fleet: which team type the workload justifies given broadcast and selection overhead, how fuse values are derived from observed successful runs, and what the organisation does when safety nets start firing routinely.
## The framing that matters An unattended multi-agent team is an autonomous loop that spends money per iteration. The engineering question is not "which termination condition do I use" but "what does the system do when the agents fail to converge". A good answer separates three layers. ## Layer 1 — the success path Exactly one stop should mean *done*. In practice that is a `TextMentionTermination` on a token the agents are explicitly instructed to emit, or a `HandoffTermination(target="user")` in a handoff-based team where finishing means returning to a human, or an agent-specific stop such as `SourceMatchTermination`. Choose a token that cannot appear in user input, and put the instruction to emit it in the responsible agent's system message rather than hoping the model volunteers it. ## Layer 2 — the fuses Everything else that can stop a run is a failure to reach layer 1, and each fuse covers a different runaway shape: - `MaxMessageTermination(n)` — bounds the conversation length; the cheapest sanity bound. - `max_turns` on the team — an independent cap that applies even if the condition object is misconfigured. - `TimeoutTermination(seconds)` — bounds wall clock, which is what your caller's SLA actually cares about. - `TokenUsageTermination(...)` — bounds spend directly, but only if the model client reports usage; if it does not, this fuse silently never fires, which is why it must be ORed with a message cap rather than trusted alone. - `ExternalTermination()` — lets an operator or supervising service stop a specific run by calling `set()`. Compose them with `|` so any one is sufficient. Also remember where the checks happen: conditions are evaluated at message boundaries, so a single agent turn with many tool calls runs to completion before any fuse can fire. Fine-grained cost control therefore also needs per-call limits on the model client and, for tool-heavy agents, sane tool timeouts. For a hard abort mid-run, cancel the `CancellationToken` you passed to `run()`. ## Layer 3 — the framework's own cost drivers AutoGen teams broadcast messages, so every participant re-reads a growing thread on its turn. Total spend therefore grows super-linearly in turns; halving the turn cap is worth more than it looks. On top of that, the team type you chose has its own overhead: a positional team adds none, `SelectorGroupChat` adds one selection model call per turn with history in its prompt, and `MagenticOneGroupChat` adds an orchestrator that maintains task and progress ledgers and re-plans when it detects the team is not making progress — bounded by `max_stalls` (default 3), after which it revises its plan rather than looping silently. Choosing the cheapest team type that can express the workload is a bigger lever than tuning any single condition. Code-executing participants deserve their own bounds; unattended execution of model-written code is a blast-radius question, not a cost question, and it is answered with sandboxing and approval gates rather than with termination conditions. ## Layer 4 — the feedback loop The part most teams skip. Every run returns a `TaskResult` with a `stop_reason`; log it as a dimension alongside turn count and token usage. Then the fleet answers questions like: what fraction of runs end on the success signal versus a fuse? Which fuse? Did that shift after a prompt or model change? A team where the message cap fires on most runs is not "safely bounded" — it is failing, and the cap is hiding the failure behind truncated but plausible output. Cutting a run off mid-conversation and returning its partial result to a caller as if it were an answer is the real production hazard here. ## The judgment call Caps are a floor on safety, not a design. If safety nets fire routinely, the response is to change the design — fewer participants, a cheaper team type, a smaller task per run, an explicit completion protocol, or a human gate — not to raise the cap. Conversely, caps set too tight look like quality regressions and push people to disable them. Pick fuse values from the observed distribution of successful runs, with headroom, and revisit them whenever prompts, models or the roster change.
- Why is a token budget alone an insufficient cost control for a team?Conditions are checked between messages, so one agent turn — with its tool calls and a long completion — finishes before the budget is re-evaluated, meaning you can overshoot within a turn. And if the model client does not report usage, the condition never fires at all. OR it with a message cap and a timeout, and enforce per-call limits at the model client or gateway.
- Most runs of your team end with the message cap firing rather than the success signal. Is raising the cap the right response?Usually not. That pattern says the agents rarely reach an explicit completion, so the cap is hiding non-convergence behind truncated output. Investigate first: is the completion signal actually instructed and unambiguous, are there too many participants, is the task too large for one run? Raise the cap only when the successful-run distribution shows it was set below normal work.
- How would you let an operator stop one misbehaving run without restarting the service?Compose an `ExternalTermination()` into that run's condition and expose a control path that calls its `set()`; the run stops at the next check and returns a TaskResult with a stop_reason you can attribute to the operator. For an immediate abort that does not wait for a clean boundary, cancel the `CancellationToken` passed into `run()` or `run_stream()`.
- Which lever reduces cost more: tuning termination values or choosing a different team type?Usually the team type. Broadcast context makes spend grow faster than turn count, and a model-driven selector adds a full extra model call per turn while an orchestrator-style team adds ledger and re-planning calls on top. Moving deterministic routing out of the model, or dropping to a positional team where the flow is fixed, changes the cost curve rather than trimming its tail.
saying these in an interview costs you the question
- Treats a message cap as evidence the run succeeded
- Relies on a token budget the model client never reports
- Assumes a condition can stop an agent mid-turn
- Raises caps whenever safety nets fire instead of investigating
- Ignores that broadcast context makes each turn more expensive than the last