skip to content

How does a supervisor agent decide which specialist runs next each turn?

level: middleimportance: must knowfreq 62%

answer

  1. one decision per turn
  2. the roster lives in the prompt
  3. scope lines, not just names
  4. constrain output to known names
  5. finish is one of the choices

basics

~20 s

A supervisor is an LLM prompted with the roster of specialists, a one-line scope for each, and the conversation so far. Every turn it emits one constrained choice - the next specialist's name, or a finish signal - and the runtime dispatches accordingly.

solid answer

~50 s

The supervisor is itself a model call whose only job is a **routing decision**. Its prompt carries the goal, the running conversation or a condensed ledger of what has happened, and a roster where each specialist gets a name plus a one-line "use this when..." scope with a couple of concrete example requests. The output is constrained to a small closed set - the agent names plus an explicit finish value - so the dispatcher can match it rather than parse prose. After the chosen specialist returns, control comes back to the supervisor and the loop repeats, which is what makes this topology centralized: the supervisor re-decides every turn instead of committing to a plan up front. In a city 311 system, the same supervisor routes one message to the potholes agent, the next to permits, and then declares the case resolved.

code

json · 15 lines
json
{
  "name": "route_next",
  "description": "Choose the single specialist that should act next, or finish.",
  "input_schema": {
    "type": "object",
    "properties": {
      "next": {
        "type": "string",
        "enum": ["potholes", "permits", "sanitation", "housing", "FINISH"]
      },
      "reason": { "type": "string", "maxLength": 200 }
    },
    "required": ["next", "reason"]
  }
}

go deeper

for a junior

Be able to say plainly that a supervisor is an agent that only chooses who acts next, that control comes back to it after each specialist, and that the specialists' descriptions live in its prompt.

for a middle

Explain the anatomy of the routing prompt - goal, roster with scope lines, state, policy - and why the decision is a constrained value from a closed set rather than free text. Name finish as one of the options.

for a senior

Show how you diagnose bad routing in production: label a sample of requests with the correct specialist, score the router's first choice, and read traces for oscillation or self-answering. Tie fixes to descriptions and state shape, not to swapping models.

for a principal

Own the tradeoff between routing quality and roster size. Argue when a router is the right control plane at all versus a deterministic workflow, and set the policy for how completion criteria are written and audited across teams.

## What a supervisor actually is In a supervised multi-agent system there is one agent - the supervisor, sometimes called the router or lead - that never does the domain work itself. Its entire job is to look at the current state of the task and answer one question: *who acts next?* The specialists (a potholes agent, a permits agent, a sanitation agent, a housing agent in a city 311 triage system) each own a narrow slice of capability and tools. The supervisor owns the control flow. The defining property is that control returns to the supervisor after every step. A specialist runs, produces a result, and the loop comes back to the router, which decides again. That is different from a chain (fixed order, decided at build time) and different from peer-to-peer transfer (the specialist itself decides who is next). Centralization is the whole point: one component knows the goal, sees the progress, and can be audited. ## The routing prompt A routing prompt has four parts, and interviewers probe each. 1. **The goal.** The user's request, restated stably so it survives many turns. Supervisors that lose the goal start routing to plausible-but-irrelevant specialists. 2. **The roster.** Each specialist gets a name and a scope line: what it handles, what it does *not* handle, and one or two concrete example requests. Names alone are the classic failure - "housing" and "permits" both sound right for "can I convert my garage into a rental?" until the scope lines disambiguate them. 3. **The state.** Either the raw conversation, or better, a condensed ledger: what has been tried, what each specialist returned, what remains open. 4. **The routing policy.** Route to exactly one agent per turn; do not answer the user directly; do not re-dispatch an agent that already answered unless there is new information; finish when the stated completion criteria are met. ## Constraining the decision The supervisor's output should be a structured value from a closed set, not free prose. Providers expose this as constrained or structured output - an enum of the known agent names plus a `FINISH`-style value, often alongside a short free-text reason for the trace. Two things go wrong without it: the model invents an agent name the dispatcher cannot match, and the model starts *answering* instead of routing, quietly turning your multi-agent system into a single agent that hallucinates specialist knowledge. Keeping a one-line rationale field next to the choice is cheap and makes traces reviewable, but the rationale must never be the thing the dispatcher parses. ## Termination is a routing option A supervisor-routed system ends when the supervisor says it ends. That is a real design responsibility, not an afterthought. If "done" is not one of the choices the router can emit, the model has no way to express completion and will keep dispatching until an external step cap fires - which looks to the user like a system that never converges. Give the supervisor explicit completion criteria ("finish when the citizen's request has a case number and an owning department") and pair the model-level finish signal with a hard cap as a safety net, not as the primary stop. ## How routing goes wrong - **Overlapping scopes** cause oscillation: the supervisor sends the same request to permits, gets a partial answer, sends it to housing, gets a partial answer, and ping-pongs. Fix the scope lines before you fix the model. - **Roster growth** degrades selection accuracy. A router choosing among four specialists is reliable; among twenty, description quality dominates and errors climb. - **Stale state** makes the supervisor re-dispatch work already done. A ledger of completed steps is a stronger input than a raw transcript. - **Silent self-answering** - the router replies to the user itself. Detect it in traces by counting turns where no specialist ran. ## What good looks like A healthy supervised run shows a short, legible trace: one routing decision, one specialist turn, one routing decision, one specialist turn, then finish - typically a handful of hops, each with a recorded reason. Instrument routing accuracy directly by labelling a sample of production requests with the specialist that *should* have handled them and scoring the router's first choice against that label. That single metric tells you whether your problem is the router prompt or the specialists themselves.

  • Why constrain the supervisor's output to an enum of agent names rather than letting it write the name in prose?
    Because the dispatcher has to match the value deterministically. Free prose lets the model name an agent that does not exist, hedge between two, or slide into answering the user itself instead of routing. A constrained value plus a separate short rationale field keeps the trace readable while making the decision machine-checkable.
  • The supervisor keeps bouncing a request between two specialists. What do you look at first?
    The scope lines, not the model. Ping-pong almost always means two specialists' descriptions both plausibly match the request. Rewrite them with explicit exclusions and concrete example requests, and add a rule that an agent already consulted on the same open question is not re-dispatched without new information. Cap consecutive routes as a backstop.
  • Should the supervisor see every specialist's full reply, or a summary?
    A summary, in almost every case. The router only needs enough to decide who acts next and whether the goal is met. Passing full replies through inflates the supervisor's context every turn, raises the cost of each routing decision, and degrades routing quality as the window grows. Define a short return contract per specialist.

saying these in an interview costs you the question

  • Says the supervisor does the domain work rather than dispatching
  • Assumes agent names alone are enough for correct routing
  • Claims the supervisor picks the whole sequence once up front
  • Leaves no finish option, so runs only end on a step cap
  • Ignores that overlapping specialist scopes cause oscillation

context