In a multi-agent LLM system, how does an orchestrator agent differ from a subagent?
answer
- one owns the goal, one owns a task
- two windows, not one
- what crosses the boundary is small
- the brief must stand alone
- worker burns tokens, returns a paragraph
basics
~20 sThe orchestrator owns the goal: it splits the work, delegates, and assembles the answer. Each subagent runs its own loop in a separate context window, sees only its assigned task, and hands back a short summary rather than its full transcript.
solid answer
~50 sBoth are the same kind of thing — a model in a loop with tools — so the difference is scope and context, not capability. The **orchestrator** holds the user's goal, decides how to split it, writes a self-contained brief for each piece, and synthesizes the final answer. A **subagent** is spawned for one bounded piece of that work. Critically, it starts a fresh context window containing only its brief and its own tools; it never sees the orchestrator's conversation, and the orchestrator never sees its intermediate steps. What crosses back is a short condensed result — a worker may burn tens of thousands of tokens searching and return one or two thousand. That asymmetry is the whole design: the expensive, noisy exploration is quarantined, and only the distilled finding lands in the window that does the reasoning about the goal.
go deeper
Be able to say plainly that the orchestrator owns the goal and the final answer while each subagent works on one task in its own separate context, returning a short summary.
Explain why the separate window matters: the orchestrator never accumulates the worker's intermediate tool output, so its context stays small and on-topic. Be ready to describe what belongs in a delegation brief.
Show you have designed the boundary — an explicit output contract, references for detail that does not fit the summary, and treating worker output as untrusted. Name what breaks when briefs are underspecified.
Argue about where decision authority sits. Advisory workers with a single deciding orchestrator compose; autonomous workers that each decide the answer's shape produce contradictions no synthesis step can repair.
## Two roles, not two kinds of software An *agent* here means a language model running in a loop: it decides an action, calls a tool, reads the result, and repeats until it is done or out of budget. In a multi-agent system, an **orchestrator** (also called a lead or coordinator) and a **subagent** (worker) are both exactly that. They are frequently the same model, running the same harness code. What distinguishes them is what each one is responsible for and, above all, what each one can see. The orchestrator is responsible for the *goal*. It interprets the user's request, decides how the work divides, decides who does what, decides when enough has been done, and writes the final answer. The subagent is responsible for one *bounded piece* — "find what these five papers say about dosing intervals and report back" — and nothing else. It has no view of the overall goal beyond what its brief tells it, and no authority to decide the shape of the answer. ## Separate context windows are the point A context window is the finite span of tokens a model can attend to on a single call. Every message, tool definition, tool result and reasoning trace consumes part of it. In a single-agent design all of that accumulates in one window: the model's forty-second search result sits next to the user's original request, and the model must attend to both. When the orchestrator spawns a subagent, the subagent begins a *new* conversation. It contains the task brief, the subagent's own tool definitions, and whatever the subagent then discovers. It does not contain the orchestrator's history. Conversely, the orchestrator's window never accumulates the subagent's intermediate observations — the dead ends, the tool errors, the raw page dumps. This is why subagents are described as *context-isolated*. ## The delegation brief must be self-contained Because the subagent cannot see the parent conversation, everything it needs has to be written down explicitly: the objective, the scope boundaries ("only these sources"), the expected output shape, and the tools it may use. A brief like "research this further" produces workers that duplicate each other's effort or drift onto different interpretations of the same task. Under-specified briefs are the single most common cause of bad multi-agent output, and the failure is invisible from the outside — you see a plausible summary that answers a subtly different question. ## The return contract is a summary, not a transcript What comes back is a condensed result: findings, an answer, a short structured report — deliberately not the working. A literature-review worker that read forty abstracts returns a 300-word finding with its citations, not the forty abstracts. Handing the whole transcript back would re-import exactly the noise the split was meant to keep out, and would leave you paying multi-agent prices for single-agent context behaviour. The practical consequence is that anything the orchestrator will need later must appear in the summary. Detail the worker saw but did not report is gone. Well-designed systems make the summary format explicit in the brief, and often have workers persist their raw output somewhere the orchestrator can fetch on demand rather than carrying it inline. ## What stays with the orchestrator Decisions, the plan, the user-facing interaction, and the synthesis all stay with the orchestrator. Subagents advise; the orchestrator decides. Keeping decision authority in one place is what makes the results composable — a set of workers that each independently decided how the answer should look produces contradictions no summary can reconcile. ## Common misreadings *"A subagent is just a tool call."* Superficially similar — you send input, you get output — but a tool is deterministic code with a fixed contract, while a subagent is a whole agentic loop that can plan, call tools, and fail in model-shaped ways. Treat its output as untrusted content, the same as any other model output. *"Subagents must be smaller or cheaper models."* Common, not required. Mixing a strong orchestrator with cheaper workers is one cost strategy, but the role split is about context and responsibility, not model size. *"Multi-agent means parallel."* Parallelism is a frequent *consequence* — independent subtasks can run at once — but a system that spawns one worker at a time is still multi-agent, and plenty of work is strictly sequential and cannot be parallelized at all.
- If the subagent returns only a summary, how does the orchestrator get at detail it later discovers it needs?Either it re-delegates with a sharper brief, or the worker persists its raw output somewhere addressable — a file, a record, an artifact id — and returns the reference alongside the summary so the orchestrator can fetch on demand. Designing for that up front is cheaper than re-running the search, because a re-run costs the full worker loop again and may not reproduce the same findings.
- Should a subagent's summary be trusted more than a tool result?Less, if anything. A tool result is deterministic output from code you control; a subagent summary is model-generated prose that can be confidently wrong, can omit what it did not find important, and can carry injected instructions from untrusted content the worker read. Treat it as untrusted input to the orchestrator, and prefer summaries that carry verifiable references over bare assertions.
- Does the orchestrator have to be an LLM at all?No. If the split is fixed and known ahead of time, plain code can fan work out to model-driven workers and assemble the results — that is a workflow, and it is cheaper, faster and far more predictable. You need a model in the orchestrator seat only when the decomposition itself has to be decided at runtime from what the work reveals.
saying these in an interview costs you the question
- Says subagents share the orchestrator's context window
- Calls a subagent just another tool with a fixed contract
- Assumes multi-agent automatically means parallel execution
- Returns the worker's full transcript to the orchestrator
- Writes vague one-line briefs and expects consistent worker output