skip to content

Orchestrator-Worker

A lead agent decomposes the goal, fans the pieces out to parallel workers with isolated contexts, then synthesizes their results. It is the workhorse pattern for research and large-codebase tasks, and the one most likely to come up in an interview.

part ofMulti-agent LLM systemsoverview, primer and where to startread it →
on this pageshow

questions

5

In an orchestrator-worker agent system, what stays in the lead agent's context?

level: middleimportance: must knowfreq 75%

answer

  1. two windows, one wall
  2. the lead never sees the raw material
  3. only condensed returns cross back
  4. grows with subtask count, not corpus size
  5. workers are blind to each other

basics

~20 s

The lead holds the goal, its decomposition into subtasks, and the short results workers hand back. Each worker's own reasoning, tool calls and raw source material stay inside that worker's separate window and never reach the lead.

solid answer

~50 s

The lead agent owns the **goal, the decomposition and the synthesis**. It splits the objective into subtasks, spawns one worker per subtask, and each worker runs in its **own fresh context window** with its own instructions, its own tool budget and its own transcript. The expensive material — full documents, page after page of tool output, the worker's step-by-step reasoning — lives and dies inside the worker. What crosses back is only a **condensed return**: typically a few hundred to a couple of thousand tokens of findings in the shape the lead asked for. So the lead's context grows with the *number* of subtasks, not with the *volume of material* the system read. That is the structural point of the pattern: the system can process far more input than one window holds, at the price that the lead never sees the evidence first-hand.

go deeper

for a junior

Be able to say plainly that each worker gets its own separate context window and that only a short summary comes back to the lead. Naming the two roles and the boundary between them is enough at this level.

for a middle

Explain the mechanics: what the lead's window accumulates, what dies with the worker, and why total input can exceed any single window. Interviewers expect you to state the return-contract idea unprompted.

for a senior

Show the judgment side — that the lead's knowledge is only as good as the returns, that provenance must be designed in, and that fan-out width has to be capped before the lead's own window becomes the bottleneck.

for a principal

Own the trade: isolation buys capacity and parallelism with a multiplied token bill and a lead that cannot audit evidence directly. Be ready to argue when that price is worth paying and where you would keep writes single-threaded.

## The shape of the pattern Orchestrator-worker (also called lead/subagent, or fan-out/gather) is a hierarchical topology with exactly two roles. One **lead agent** receives the goal and never leaves the driver's seat: it decides how to cut the goal into pieces, launches one **worker agent** per piece, waits, and writes the final answer from what comes back. Workers do not talk to each other, do not talk to the user, and do not decide what happens next. They answer one bounded question and terminate. A concrete case: a newsroom editor agent is asked what eight city councils decided about short-term rentals this month. It fans out to eight reporter workers, one per council transcript, each asked for a 200-word brief with two quotes. The editor then writes the cross-city story. ## Two kinds of context, and the wall between them The lead's window accumulates: the user's request, the plan, the worker returns, and the lead's own reasoning about them. It is long-lived and it is the scarce resource. Each worker's window is created empty, filled with exactly what the lead handed it plus whatever the worker's own tools pull in, and then discarded. A worker that reads a 60,000-token transcript burns 60,000 tokens *of its own window*. The editor never pays for those tokens and never sees that text. It sees 200 words. This is why the pattern scales to corpora no single window can hold. Eight transcripts at 60k tokens each cannot fit one 200k window; eight isolated workers handle them comfortably and the lead assembles 1,600 words of briefs. ## What the lead can and cannot know The cost of the wall is **epistemic**. The lead's picture of the world is entirely mediated by the returns. It cannot re-read a passage a worker glossed over, cannot check a quote against the source, and cannot tell a confident brief from a hallucinated one without spending a second call to go look. Anything the worker omitted is, from the lead's perspective, non-existent. Workers are also blind to each other. Worker 3 knows nothing of worker 5's findings unless the lead passes it forward in a later round. That independence is what makes the fan-out parallelizable, and it is also why two workers will happily produce overlapping or contradictory briefs. ## Practical consequences - **The return contract is load-bearing.** Because the return is all the lead ever gets, its format, length cap and required fields are part of the system's design, not a nicety. Ask for structure — claim, evidence, source, confidence — rather than free prose. - **Keep provenance in the return.** A brief that carries the document id and line range lets the lead spot-check or re-dispatch; one that carries only conclusions cannot be audited. - **The lead can still fill up.** Fan out to fifty workers with 2,000-token returns and you have re-created the context problem one level up. Deep fan-outs usually need a second-stage merge, or tighter caps. - **Multi-round fan-outs are normal.** A lead often runs a cheap scoping round, reads the returns, then fans out again with sharper subtasks. ## Cost Isolation is not free. Every worker pays its own system prompt, tool definitions and reasoning overhead, and the lead pays a full pass over the returns. A fan-out of eight commonly costs several times what one agent reading the same material sequentially would cost — the trade is latency and window capacity bought with tokens. As of mid-2026 that multiplier is treated as a first-class design constraint rather than an implementation detail, and it is the main reason teams reserve the pattern for genuinely parallel, read-heavy work. ## The rule of thumb that survived The pattern holds up best when workers **contribute intelligence rather than actions**: they read, search, analyse and report, while writes and side effects stay with the lead or a single designated writer. Parallel workers all mutating the same repository, ticket queue or document is the failure shape that repeatedly did not work in practice.

  • If the lead never sees the source text, how can it verify a worker's claim?
    Only by paying for it. Either require provenance in the return — document id, line range, verbatim quote — so the lead can re-dispatch a cheap targeted read, or run a verification worker with a clean context over the specific claim. Both cost an extra call; a lead that simply trusts the brief inherits every worker hallucination silently.
  • What happens to the lead's context if you fan out to fifty workers?
    You re-create the context problem one level up. Fifty returns at a couple of thousand tokens each will fill a large window on their own, and the lead's synthesis quality degrades as the window lengthens. The fixes are tighter return caps, a hierarchical merge stage that condenses batches of returns before the lead reads them, or fewer, larger subtasks.
  • Can workers share findings with each other mid-run?
    Not in the pure topology — that is what makes the fan-out parallel and reproducible. If workers genuinely need each other's output, either the dependency belongs in an earlier sequential round, or the workload is not a fan-out at all. A common middle ground is a shared read-only artifact the lead writes between rounds.

saying these in an interview costs you the question

  • Thinking workers share one conversation history with the lead
  • Assuming the lead can re-read what a worker read
  • Believing isolation is free rather than a token trade
  • Letting workers return raw dumps instead of condensed findings
  • Expecting workers to coordinate with each other automatically

context

open as a page

Eight parallel research workers return overlapping, uneven briefs — what do you fix first?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Fix the worker contract before anything else. Each worker needs an explicit objective, an exclusive slice of the sources, an allowed tool set, and a required output shape with a length cap. Overlap and unevenness are almost always a lead that copied the goal instead of partitioning it.

open as a page

Two of eight parallel worker agents fail — should the orchestrator synthesize, retry, or abort?

level: middleimportance: should knowfreq 44%

basics

~20 s

It depends on whether the missing slices are load-bearing. Retry the two once with a bounded budget; if they still fail, synthesize from the six and state the gap explicitly, and abort only when the task is invalid without complete coverage. Never silently answer from partial data.

open as a page

When two worker agents return contradictory briefs, how should the orchestrator synthesize them?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Treat contradiction as a signal to resolve, not noise to average. Compare the returns' evidence and provenance, prefer the claim with a locatable source, and if neither is decisive re-dispatch a narrow worker to settle it or report the disagreement explicitly rather than silently picking one.

open as a page

How do you decide how many parallel workers an orchestrator should fan out to?

level: principalimportance: should knowfreq 38%

basics

~20 s

Let the work decide, not the model. Width should follow the number of genuinely independent slices, then be capped by three ceilings: the token budget, which grows roughly linearly with width; the lead's own window, which must hold every return; and provider concurrency limits.

open as a page