skip to content

Multi-Agent Orchestration

Splitting work across an orchestrator and subagents that each keep their own context and hand back a short summary. Interviewers ask when that isolation is worth its coordination cost and token bill.

part ofAI agentsoverview, primer and where to startread it →
on this pageshow

questions

5

Why is context isolation, not extra reasoning, the main gain from LLM subagents?

level: middleimportance: must knowfreq 70%

answer

  1. same model, so no new intelligence
  2. the win is smaller windows
  3. noise stays where it was generated
  4. context rot, lost in the middle
  5. measure the compression ratio

basics

~20 s

Spawning subagents adds no reasoning capacity — it is the same model. The gain is that each worker's noisy intermediate output stays in its own window, so the orchestrator attends to a small set of distilled findings instead of a bloated transcript full of dead ends.

solid answer

~50 s

Running five subagents does not give you five times the intelligence: they are typically the same model, so nothing new arrives in the way of reasoning ability. What changes is the token diet each call sees. A worker that greps a corpus, hits three dead ends and reads twenty documents might burn 50k tokens; it returns a 1–2k summary, and the orchestrator's window absorbs only that. This matters because model quality degrades as context fills — the effect usually called *context rot*, with retrieval accuracy dipping for material buried mid-window (*lost in the middle*), and irrelevant or contradictory material actively pulling answers off course. Isolation keeps every window small and high-signal. The corollary is diagnostic: if a task does not generate much intermediate noise, splitting it buys you almost nothing and you have paid the coordination cost for free.

go deeper

for a junior

Know that subagents are usually the same model, so splitting adds no intelligence. The point is that each one works in its own smaller context and reports back briefly.

for a middle

Explain why smaller windows produce better answers — context rot, lost-in-the-middle retrieval decay, and earlier mistakes conditioning later steps. Contrast that with the false 'more agents, more reasoning' story.

for a senior

Show the decision rule you apply: estimate the subtask's compression ratio and its coupling, and reach for compaction, context editing or offloading by reference before paying for isolation.

for a principal

Own the information-loss argument. Isolation is a lossy interface you are imposing on your own system; decide deliberately what must survive the boundary and what your system does when two workers return incompatible conclusions.

## The claim, stated precisely When an orchestrator delegates to several subagents, the improvement you measure does not come from having "more minds on the problem." The workers are usually instances of the same model, with the same weights and the same reasoning ability. Nothing about spawning a second copy makes either copy smarter. What multi-agent buys is a *context* property: work that would have piled up in one window is partitioned across several, and only condensed results cross the boundaries. This distinction is the single most testable thing about the topic, and it is where weak answers collapse. "Five agents means five perspectives" sounds right and is mostly wrong — five identical models given the same brief produce five correlated outputs, not five independent perspectives. ## Why context size degrades quality Attention is finite and roughly competitive: every token in the window competes for the model's attention against every other. Three consequences follow, and they have standard names in the field. **Context rot** — as a window fills, accuracy on the material it contains declines, even well below the nominal limit. A model that answers reliably at 10k tokens of context is measurably shakier at 150k, on the same question. **Lost in the middle** — retrieval of a specific fact is best when it sits near the beginning or the end of the window and worst when it is buried in the middle. Long transcripts bury things by construction. **Context poisoning / distraction** — a wrong intermediate conclusion, a failed tool call, or a page of irrelevant scraped text does not sit inertly. It conditions everything that follows. An agent that guessed wrong at step four will keep referring to that guess at step forty. A long single-agent run accumulates all three. Every dead-end search, every oversized tool result, every abandoned hypothesis stays in the record. ## What isolation actually does about it Give the same work to an orchestrator and three workers and the arithmetic changes. Each worker starts clean, holds only its brief and its own findings, and finishes before it has time to bloat. The orchestrator never sees the twenty documents — it sees three paragraphs. Its window stays small, so the reasoning that matters most (what does all this mean for the user's actual question?) happens under the best conditions available, not at the end of a 150k-token slog. A concrete shape: a scientific-literature-review orchestrator splits a question across five workers, each assigned a disjoint set of papers. Each reads its set — call it 40–60k tokens of full text and search results — and returns a 300-word finding with citations. The orchestrator ends up with roughly 2k tokens of high-signal input covering material that would have overflowed a single window entirely. The workers' dead ends, retracted PDFs and irrelevant hits never touch it. ## Isolation is one tool among several Context isolation sits in the same family as the other context-engineering techniques, and interviewers like to hear you place it: - **Compaction** — summarize the conversation so far and restart from the summary. Same idea applied to time rather than to workers, and lossy in ways you do not control. - **Context editing** — clear stale tool results or old reasoning out of the window rather than summarizing them. - **Just-in-time retrieval** — hold identifiers or file paths instead of contents, and fetch at the moment of use. - **Structured note-taking** — write findings to durable storage and re-read the notes rather than re-reading the transcript. Subagent isolation is the most expensive of these, because it duplicates system prompts, tool definitions and task context per worker. Reach for the cheaper ones first: if compaction or offloading tool output by reference solves your context pressure, you do not need a multi-agent system at all. ## The diagnostic that follows If isolation is the mechanism, then the payoff scales with how much noise the subtask generates. That gives a usable test before you split anything: *how many tokens does this subtask consume relative to what it needs to report back?* A worker that reads a hundred pages to answer one question has a compression ratio of maybe 50:1 and is an excellent candidate. A worker that makes one API call and returns the response verbatim compresses nothing — you have paid for an extra model round-trip and a coordination step to achieve exactly nothing. This is also why multi-agent shines on breadth-first search and exploration, and disappoints on tightly coupled work. Exploration is noise-generating by nature; a sequential refactor where each step depends on the last is not, and splitting it merely means every worker lacks the context the previous step produced. ## Two honest caveats First, isolation costs information. Anything a worker saw but did not report is unrecoverable to the orchestrator, and workers are poor judges of what will matter later. Second, isolated workers cannot coordinate mid-flight — two of them can independently reach incompatible conclusions, and no amount of summary quality lets the orchestrator tell which one was right.

  • If isolation is the mechanism, how do you decide up front whether a given subtask is worth splitting out?
    Estimate its compression ratio: tokens consumed versus tokens it must report back. Reading a hundred documents to produce a paragraph compresses roughly 50:1 and is an excellent candidate. A step that makes one call and returns the payload compresses nothing, so you pay an extra model round-trip plus coordination for zero context benefit. High-noise, high-compression, low-coupling work is where the gain lives.
  • What information do you lose by isolating a worker's context, and how do you mitigate it?
    Everything the worker saw but chose not to report, and workers judge relevance badly against a goal they cannot see. Mitigate by specifying the output contract in the brief so the important dimensions are named explicitly, and by having workers persist raw output to durable storage and return references — so the orchestrator can fetch detail on demand instead of re-running the search.
  • Isn't compaction a cheaper way to get the same benefit?
    Often, yes, and it should be tried first. Compaction summarizes a filling window and restarts from the summary; context editing clears stale tool results outright. Both keep one window small without duplicating system prompts and tool definitions across workers. Subagent isolation earns its extra cost mainly when subtasks are genuinely independent and can run concurrently — otherwise you are buying compaction at several times the price.

saying these in an interview costs you the question

  • Claims more agents means more reasoning capacity or IQ
  • Says identical models given the same brief give diverse perspectives
  • Ignores that isolation loses everything the worker did not report
  • Splits tightly coupled sequential work and expects a gain
  • Reaches for subagents before trying compaction or offloading by reference

context

open as a page

In a multi-agent LLM system, how does an orchestrator agent differ from a subagent?

level: juniorimportance: should knowfreq 55%

basics

~20 s

The orchestrator owns the goal: it splits the work, delegates, and assembles the answer. Each subagent runs its own loop in a separate context window, sees only its assigned task, and hands back a short summary rather than its full transcript.

open as a page

How would you choose between LangGraph, CrewAI and a vendor agent SDK?

level: middleimportance: should knowfreq 50%

basics

~20 s

Match the tool to the control you need. LangGraph gives explicit graph control flow with durable checkpointed state for production. CrewAI reaches a working prototype fastest with role-based crews. Vendor SDKs are the shortest path to that provider's newest capabilities.

open as a page

Why do multi-agent systems need a single-writer rule for each shared artifact?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Isolated workers cannot see each other's decisions, so two of them editing the same artifact make incompatible implicit choices that the orchestrator cannot reconcile from summaries. Keep writes with one designated agent; let the rest return read-only findings or proposals.

open as a page

Multi-agent research systems burn ~15x the tokens of a chat — when does that pay?

level: principalimportance: should knowfreq 50%

basics

~20 s

It pays when the work is read-heavy, splits into independent parallel pieces, and the result can be checked — research, broad search, breadth-first exploration. It does not pay for latency-sensitive, high-volume or tightly sequential tasks where a single agent or plain workflow is adequate.

open as a page