skip to content

Orchestration Topologies

The wiring diagrams: hierarchical orchestrator-worker, peer-to-peer, pipeline, and fan-out/gather, along with the routing and delegation decisions each implies. Interviewers ask you to pick a topology for a concrete workload and defend how centralized it is.

part ofMulti-agent LLM systemsoverview, primer and where to startread it →
on this pageshow

explore

questions

19

In multi-agent debate, what do extra rounds buy beyond a single critique pass?

level: middleimportance: must knowfreq 62%

answer

  1. gains front-loaded, then flatten
  2. round two carries most of the lift
  3. convergence is not correctness
  4. agents herd toward the confident voice
  5. rounds are sequential, so latency stacks

basics

~20 s

Extra debate rounds let each agent see the others' revised arguments and change its own, which corrects confident first-shot errors. Most of the lift lands in the first one or two exchanges; after that agents converge and mostly restate a consensus that may be wrong.

solid answer

~50 s

A single critique pass gives the proposer one outside opinion; debate iterates it. Each agent answers, reads the other answers, and revises, so a claim has to survive repeated challenge instead of one review. The published multi-agent debate work (Du et al., framed as a "Society of Minds") shows the gain concentrated in the early exchanges on reasoning and factuality tasks, then flattening. Extra rounds mostly buy convergence, not correctness: agents drift toward whichever position is stated most confidently, and if they share a base model and prompt they can herd onto the same wrong answer. Practically, I cap rounds at two or three, stop early when answers stop changing, and force diversity by giving debaters different prompts or different base models. Each round multiplies both tokens and wall-clock latency, because rounds are strictly sequential.

go deeper

for a junior

Be able to say what a debate round is: several agents answer separately, then read each other and revise. Knowing that this costs one full generation per agent per round already puts you ahead of a vague "multiple agents check each other" answer.

for a middle

Explain the mechanism and the shape of the curve — most of the improvement arrives in the first exchange and flattens after two or three rounds — and name herding as the reason agreement does not prove correctness. Interviewers expect you to state the token and latency multiplier without prompting.

for a senior

Show you have run this: a round cap, an early-stop rule when answers stop changing, deliberate decorrelation of debaters, and a rule that a deterministic verifier always beats an argument. Be ready to say how you would detect a debate that converged on a confident wrong answer in production.

for a principal

Own the framing that debate is one way to spend a token budget, competing with a stronger single model, better grounding, or a verifier. Expect to defend applying it only to a selected slice of traffic and to explain what you would measure before making it the default path.

## What multi-agent debate is In multi-agent debate, several model instances are asked the same question and produce answers independently. Their answers are then shown to each other, and every agent is asked to reconsider in light of what the others said. That cycle — answer, read peers, revise — is one round. After a fixed number of rounds, the final answers are collapsed into one output, either by taking the majority position or by handing the transcript to a separate judge agent. The idea entered the field with work on improving factuality and reasoning through multi-agent debate (Du et al.), often described under the "Society of Minds" framing: reasoning emerges from disagreement between independent processes rather than from one process thinking longer. ## A single critique pass versus a debate A single critique pass is asymmetric: one agent proposes, one agent reviews, and the proposer either accepts the review or does not. Only one opinion ever enters the system, and if the critic is wrong the error is now laundered as a correction. Debate is symmetric and iterated. Every participant is both proposer and critic. A wrong claim has to survive being contradicted, then survive the contradiction being contradicted. That is the mechanism: repeated challenge filters claims that only one agent's particular sampling path supported. ## Where the gain actually comes from The first exchange carries most of the value, because that is the moment an agent first encounters evidence it did not generate. An agent that hallucinated an intermediate step in an arithmetic or multi-hop factual chain frequently abandons it once two peers show a different chain. This is why debate helps most on tasks where errors are idiosyncratic — each run fails differently — and helps least on tasks where the model has a systematic bias that every instance shares. By the second and third round, the answers being exchanged are already heavily influenced by the previous round, so the marginal new information per round shrinks. Published curves flatten quickly. Beyond three rounds you are usually paying full price for restatement. ## Convergence is not correctness The most important failure mode is herding. Debate reliably produces agreement; it does not reliably produce truth. Agents tend to defer to the most assertively phrased position, and models that share training data and prompt framing share priors, so a confident wrong answer can pull the whole panel. A transcript that ends in unanimous agreement therefore tells you the process terminated, not that the answer is right. If you report confidence to a downstream consumer, do not derive it from unanimity alone. A related pathology is instant collapse: on round one every agent already agrees, so the remaining rounds are pure cost. Detect this and terminate. ## Keeping disagreement real Several levers keep the panel from being one voice repeated: - Have every agent commit to a private first answer before seeing any peer output, so the first round is genuinely independent. - Vary the debaters: different system prompts, different personas with different priorities, or genuinely different base models. Different models are the strongest source of decorrelated error; temperature alone is the weakest. - Require each revision to name the specific claim it is disputing, rather than allowing a bare concession. - Give the debaters access to a retrieval or execution tool so at least some claims can be settled against evidence instead of rhetoric. ## When debate is the wrong tool If the task has a cheap deterministic verifier — the code compiles or it does not, the SQL runs, the JSON validates, the invariant holds — run the verifier. Verification is orders of magnitude cheaper than argument and is actually correct rather than persuasive. Debate earns its keep on judgement calls and on reasoning chains where no checker exists. Debate is also weak on open-ended generation with no ground truth. There, the panel converges on the most conventional phrasing, which reads like consensus and is really regression to the mean. ## The cost, stated plainly Debate multiplies generations. Three agents over three rounds is nine full generations before you even run a judge, and each round's prompt carries the previous rounds' transcript, so per-call input tokens grow as the debate proceeds. Latency is worse than the token count suggests, because rounds cannot be parallelised — only the agents within a round can. A structure that raises accuracy by a few points while multiplying cost by an order of magnitude has to be justified against simply spending that budget on a stronger single model or better grounding. ## What a good answer sounds like Name the mechanism (independent answers, then mutual revision), say where the gain concentrates and why it flattens, name herding as the reason convergence is not evidence, and finish on the round cap plus early-stop rule and the sequential-latency cost. That is the full shape an interviewer is listening for.

  • What signal would you use to stop a debate early?
    Stop when the round-over-round answers stop changing — no agent revised its position, or the diff between rounds is cosmetic. Some systems ask each agent for a short structured verdict alongside its argument so consensus is machine-checkable rather than judged by a model. Always pair the early-stop rule with a hard round cap, so a pair of agents that keep flip-flopping cannot run forever.
  • How do you stop a debate collapsing into agreement on the first round?
    Make the first answer private: every agent commits before seeing any peer output. Then decorrelate the debaters — different system prompts, different priorities, ideally different base models, since same-model agents share priors and agree too cheaply. Requiring each revision to cite the specific claim it disputes, rather than allowing a bare concession, also keeps the disagreement substantive.
  • Would you use debate on a task that has a unit-test suite as ground truth?
    No. When a deterministic verifier exists, run it — the tests settle the question for a fraction of a single generation, and they are actually right rather than merely persuasive. Debate is for judgement calls and reasoning chains with no checker. A sensible hybrid gives the debaters the verifier as a tool so their arguments are grounded in run results instead of rhetoric.

It is like a review panel that keeps talking until everyone agrees: the first exchange surfaces the real objections, and the later rounds mostly manufacture unanimity — including, sometimes, unanimity about something false.

saying these in an interview costs you the question

  • Claims accuracy keeps rising with every additional debate round
  • Treats unanimous agreement as evidence the answer is correct
  • Uses debate on tasks that already have a cheap deterministic verifier
  • Gives every debater the same model and prompt and still expects diversity
  • Ignores that rounds are sequential, so latency scales with round count

context

open as a page

In an agent handoff, what actually changes when one agent transfers control to a peer?

level: middleimportance: must knowfreq 68%

basics

~20 s

The active agent swaps. The transfer is exposed to the model as an ordinary tool call, and once it fires, the next turn runs under the peer's system prompt, tool set and model — on the same live conversation, with no third party in between.

open as a page

In an orchestrator-worker agent system, what stays in the lead agent's context?

level: middleimportance: must knowfreq 75%

basics

~20 s

The lead holds the goal, its decomposition into subtasks, and the short results workers hand back. Each worker's own reasoning, tool calls and raw source material stay inside that worker's separate window and never reach the lead.

open as a page

How does a supervisor agent decide which specialist runs next each turn?

level: middleimportance: must knowfreq 62%

basics

~20 s

A supervisor is an LLM prompted with the roster of specialists, a one-line scope for each, and the conversation so far. Every turn it emits one constrained choice - the next specialist's name, or a finish signal - and the runtime dispatches accordingly.

open as a page

How do you detect and break a handoff loop between two agents in a swarm?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Log every transfer as a structured event, then enforce a hop budget and a repeated-pair check on that log. When the budget is spent or the same two agents trade the conversation twice, stop transferring and escalate to a human or a generalist rather than letting the cycle continue.

open as a page

Eight parallel research workers return overlapping, uneven briefs — what do you fix first?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Fix the worker contract before anything else. Each worker needs an explicit objective, an exclusive slice of the sources, an allowed tool set, and a required output shape with a length cap. Overlap and unevenness are almost always a lead that copied the goal instead of partitioning it.

open as a page

What role does a triage agent play in a handoff-based support swarm?

level: juniorimportance: should knowfreq 48%

basics

~20 s

A triage agent is the swarm's front door. It holds the opening turns, works out what the customer actually needs, and then transfers the conversation to the specialist peer that owns that need instead of answering itself.

open as a page

Why give a critic agent a fresh context instead of the author agent's thread?

level: middleimportance: should knowfreq 45%

basics

~20 s

A critic started in a fresh context judges the artifact rather than the story that produced it. Continuing the author's thread anchors the critic on the author's assumptions and inherits a long, degraded window, so it tends to ratify work instead of finding defects.

open as a page

Two of eight parallel worker agents fail — should the orchestrator synthesize, retry, or abort?

level: middleimportance: should knowfreq 44%

basics

~20 s

It depends on whether the missing slices are load-bearing. Retry the two once with a bounded budget; if they still fail, synthesize from the six and state the gap explicitly, and abort only when the task is invalid without complete coverage. Never silently answer from partial data.

open as a page

In a supervisor system, how does agent-as-tool differ from agent-as-graph-node?

level: middleimportance: should knowfreq 44%

basics

~20 s

Agent-as-tool exposes each specialist as a callable the supervisor's model invokes, so control always returns and the reply lands in the supervisor's context as a tool result. Agent-as-graph-node makes each specialist a node in an explicit state machine, with edges deciding where control goes and shared state carrying the work.

open as a page

When does ensembling independent LLM samples stop improving accuracy?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Aggregation only cancels errors that are independent. Samples drawn from one model on one prompt fail in the same direction, so five runs agreeing means the model is consistent, not correct — and more samples then only sharpen a biased estimate.

open as a page

How do you design a judge agent that picks between two agents' answers?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Give the judge the rubric and the underlying evidence, not just the two arguments, so it checks claims instead of rating rhetoric. Constrain its output to a verdict plus cited support, blind and shuffle the candidates, and route low-confidence cases to a human.

open as a page

At an agent handoff, should the receiving agent get the whole conversation history?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Usually not the raw whole. Pass a structured handoff summary — reason, verified identity, established facts, what was already tried — plus the last few verbatim turns. Full transcripts cost tokens and invite the peer to re-litigate work the previous agent finished.

open as a page

When two worker agents return contradictory briefs, how should the orchestrator synthesize them?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Treat contradiction as a signal to resolve, not noise to average. Compare the returns' evidence and provenance, prefer the claim with a locatable source, and if neither is decisive re-dispatch a narrow worker to settle it or report the disagreement explicitly rather than silently picking one.

open as a page

Why does a routing supervisor's context balloon by turn 15, and how do you fix it?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Every specialist reply flows back through the supervisor, so its window accumulates work it never needed to read. Cost per routing decision climbs, routing quality degrades with length, and the run eventually stalls. Fix it with short return contracts, a task ledger, and externalized artifacts.

open as a page

A debate ensemble costs 10x tokens for 4 points of accuracy — do you ship it?

level: principalimportance: should knowfreq 46%

basics

~20 s

Answer with three checks: is the 4 points real on your traffic and outside the noise band, is an error costly enough to be worth 10x, and would the same budget spent on a stronger model, better grounding or a verifier buy more. Then apply it selectively, not everywhere.

open as a page

When does a decentralized handoff swarm beat a central orchestrator for a workload?

level: principalimportance: should knowfreq 42%

basics

~20 s

Swarms win where one user-facing conversation moves through specialist domains in sequence and each specialist can judge the next owner locally. They lose where routing must be auditable, policy is centrally governed, or the flow needs results combined rather than control passed on.

open as a page

How do you decide how many parallel workers an orchestrator should fan out to?

level: principalimportance: should knowfreq 38%

basics

~20 s

Let the work decide, not the model. Width should follow the number of genuinely independent slices, then be capped by three ceilings: the token budget, which grows roughly linearly with width; the lead's own window, which must hold every return; and provider concurrency limits.

open as a page

When does a supervisor-of-supervisors beat one flat supervisor over all agents?

level: principalimportance: should knowfreq 33%

basics

~20 s

Nest when the roster has grown past what one router can choose from reliably and the specialists cluster into genuine domains with separate ownership or permissions. Nest reluctantly: each extra level adds a model call, a summarization boundary and a harder failure to attribute.

open as a page