Multi-agent LLM systems
Getting several agents to work as one system: who plays which role, how they talk, which topology routes the work, what state they share, and what happens when they disagree or stall. Interviewers ask because multi-agent designs multiply capability and failure modes at the same rate.
on this pageshowhide
guide
overview
~1 minMulti-agent systems questions test whether you can split one job across several model-driven agents and still reason about the whole. Interviewers care less about which framework you have used than about whether you see the bill: every extra agent adds tokens, latency, a boundary where meaning can drift, and one more place for a failure to hide. A good answer justifies the second agent before it designs the tenth, and at the senior end several questions are really asking when a single agent would have done better. The subject splits into six sections. [Agent roles and specialization](/topics/found-multi-agent-systems-agent-roles) covers what makes a role real rather than cosmetic. [Orchestration topologies](/topics/found-multi-agent-systems-orchestration-topologies) is the largest: orchestrator-worker fan-out, supervisor routing, peer handoffs, and debate or critique ensembles, with the question of when each fits. [Inter-agent communication](/topics/found-multi-agent-systems-communication-protocols) is about how work crosses from one agent to another and what happens when a message arrives broken. [Shared state and context](/topics/found-multi-agent-systems-shared-state) covers the durable store agents coordinate through and who may change it. [Coordination and conflict resolution](/topics/found-multi-agent-systems-coordination-conflict) treats a fleet of agents as the distributed system it is: races, stalls, contradictions, restarts. [Multi-agent evaluation](/topics/found-multi-agent-systems-evaluation) asks how to score the whole and find the part to fix. Start with roles, because every topology is a way of arranging them. Topologies come next, then communication and shared state, the two channels every topology uses. Coordination and evaluation come last, since both assume you can already draw the system you are defending.
primer
A few ideas run under every section. Once they are in place, most questions below become a matter of applying them to a specific design. - **An agent is what it can see and what it can do.** Its context window and its tool scope separate it from its neighbours. A new persona laid over identical context and identical tools changes the wording, not the capability. Most of the measurable benefit of splitting work comes from isolation: a worker explores in its own window and hands back something much smaller than what it read. - **Every boundary is a lossy, priced channel.** Moving work between agents costs input tokens on the receiving side and risks the receiver reading the task differently from the sender. A typed contract makes the boundary checkable; free prose leaves it to interpretation. - **Parallelism suits independent reading, not shared writing.** Fan-out works when slices do not overlap and the results can be combined afterwards. When several agents must change the same artifact, you have built a coordination problem, and adding agents makes it worse. - **Errors compound and correlate.** A chain of stages multiplies its failure chances, and downstream agents tend to trust what upstream handed them. Agents built on the same model and prompt also tend to fail in the same way, so agreement between them is weaker evidence than it looks. Redundancy helps only against errors that are independent. - **Durable state lives outside the context window.** Windows are small and disappear with the process; an external store and a log give restartability, inspection and one version of the truth. The design question that follows is who is allowed to write. - **Control needs an owner.** Something has to decide who acts next, when a step has run too long, when a loop must be broken and when the task is finished. Left to the agents themselves, those decisions tend to turn into ping-pong and silent stalls. - **You cannot fix what you cannot attribute.** An end-to-end score tells you whether the system works. Only a complete trace tells you which agent to change.
- Orchestrator
- The lead agent that splits a task into subtasks, dispatches them to worker agents in their own contexts, and combines what they return.
- Supervisor
- A routing agent that, on each turn, picks which specialist should act next or declares the task complete, keeping control centralized.
- Handoff
- A transfer of the live conversation from one agent to a peer, which then continues under its own instructions and tools.
- Fan-out and gather
- Dispatching independent subtasks to several agents in parallel, then collecting and merging their results in one place.
- Tool allowlist
- The explicit set of tools an agent may call. It bounds what a role can actually do, independent of what its prompt says.
- Context isolation
- Running each agent in a separate context window so its exploration, dead ends and raw material do not crowd other agents' windows.
- Handoff contract
- The agreed shape of what passes between two agents: required fields, types and limits, checked at the boundary instead of trusted.
- Artifact store
- Shared durable storage for the files and results agents produce, referenced by path or id rather than copied into every context.
- Append-only event log
- A record of every state change in order, never edited in place, used to rebuild state, audit decisions and resume after a crash.
- Single-writer discipline
- A rule that one owner performs all writes to a given piece of shared state, while other agents may only propose changes.
- Hop budget
- A cap on how many handoffs or routing steps one task may take before the system stops and escalates, used to break loops.
- Correlated errors
- Mistakes that several agents make together because they share a model, prompt or context; voting and debate cannot cancel them.
The sections describe one system from different angles, and they depend on each other in a fairly fixed order. **Roles are the parts; topologies are the wiring.** A topology only makes sense once you know what each node is allowed to see and do. Orchestrator-worker, supervisor routing, peer handoff and debate differ mainly in who holds control and where results come back to. The central designs keep one agent accountable at the cost of that agent's context filling up; the handoff designs spread control out at the cost of auditability. Questions in [orchestration topologies](/topics/found-multi-agent-systems-orchestration-topologies) usually ask you to pick one for a workload and defend how centralized it is. **Communication and shared state are the two channels.** Agents either pass information directly (a chained prompt, a handoff payload, a message on a queue, a call to a remote agent over a cross-party protocol such as A2A) or leave it somewhere others can read. Direct messages are simple and ephemeral; a shared store is durable but raises the question of write authority. Most real systems use both, with small messages carrying references into the store. **Coordination is what the channels need to stay sane.** Once agents run concurrently, the distributed-systems problems arrive: two writers, a worker that never returns, a crash halfway through, two agents that each believe they succeeded. The fixes in [coordination and conflict resolution](/topics/found-multi-agent-systems-coordination-conflict), such as single ownership, deadlines, checkpoints and independent verification, borrow from ordinary distributed systems, adjusted for participants that do not reliably follow a protocol. **Evaluation closes the loop.** [Multi-agent evaluation](/topics/found-multi-agent-systems-evaluation) depends on everything above being observable. If handoffs are typed and every turn is traced, a failed run can be replayed and blamed on one agent; if they are prose and untraced, all you have is a pass rate.
- Agent Roles and Specialization →
Every design is an arrangement of roles, so first learn what makes a role more than a persona: its context, its tools and its contract.
- Orchestrator-Worker →
The most common topology and the clearest case of context isolation paying off; the other topologies are easiest to read as variations on it.
- Inter-Agent Communication →
Every topology moves work across boundaries; learn what each kind of message costs and how a malformed one is caught.
- Shared State and Context →
The durable half of communication, and where questions about write authority, auditing and restartability begin.
- Coordination and Conflict Resolution →
With the channels understood, study how concurrent agents race, stall and contradict each other, and the controls that contain it.
- Multi-Agent Evaluation →
Last, because scoring and blaming a multi-agent run assumes you know every part that could have caused the failure.
Proposing several agents without first saying why one agent with good tools would fall short; interviewers often want that baseline defended.
Treating role prompts as the specialization while every agent shares the same context and tools; the split then adds cost and no capability.
Passing free-text summaries between agents and trusting them, so a missing or misread detail is silently filled in downstream.
Letting several agents write the same shared artifact directly, then trying to fix the resulting conflicts with retries.
Counting agreement between agents on one model and one prompt as independent confirmation; they tend to be wrong together.
Leaving termination to the workers: no deadline, no hop budget, no loop detection, so one stuck agent blocks the whole run.
Routing every specialist's full output back through the lead or supervisor until its own window becomes the bottleneck.
Reporting only an end-to-end pass rate, or comparing a multi-agent system to one agent on unequal token budgets.
The same few decisions come up in almost every design question here, and naming the one you are making usually earns more credit than the diagram. - **One agent versus several.** More agents buy parallel exploration and clean, separate contexts; they cost tokens, latency and debuggability. The honest default is the smallest number that the task's independent parts justify. - **Central control versus local handoff.** A lead or supervisor keeps decisions auditable and policy in one place, but its context becomes a bottleneck. Peer handoffs keep each conversation lean and suit sequential flows, but routing becomes harder to predict and to govern. - **Direct messages versus shared state.** Messages are simple and leave no trace unless you log them; a store survives restarts and can be inspected, but needs rules about who may write. - **Full history versus summary at a boundary.** Passing everything preserves detail and costs tokens and focus; passing a summary is cheap and risks dropping the one fact that mattered. - **Redundancy versus cost.** Debate, voting and independent critics catch some errors at a multiple of the price, and only when the agents can actually fail independently. A deterministic check, where one exists, is usually cheaper.
A handful of shapes recur across the sections under different names; spotting them is the quickest way to place a new question. - **Isolate, then compress.** A worker, a subagent or a private scratchpad does messy work out of sight and releases only a short, checked result. It appears in orchestrator-worker designs, in shared-state promotion rules and in context budgeting. - **Contract at every boundary.** Typed payloads, worker briefs with an explicit scope and output shape, and structured handoff summaries are one idea: make what crosses the boundary checkable. - **One owner per decision.** A single writer for shared state, an orchestrator that owns termination, a supervisor that owns routing: conflicts shrink when exactly one party decides. - **Budget and escalate.** Deadlines on subtasks, hop limits on handoffs, retry caps on failed workers and round limits on debate all bound the work and pass whatever is left over to a person or a fallback path. - **Verify independently.** Fresh-context critics, judges given the evidence rather than the arguments, and checks against the original request all avoid asking an agent to grade itself. - **Record everything, replay later.** Event logs, checkpoints and full traces serve restart, audit and failure attribution with the same data.
explore
- Agent Roles and Specialization4 questions
- Orchestration Topologies19 questions
- Orchestrator-Worker5 questions
- Supervisor and Hierarchical Routing4 questions
- Handoff and Swarm5 questions
- Debate, Critique and Ensembling5 questions
- Inter-Agent Communication5 questions
- Shared State and Context4 questions
- Coordination and Conflict Resolution5 questions
- Multi-Agent Evaluation4 questions
questions
page 2 of 2When two worker agents return contradictory briefs, how should the orchestrator synthesize them?
basics
~20 sTreat contradiction as a signal to resolve, not noise to average. Compare the returns' evidence and provenance, prefer the claim with a locatable source, and if neither is decisive re-dispatch a narrow worker to settle it or report the disagreement explicitly rather than silently picking one.
Why does a routing supervisor's context balloon by turn 15, and how do you fix it?
basics
~20 sEvery specialist reply flows back through the supervisor, so its window accumulates work it never needed to read. Cost per routing decision climbs, routing quality degrades with length, and the run eventually stalls. Fix it with short return contracts, a task ledger, and externalized artifacts.
Your five-persona agent crew performs no better than one agent — how do you diagnose it?
basics
~20 sCheck whether the roles differ in anything except wording. If all five share the same context and the same tools, the personas are decoration — the measurable gains in multi-agent systems come from context isolation and tool scoping, not from character descriptions.
When does adopting an agent protocol beat plain HTTP or a queue between agents?
basics
~20 sAn agent protocol pays when the two ends are owned by different parties: it supplies discovery, a shared task contract and negotiated capabilities across a boundary. Inside one codebase, where both agents deploy together, a typed message on a queue or a direct call is usually the better default.
When is multi-agent arbitration worth its cost before an irreversible action?
basics
~20 sOnly when the action is genuinely irreversible, no deterministic check exists, and the agents can fail independently. Requiring several agents to agree multiplies cost and latency, and buys nothing against correlated errors — agents sharing a model, prompt and context tend to be wrong together.
How would you design a fair evaluation comparing a multi-agent pipeline to one agent?
basics
~20 sHold the token and cost budget equal, use the same task set and the same grading, repeat each task many times because both systems are nondeterministic, and report cost and latency per solved task with variance rather than one pass rate from one run.
A debate ensemble costs 10x tokens for 4 points of accuracy — do you ship it?
basics
~20 sAnswer with three checks: is the 4 points real on your traffic and outside the noise band, is an error costly enough to be worth 10x, and would the same budget spent on a stronger model, better grounding or a verifier buy more. Then apply it selectively, not everywhere.
When does a decentralized handoff swarm beat a central orchestrator for a workload?
basics
~20 sSwarms win where one user-facing conversation moves through specialist domains in sequence and each specialist can judge the next owner locally. They lose where routing must be auditable, policy is centrally governed, or the flow needs results combined rather than control passed on.
How do you decide how many parallel workers an orchestrator should fan out to?
basics
~20 sLet the work decide, not the model. Width should follow the number of genuinely independent slices, then be capped by three ceilings: the token budget, which grows roughly linearly with width; the lead's own window, which must hold every return; and provider concurrency limits.
When does a supervisor-of-supervisors beat one flat supervisor over all agents?
basics
~20 sNest when the roster has grown past what one router can choose from reliably and the specialists cluster into genuine domains with separate ownership or permissions. Nest reluctantly: each extra level adds a model call, a summarization boundary and a harder failure to attribute.
showing 31–41 of 41