skip to content

When do multiple LangGraph agents beat one agent with more tools?

level: principalimportance: should knowfreq 42%

answer

  1. one agent is the null hypothesis
  2. split on measured tool-selection failure
  3. different model, different blast radius
  4. every hop is a model call
  5. topology is cheap, schema is not

basics

~20 s

Split when a single prompt can no longer carry the job: too many tools to select from reliably, genuinely different models or permissions per role, or work that must run in parallel. Otherwise one agent with a good tool set is cheaper, faster and far easier to debug.

solid answer

~50 s

Default to one agent. In LangGraph the cost of splitting is concrete: every delegation adds a model call, transitions are serialized, context is lost or duplicated at each handoff, and evaluation now has to attribute failures across agents. Splitting pays when one of a few things is true — tool selection accuracy is degrading because the tool list outgrew what the model handles well; roles need different models, temperatures or credentials, so a coding agent runs a different model than a cheap classifier; work is genuinely parallel and you want branches running concurrently; or an agent needs a narrower permission scope than its neighbours, which is a security boundary a shared prompt cannot enforce. The reassuring part is that in LangGraph agents are just nodes, so promoting a node to a subgraph later is a refactor of node boundaries. The commitment you make early is the state schema, not the topology — design that to survive the split.

go deeper

for a junior

Know that more agents is not automatically better, and that each handoff means another model call. Say that one agent with the right tools is the starting point.

for a middle

Name the concrete triggers — too many tools to select from, different models per role, work that can run in parallel — and the costs on the other side.

for a senior

Bring measurement into it: what you track before splitting, what filtering you try first, and how you attribute failures once several agents share a run.

for a principal

Own the reversal criterion and the durable commitment. Argue that topology follows the evidence while the state schema and the trust boundaries are the decisions you must get right early.

## Start from the null hypothesis The honest default is one agent with a well-chosen tool set. It makes one model call per step, keeps all context in one place, produces one transcript to read, and has one failure to explain. Multi-agent designs are frequently adopted because they are architecturally satisfying rather than because a measured problem demanded them, and a principal-level answer says so out loud before listing the cases where splitting wins. ## The cases that genuinely justify a split **Tool-selection accuracy has degraded.** There is a point where a single model choosing among a long tool list starts picking wrong or hallucinating arguments. The number depends on the model and on how distinguishable the tools are, so the trigger is measured, not memorized: track wrong-tool rate as the list grows. The fix is to partition the list so each agent sees only tools relevant to its role. Note the cheaper alternative first — dynamic tool filtering by task, which reduces the visible list without adding a hop. **Roles need different model configuration.** A router wants a small, cheap, deterministic model with structured output. A code writer wants a strong model with a long context. A summarizer wants a cheap one with a tight token budget. One agent means one model per turn, so heterogeneous requirements are a real argument for separate nodes — and in LangGraph this is nearly free to express, since each node binds its own model. **Different blast radius.** An agent that can write to production systems should not share a prompt, and therefore a decision surface, with one that browses arbitrary web content. Separating them puts a structural boundary between untrusted input and privileged tools rather than a sentence of instructions. This is the argument least often made and hardest to refute. **Real parallelism.** If five independent lookups can run at once, a graph that fans out to concurrent branches wins on wall-clock time. One agent is inherently sequential across turns. **Distinct evaluation.** When you need to say "retrieval is at 82%, synthesis is at 91%", the components have to be separately invocable. Splitting for measurability is legitimate — just be honest that it is an observability decision, not a capability one. ## The costs to put on the other side - **Model calls.** A router plus workers roughly doubles the calls for the same work. At scale that is the budget line item, not a rounding error. - **Latency.** Handoffs are serialized supersteps. Every hop is a full round trip, and users feel it. - **Context loss at boundaries.** Either you share the transcript, and prompts grow with every hop, or you summarize, and detail is dropped. There is no third option that costs nothing. - **Debuggability.** Failure attribution across a graph of agents is materially harder than reading one transcript, and it gets worse when workers have private state. - **Prompt surface.** Each agent has a prompt, and each prompt drifts independently. Five agents is five things to regression-test after a model upgrade. ## A workable decision order 1. One agent, all tools. Measure wrong-tool rate, latency, cost per task, and end-quality. 2. If tool selection is the problem, filter the tool list per task before splitting anything. 3. If prompts have become contradictory — the same instructions trying to serve two jobs — split by role, because that is a genuine prompt-capacity limit. 4. If the split is happening, choose the topology by how variable routing is: fixed sequences want plain edges; variable delegation wants a router node; well-known transitions want direct handoffs between workers. 5. Re-measure. If the multi-agent version is not better on the metric that motivated the split, revert — this is the step teams skip. ## What LangGraph specifically changes about the decision Because every agent is a node and every compiled graph is usable as a node, the topology is late-binding: a function node can be promoted to a subgraph without rewriting the parent. What is *not* cheap to change is the state schema. Channels, reducers and what counts as public versus private are the interface everything else is written against, and reworking them touches every node and every persisted checkpoint. So the principal-level advice is to keep the state schema clean and explicit from day one, and to stay relaxed about the agent count — the topology can follow the evidence, the schema mostly cannot.

  • What would you measure before deciding to split one agent into several?
    Wrong-tool-selection rate, end-to-end task success, cost per completed task, and p95 latency, all on a fixed evaluation set. Splitting is justified when one of those degrades as tools or responsibilities pile up. Afterwards, re-run the same set on the multi-agent version — if the motivating metric did not improve, the split added cost and nothing else.
  • Is there a cheaper intervention than splitting when the tool list gets too long?
    Usually yes: filter the tools exposed for a given task or phase, so the model sees five relevant tools instead of thirty. That recovers most of the selection accuracy without adding a routing model call or a handoff, and it keeps the single readable transcript. Split only when filtering is not enough or the roles differ in model or permissions too.
  • Which early decision is hardest to reverse in a LangGraph multi-agent system?
    The state schema. Channels, their reducers, and the public-versus-private split are what every node and every persisted checkpoint is written against, so changing them ripples everywhere. Topology is comparatively cheap — a function node can be promoted to a subgraph, or a router removed in favour of direct handoffs, without rewriting the agents themselves.

saying these in an interview costs you the question

  • Treats multi-agent as automatically more capable than one agent
  • Cannot name a metric that would justify the split
  • Ignores that each handoff costs a model call and a round trip
  • Splits by team org chart rather than by measured failure
  • Assumes the state schema is as easy to change as the topology

context