In LangGraph, how does a supervisor route to worker agents, and what does each hop cost?
answer
- star topology, one router node
- LLM decides, Command carries the choice
- workers edge back to the centre
- two model calls per unit of work
- recursion limit is the crash barrier
basics
~20 sA supervisor is an LLM node that reads shared state and returns a Command naming which worker node runs next; each worker hands control back to the supervisor. Every delegation therefore costs an extra model call, on a message list that keeps growing.
solid answer
~50 sYou build it as an ordinary `StateGraph`: one `supervisor` node plus one node per worker, an edge from `START` to the supervisor, and an edge from each worker back to the supervisor. The supervisor node calls a model with structured output over the worker names and returns `Command(goto=chosen)`, or `Command(goto=END)` when the task is done; each worker is usually itself a compiled agent graph attached with `add_node`. The cost is real: N delegations mean roughly 2N model calls, all serialized, and because the supervisor re-reads the shared history its prompt grows with every worker turn. A supervisor that never chooses `END` loops until the run hits LangGraph's recursion limit (default 25 supersteps) and raises `GraphRecursionError`. Mitigations: a cheap model for the supervisor, a hop counter in state, and workers that write summaries back rather than raw tool chatter.
code
python · 32 linesfrom typing import Literal
from langgraph.graph import END, START, MessagesState, StateGraph
from langgraph.types import Command
MEMBERS = ["researcher", "writer"]
def supervisor(state: MessagesState) -> Command[Literal["researcher", "writer", "__end__"]]:
last = state["messages"][-1].content
if "final draft" in last:
return Command(goto=END)
goto = "writer" if "researched" in last else "researcher"
return Command(goto=goto)
def researcher(state: MessagesState) -> dict:
return {"messages": [{"role": "assistant", "content": "researched the topic"}]}
def writer(state: MessagesState) -> dict:
return {"messages": [{"role": "assistant", "content": "final draft"}]}
builder = StateGraph(MessagesState)
builder.add_node("supervisor", supervisor)
builder.add_node("researcher", researcher)
builder.add_node("writer", writer)
builder.add_edge(START, "supervisor")
for member in MEMBERS:
builder.add_edge(member, "supervisor")
graph = builder.compile()go deeper
Be able to draw the shape: one router node in the middle, worker nodes around it, control returning to the middle after each worker. Say that the router is itself an LLM call.
Explain how the router picks — a model call with structured output over the worker names, returned as a Command goto — and that workers can be compiled agent graphs attached as nodes.
Quantify the cost: two model calls per unit of work, serialized latency, a prompt that grows every hop, and a recursion limit that turns a non-terminating router into a crash. Bring concrete mitigations.
Own the decision of whether the routing step deserves its own model call at all, and where hierarchy starts paying for itself. Be ready to argue for direct worker-to-worker handoff on the transitions that never vary.
## The shape Supervisor is not a LangGraph primitive — it is a topology you build from the primitives, which is why interviewers like it: you cannot recite an API, you have to describe a design. The canonical shape is a star. One node named `supervisor` and one node per worker (`researcher`, `coder`, `writer`). `START` points at the supervisor. Each worker has an edge back to the supervisor. The supervisor's own outgoing transition is dynamic: it returns `Command(goto="researcher")`, or `Command(goto=END)`. Because it routes with a Command rather than declared edges, annotate it `-> Command[Literal["researcher", "coder", "writer", "__end__"]]` so the star is still visible when you draw the graph. Workers are usually compiled graphs, not plain functions. A prebuilt agent (`create_react_agent` in `langgraph.prebuilt`) returns a compiled graph, and a compiled graph is accepted by `add_node` directly as long as it shares the parent's state keys — with a `MessagesState`-shaped parent, it does. ## Why the supervisor exists at all The supervisor turns routing into an explicit, inspectable reasoning step. A single agent with fifteen tools makes its routing decision implicitly inside tool selection, where you cannot see it, cannot give it its own prompt, and cannot swap its model. Pulling it out means you can give the supervisor a small, cheap, strongly-constrained model with structured output over a closed set of worker names, and log every routing decision as a first-class event. It also gives each worker its own system prompt and its own tool subset. That is often the actual win: the coder never sees the search tools, so it cannot pick one by mistake. ## What each hop costs Be precise about this in an interview, because the naive answer ("a bit more latency") is wrong by a factor. 1. **Two model calls per unit of work.** One supervisor call to choose, one (or more) worker call to act. A five-step task that a single agent does in five calls costs ten or more here. 2. **Serialized latency.** The star topology is strictly sequential: worker, supervisor, worker, supervisor. Nothing overlaps unless you deliberately fan out. 3. **Prompt growth.** With one shared message channel, the supervisor re-reads the entire accumulated history on every hop, and so does each worker. Token spend grows roughly with the square of the number of hops, and the growth is invisible until the bill arrives or a context window overflows mid-run. 4. **Checkpoint size.** If the run is persisted, every superstep writes a snapshot of that growing state. ## The termination problem The supervisor decides when to stop, and models are bad at deciding they are finished. The classic failure is a supervisor that keeps handing the task back to a worker that keeps producing near-identical output. LangGraph bounds this structurally: a run has a recursion limit — 25 supersteps by default, overridable per invocation via the config — and exceeding it raises `GraphRecursionError` rather than burning budget forever. Treat that as a crash barrier, not as your termination logic. Real termination logic is an explicit `END` in the supervisor's decision set, plus a hop counter in state that the supervisor prompt is told about, plus a task-completion criterion that is checkable rather than vibes-based. ## Hierarchical supervisors When the worker list grows past what one model routes reliably, the same shape nests: a top supervisor whose "workers" are themselves compiled supervisor graphs, each with its own team. Every level adds a model call per hop, so a three-level hierarchy costs three routing calls before any work happens. That is why hierarchy is a response to a routing-accuracy problem, not a way to organise code — if the top model can pick correctly among ten workers, a flat star is cheaper and easier to debug. ## Making it cheaper - Give the supervisor a small model with structured output constrained to the worker names; it is a classification task, not a reasoning task. - Have workers return a summarized result into the shared channel and keep their intermediate tool traffic in their own subgraph state, so the supervisor's prompt grows by a paragraph per hop instead of a transcript. - Consider handoffs directly between workers (a worker returning `Command(goto="writer", graph=Command.PARENT)`) for the transitions that are always the same — that removes a supervisor call per hop, at the cost of decentralizing the routing logic. - Instrument hops per task. If the median run takes two hops, the supervisor is overhead; if it takes nine, your termination criterion is broken.
- What stops a supervisor loop that never selects END?LangGraph counts supersteps against a recursion limit — 25 by default, settable per invocation in the run config — and raises `GraphRecursionError` when it is exceeded. That is a crash barrier, not termination logic. Real designs also carry a hop counter in state that the supervisor prompt can see, and a checkable completion criterion, so the run ends deliberately rather than by hitting the ceiling.
- How would you cut the token cost of a long supervisor run without changing the topology?Stop letting raw worker traffic into the shared channel. Give each worker its own state and write back a short summary, so the supervisor's prompt grows by a paragraph per hop rather than a transcript. Then downgrade the supervisor to a small model with structured output over the worker names — routing is classification, and the cheap model usually matches the expensive one on it.
- When is a hierarchical supervisor worth the extra routing call?When routing accuracy, not code organisation, is the bottleneck — typically once one router must choose among more workers than the model picks reliably, or when teams have genuinely different state shapes. Each level adds a model call before any work happens, so a flat star stays cheaper and far easier to debug whenever the top model can still route correctly.
saying these in an interview costs you the question
- Thinks supervisor is a built-in LangGraph node type
- Ignores that each delegation adds a whole extra model call
- Assumes the model reliably decides on its own when to stop
- Treats the recursion limit as the termination strategy
- Says workers routing back to the supervisor needs conditional logic