When does a supervisor-of-supervisors beat one flat supervisor over all agents?
answer
- depth is debt, not tidiness
- each router should face a small menu
- clusters must be real, not org-chart shaped
- every level adds a lossy summary
- failure attribution gets harder per level
basics
~20 sNest when the roster has grown past what one router can choose from reliably and the specialists cluster into genuine domains with separate ownership or permissions. Nest reluctantly: each extra level adds a model call, a summarization boundary and a harder failure to attribute.
solid answer
~50 sA flat supervisor works well over a handful of specialists and degrades as the roster grows - descriptions overlap, the routing prompt bloats, and selection accuracy falls. Hierarchy fixes that by giving each router a small menu: a city-hall lead routes to a Services supervisor over four agents and a Compliance supervisor over three, and each sub-supervisor owns its own vocabulary, completion criteria and permission scope. The gains are real when the clusters are real - when the domains are separable, owned by different teams, or need different authority. The costs are equally real: every hop is another model call and another summary, so latency and token spend multiply, information is lost at each boundary, and when the run fails you must first work out *which level* failed. My rule of thumb is to sharpen descriptions and prune the roster first, and only nest when a genuine domain boundary - not just headcount - already exists.
go deeper
Know that supervisors can be nested - a top supervisor routing to team supervisors, each with its own specialists - and that this keeps any single routing choice small.
Explain why a flat roster degrades as it grows and how grouping into teams restores a small menu per decision, plus the extra model call each level adds on the critical path.
Argue the tradeoff with operational detail: latency and token multipliers per hop, information loss at each summary boundary, per-level tracing, and making sub-supervisors return control instead of finishing.
Own the decision. Tie depth to real domain, ownership and permission boundaries rather than headcount or org shape, name the metrics that justify or retire a level, and be willing to collapse a hierarchy that did not move routing accuracy.
## The flat supervisor's ceiling One router over four specialists is a solved problem. One router over twenty is not. Three things degrade together as the roster grows. **Selection accuracy falls.** More options means more near-neighbours, and near-neighbours are where routers guess. Twenty specialists in one municipal system will contain several pairs whose scopes overlap at the edges, and the router will oscillate between them. **The routing prompt bloats.** Every specialist's scope line, exclusions and examples sit in the prompt for every decision, so the fixed cost of routing rises with roster size regardless of the request. **Ownership blurs.** Twenty descriptions maintained by one team drift; twenty maintained by five teams drift faster, because nobody owns the disambiguation between two teams' agents. ## What a second level buys Hierarchy - a supervisor whose "specialists" are themselves supervisors - addresses all three by making every routing decision small again. A worked shape: a city-hall lead sits over two team supervisors. The **Services** supervisor routes among potholes, sanitation, streetlights and parks. The **Compliance** supervisor routes among permits, inspections and code enforcement. The lead makes one coarse decision - service delivery or compliance - and each team supervisor makes a fine one among three or four options. Nobody ever chooses from seven. Beyond menu size, three things become cleanly expressible: - **Vocabulary and criteria per team.** The Compliance supervisor can carry a long prompt about statutory language and evidentiary standards that would be dead weight in the Services router. - **Authority boundaries.** Teams can hold different tool permissions and data scopes. An agent's effective permission is the intersection of the user's rights and what that branch is allowed to do, and a level boundary is a natural place to enforce the narrowing. - **Organizational ownership.** Each team supervisor and its roster can be owned, versioned and evaluated by the team that understands that domain, with only the coarse routing contract shared. ## What it costs **Latency and tokens multiply.** A request now traverses lead, team supervisor, specialist, and back up through both. Each hop is a model call on the critical path, and multi-agent orchestration already runs at a large multiple of single-agent token cost. **Information is lost at every boundary.** Each level summarizes for the one above. Two levels means two lossy compressions between the work and the decision-maker, and nuance the lead needed can vanish in a sub-supervisor's summary. **Attribution gets hard.** When the run produces a wrong answer, the question is no longer "which agent failed" but "which level failed" - did the lead route to the wrong team, did the team supervisor pick the wrong specialist, or did the specialist do the work badly? Published failure taxonomies put a large share of multi-agent failures in specification and system design rather than in the individual agents, and depth is exactly where specification errors hide. Budget for per-level tracing before you add the level. **Termination gets ambiguous.** Each supervisor believes it owns finishing. Sub-supervisors must return control rather than declaring the whole task complete, and that has to be written into their completion criteria explicitly or you get runs that end early because a team decided its part was the whole job. ## When not to nest Do the cheap things first. Rewrite overlapping descriptions with explicit exclusions. Delete specialists nobody routes to - measure it. Merge two agents whose scopes differ by a paragraph. A flat roster of eight sharp, non-overlapping specialists routes better than a two-level tree over twelve mushy ones. Also resist nesting for *organizational symmetry*. Mirroring your org chart into your agent topology is seductive and produces depth that serves reporting lines rather than requests. The clusters that justify a level are the ones a user's request actually respects - if most requests cross your proposed team boundary, the boundary is wrong. ## The judgment an interviewer is listening for That you treat depth as debt. You add a level when a measured routing-accuracy problem coincides with a real domain, ownership or permission boundary; you pay for it in latency, tokens and attribution difficulty; and you say what you would measure - first-choice routing accuracy per level, hops per request, cost per resolved request - to know whether the level earned its keep. Anyone who answers "more structure is better organized" has not run one of these in production.
- A team proposes three levels of supervision for twelve agents. What is your challenge?Twelve agents rarely need more than one level. I would ask for the routing-accuracy numbers on the current flat roster, check how many specialists are actually being routed to, and look for overlapping descriptions first. Three levels over twelve agents usually mirrors an org chart, and it buys two extra lossy summaries and two extra model calls per request in exchange for tidiness.
- How do you stop a sub-supervisor from ending the whole run when its own part is done?Write its completion criteria as returning control, not finishing the task. A team supervisor's terminal action should be handing a structured result back to its parent, and only the top-level supervisor holds a true finish signal. Enforce it in the orchestration layer too, so a sub-supervisor's finish value simply pops one level rather than terminating the session.
- What would you measure to decide whether a second level actually earned its place?First-choice routing accuracy at each level against labelled requests, hops and wall-clock latency per request, tokens and cost per resolved request, and the rate of cross-boundary requests that bounce between teams. Compare against the flat baseline you replaced. If accuracy did not move and cost did, collapse it back.
saying these in an interview costs you the question
- Adds levels to mirror the org chart rather than the requests
- Assumes deeper hierarchies are inherently better organized
- Ignores that each level adds a lossy summarization boundary
- Lets a sub-supervisor declare the whole task finished
- Cannot name a metric that would justify the extra level