skip to content

How do you decide how many parallel workers an orchestrator should fan out to?

level: principalimportance: should knowfreq 38%

answer

  1. width follows independent slices
  2. cost scales linearly, capability does not
  3. the lead's window is the real ceiling
  4. compare at equal token budget, not equal agent count
  5. leads over-fan unless capped

basics

~20 s

Let the work decide, not the model. Width should follow the number of genuinely independent slices, then be capped by three ceilings: the token budget, which grows roughly linearly with width; the lead's own window, which must hold every return; and provider concurrency limits.

solid answer

~50 s

Fan-out width is a **budget decision the lead should not be left to make freely**. Start from the decomposition: the natural width is the number of independent, non-overlapping slices the task actually has — eight transcripts, eight workers. Then apply three ceilings. **Cost**: each worker pays its own instructions, tool definitions and reasoning, so a research-style fan-out commonly runs several times the tokens of one agent reading the same material, and the multiplier scales with width. **Lead capacity**: every return lands in the lead's window, so beyond a few dozen returns you need a two-stage merge or the lead's synthesis degrades. **Concurrency**: provider rate limits and your own throughput controls bound how many workers actually run in parallel, so nominal width beyond that buys cost without buying latency. In practice teams set an explicit width cap per task class and have the lead justify exceeding it, because leads systematically over-fan when told to be thorough.

go deeper

for a junior

Know that each parallel worker costs a full share of tokens and that more workers is not automatically better. Recognising the cost multiplier is the bar here.

for a middle

Explain that width should match the number of independent, comparably sized slices, and that a fan-out's latency is set by its slowest worker, not its average one.

for a senior

Show the three ceilings — token budget, lead window capacity, provider concurrency — and be ready to argue for tightening worker contracts or adding a second round instead of widening the first.

for a principal

Own the economics: define the equal-budget comparison, set width caps and effort scales per task class, enforce them in the harness rather than the prompt, and be honest that some workloads do not justify multi-agent orchestration at all.

## Width is an economic decision The seductive property of a fan-out is that adding workers looks like adding capability. It is really adding cost linearly and capability sub-linearly. Every extra worker pays a fixed overhead — its instructions, its tool schemas, its own reasoning about a task it has just been told about — before it produces a single useful token. Interviewers ask this question because getting it wrong is one of the most common ways an agent system becomes too expensive to ship. ## Start from the work, not from a number The defensible width is the number of **genuinely independent slices**. Independence has a test: could this worker complete correctly without knowing any other worker's output? If not, it is a sequential dependency wearing a parallel costume, and widening the fan-out just produces workers that guess at each other's results. Slices should also be **comparably sized**. A fan-out's latency is the slowest worker's, so eight even slices finish far sooner than one huge slice and seven trivial ones. Uneven slicing is why a nominally parallel run often shows almost no latency benefit. ## Three ceilings **Token budget.** Cost scales close to linearly in width, and orchestrated research patterns burn substantially more than a single-agent equivalent — on the order of a small multiple for a modest fan-out, and roughly an order of magnitude over ordinary chat use for full research runs. There is a real result to hold in mind here: as of mid-2026, several evaluations found single-agent systems matching or beating multi-agent ones *at equal token budget*, which means the honest comparison is never "one agent versus eight" but "one agent with 8x the budget versus eight workers with 1x each". Width is only justified where the work is parallel and the input genuinely exceeds one window. **Lead capacity.** Every return enters the lead's context. Twenty workers at 1,500 tokens each is 30,000 tokens of briefs before the lead reasons at all, and synthesis quality degrades as that window lengthens. Past that point you need a hierarchical merge — condense batches, then synthesize the condensations — which costs another layer of information loss. That trade, not raw cost, is usually the real ceiling on width. **Concurrency.** Provider rate limits, your own bounded-concurrency controls and downstream tool quotas determine how many workers genuinely run at once. If only six can run concurrently, a width of twenty runs in waves and buys none of the latency benefit that motivated the width. ## Governing the lead Leads over-fan. Told to be thorough, a lead will spawn a worker per question it can think of. Practical governance: give the lead an explicit width cap per task class, give it a rough effort scale (a simple lookup gets one or two workers, a comparative study gets one per source, a deep survey gets ten to fifteen), and require it to state why a wider fan-out is needed. Enforce with a hard cap in the harness rather than relying on the prompt, and track spend per run so regressions show up. ## Prefer depth of contract over width of fan-out When results are poor, the instinct is to add workers. Usually the better lever is a tighter worker contract, better slices, or a verification pass on the existing returns. A second round of three sharp workers, dispatched after reading the first round, routinely beats one round of fifteen vague ones — and costs less. ## Where the ceiling moves Width is justified higher when slices are large and read-heavy (each worker's window is doing real work no single window could do), when latency matters more than cost (user-facing research where wall-clock time is the product), and when returns are small and highly structured so the lead's window stays light. It should be pulled lower when slices are small (the per-worker overhead dominates the useful work), when subtasks are interdependent, when workers perform side effects rather than only reading, and when the answer must be auditable — every hop compresses evidence, and wide fan-outs put more compression between the source and the final claim.

  • Why is "one agent versus eight workers" the wrong comparison?
    Because it holds agent count fixed instead of spend. The honest baseline is a single agent given the same total token budget the fan-out would consume. Several mid-2026 evaluations found single-agent systems matching or beating multi-agent ones at equal budget, which narrows the justification for width to cases where the work is genuinely parallel and the input exceeds any single context window.
  • What breaks first as you widen a fan-out — cost or quality?
    Usually quality, via the lead. Cost rises predictably and linearly, but the lead's synthesis degrades once the returns fill a large fraction of its window, and that degradation is silent. The mitigation is a hierarchical merge that condenses batches before the lead reads them, which buys width at the price of another lossy compression hop.
  • Results are poor. Is adding workers the right response?
    Rarely. Poor returns usually mean under-specified worker contracts, badly chosen slices or missing verification, and none of those improve with more of the same. A second round of three sharply scoped workers, dispatched after reading round one, typically beats a single round of fifteen vague ones and costs less.
  • How do you stop a lead from over-fanning in production?
    Enforce a width cap in the harness rather than in the prompt, give the lead an effort scale keyed to task class, require a stated justification for exceeding the default, and track token spend per run so regressions surface. A prompt asking the lead to be economical is not a control; a hard cap is.

saying these in an interview costs you the question

  • Treating more workers as automatically more thorough
  • Ignoring that every return lands in the lead's window
  • Comparing to one agent without equalising the token budget
  • Fanning out over interdependent subtasks
  • Setting width above the concurrency the provider allows

context