skip to content

Eight parallel research workers return overlapping, uneven briefs — what do you fix first?

level: seniorimportance: must knowfreq 62%

answer

  1. it is a specification failure, not a model failure
  2. objective, scope, tools, output shape
  3. disjoint slices are the lead's job
  4. cap the length, fix the fields
  5. tool scoping constrains more than persona text

basics

~20 s

Fix the worker contract before anything else. Each worker needs an explicit objective, an exclusive slice of the sources, an allowed tool set, and a required output shape with a length cap. Overlap and unevenness are almost always a lead that copied the goal instead of partitioning it.

solid answer

~50 s

Overlapping, uneven briefs are a **specification failure at the lead**, not a model failure at the workers. The repair is to make each dispatch a contract with four parts: a **single objective** stated in the worker's own terms; an **exclusive scope** — this one transcript, this directory, this date range, and explicitly not the neighbours'; an **allowed tool set** small enough that the worker cannot wander; and an **output shape** — for example 200 words or fewer, two verbatim quotes with locations, a confidence marker, and no follow-up questions. Add a negative clause for the common drift ("do not summarise the wider policy debate"). This is the highest-leverage knob in the whole topology, because the lead's picture of the world is built entirely from these returns: an under-specified worker is free to reinterpret the goal, and eight workers reinterpreting it independently produce eight incomparable documents.

go deeper

for a junior

Know that each worker is given its own instructions, and that vague instructions produce inconsistent results. Being able to name the parts of a good worker prompt is enough here.

for a middle

Explain the four parts of a dispatch contract and why an uncapped return quietly undoes the benefit of the fan-out. Be able to read a symptom — overlap, uneven length — back to the missing clause.

for a senior

Demonstrate that you would fix the dispatch rather than the synthesis prompt, partition scopes at dispatch time, restrict tools per worker, and calibrate expected effort so cost and latency stay bounded.

for a principal

Own the framing that most multi-agent failures are specification failures the orchestration layer is responsible for, and be ready to say what you would standardise across every worker in a fleet versus what stays per-task.

## Why the contract is the main quality lever In an orchestrator-worker system the lead cannot see what the workers saw. Everything it knows arrives through the returns. That makes the dispatch prompt — the worker's objective, scope, tools and output shape — the single place where quality is decided. Tuning the lead's synthesis prompt afterwards is downstream damage control on inputs that were already incomparable. The symptom is diagnostic. Overlap means the slices were not disjoint. Unevenness — one brief three paragraphs, another three pages, one with citations, another without — means the output shape was left to the model's discretion, and eight independent samples of "be helpful" diverge. ## The four parts of a worker contract **Objective.** One sentence, in the worker's frame, not the user's. "Extract every short-term-rental vote from this transcript" — not "help answer the user's question about rentals". A worker that has to infer the goal from a copy of the user's request will re-do the lead's job badly, in miniature. **Scope.** Name exactly what this worker owns and, where sibling workers exist, what it must not touch. Disjointness is the lead's responsibility: if it hands two workers the same corpus with different phrasings, they will both read the whole thing and return two overlapping digests at double the cost. Explicit boundaries are cheaper than de-duplication at fan-in. **Tools.** Give the minimum set the subtask needs. Tool scoping constrains behaviour at least as strongly as the persona text does: a worker without a broad web-search tool cannot drift off its assigned transcript, whatever the prompt says. It also cuts the token cost of the tool definitions each worker pays. **Output shape.** Length cap, required fields, forbidden moves. A good return is structured — claim, evidence with a locator, confidence — because the lead will have to compare eight of them and reconcile disagreements. Free prose does not compare. The cap matters as much as the fields: uncapped returns re-fill the lead's window and turn a fan-out into a slow, expensive single agent. ## Effort calibration belongs in the contract too A subtle version of unevenness is effort mismatch: one worker makes two tool calls, another makes forty on a comparable slice. The lead should state the expected size of the job — roughly how many searches, how deep to go, when to stop and report what it has. Without that, workers calibrate effort from the task text alone and the variance is large. Telling a worker to return partial findings with a gap flagged, rather than grinding, also keeps the fan-out's latency bounded by the slowest sensible worker rather than the most stubborn one. ## What good looks like A dispatch for the newsroom case might read: *you are reading the March 4 Oakland council transcript only; extract every agenda item and vote touching short-term rentals; you may use the transcript-read and quote-locate tools; return 200 words or fewer, including at most two verbatim quotes with line numbers and a one-word confidence; do not discuss other cities; do not ask follow-up questions; if the transcript is truncated, say so and return what you have.* Every clause in that closes a failure mode someone has actually hit. ## Where contracts do not save you Contracts fix specification and comparability. They do not fix genuinely interdependent work — if worker B needs worker A's answer, no amount of prompt precision makes a parallel fan-out correct, and the dependency belongs in an earlier sequential round. They also do not fix a bad decomposition: perfectly specified workers on the wrong slices produce tidy, uniform, useless briefs. And they cannot make a worker's claims true; verification is a separate step. Failure-attribution studies of multi-agent systems put a large share of observed failures in the specification-and-system-design cluster rather than in the models themselves, which is the empirical version of the same point: most of what looks like agent stupidity in a fan-out is an instruction the lead never wrote.

  • How would you stop two workers reading the same source without the lead diffing their returns afterwards?
    Partition at dispatch. The lead assigns each worker an exclusive key — a document id, a directory, a date range — and states the exclusion in the prompt. Deduplicating at fan-in is strictly worse: you have already paid twice for the reading and you now have to decide which of two divergent digests of the same text to believe.
  • Should the worker prompt specify how much effort to spend?
    Yes. Effort variance is a real failure mode: comparable slices can draw two tool calls from one worker and forty from another. State the expected shape of the job — roughly how many searches, how deep, when to stop — and instruct the worker to return partial findings with a gap flagged rather than grinding. That bounds both cost and the fan-in's latency.
  • Does specialising workers by persona help, or is it theatre?
    Mostly it is the tool allowlist doing the work, not the persona text. A "security reviewer" worker with the same tools and the same output shape as a generic one behaves similarly; the same worker restricted to reading diffs and running the test suite behaves very differently. Treat a role as persona plus tool scope plus output contract, and expect the last two to carry the weight.

saying these in an interview costs you the question

  • Copying the user's request verbatim into every worker prompt
  • Leaving output length and format to the worker's discretion
  • Fixing overlap by deduplicating at fan-in instead of at dispatch
  • Giving every worker the full tool set for convenience
  • Blaming the model when the dispatch never stated the scope

context