skip to content

When two worker agents return contradictory briefs, how should the orchestrator synthesize them?

level: seniorimportance: should knowfreq 48%

answer

  1. synthesis is not concatenation
  2. check for fake conflicts first
  3. evidence beats stated confidence
  4. agreement is not independent confirmation
  5. resolve, verify, or disclose — never average

basics

~20 s

Treat contradiction as a signal to resolve, not noise to average. Compare the returns' evidence and provenance, prefer the claim with a locatable source, and if neither is decisive re-dispatch a narrow worker to settle it or report the disagreement explicitly rather than silently picking one.

solid answer

~50 s

Synthesis is not concatenation. When briefs disagree, the lead should first check whether the disagreement is **real** or an artefact — different scopes, different time windows, or two names for the same thing produce fake contradictions constantly. If it is real, adjudicate on **evidence, not confidence**: the return carrying a verbatim quote and a locator beats the return asserting a conclusion, however assertively phrased. When neither return can settle it, the lead has two honest moves: **re-dispatch a narrow tie-break worker** against the specific sources, which costs one cheap call, or **surface the disagreement** in the final answer with both claims and their sources. What it must not do is average the two into a hedge, or take the more fluent brief because it reads better. Requiring provenance and a confidence marker in every return is what makes any of this possible at fan-in.

go deeper

for a junior

Know that the lead has to reconcile what the workers send back, and that simply pasting the briefs together is not synthesis. Recognising that conflicts must be handled at all is the bar here.

for a middle

Explain the difference between a real contradiction and a scope, vocabulary or recency artefact, and why returns need source identifiers and dates for the lead to tell them apart.

for a senior

Show the adjudication policy: rank by locatable evidence, re-dispatch a narrow verifier for high-stakes conflicts, disclose unresolved disagreement, and never average. Mention the citation chain that makes a bad answer debuggable.

for a principal

Own the information-loss argument — every hop compresses, and a two-stage merge trades more loss for a lead that can still reason. Be ready to say which classes of disagreement your product should surface to users rather than resolve.

## Why fan-in is the hard half Fanning out is mechanical. Fanning in is where the pattern's quality is actually decided, because the lead is now reasoning over second-hand evidence it cannot re-read cheaply. Eight briefs about eight city councils will contain overlaps, gaps and outright conflicts, and the lead's default behaviour — write a smooth paragraph that accommodates everything — is exactly the failure mode. ## Step one: is the contradiction real? A large share of apparent conflicts are artefacts of the fan-out itself: - **Scope mismatch.** One worker read the March minutes, another the April ones; both are right about different states of the world. - **Vocabulary mismatch.** Two workers name the same entity, policy or metric differently, or use the same word for different things. - **Granularity mismatch.** One reports a committee recommendation, another the full-council vote. - **Recency.** One source is stale and the other current, and neither return dated its evidence. All four are cured at dispatch by requiring the return to carry scope, source identifier and date. If your synthesis step is spending most of its effort disambiguating, the contract upstream is under-specified. ## Step two: adjudicate on evidence For genuine conflicts, rank by what can be checked: 1. A claim with a verbatim quote and a locator the lead can re-fetch. 2. A claim with a named source but no quote. 3. A bare assertion, however confidently worded. Model confidence is a weak signal — a worker that read a truncated document can be entirely sure and entirely wrong — so a self-reported confidence marker is useful mainly as a *tie-break* and as a flag for which claims to verify, not as the primary ranking key. Beware the **independence illusion** too. Two workers agreeing does not mean two independent confirmations if both read the same upstream source, or if both inherited the same framing from the lead's prompt. Correlated errors survive agreement. Weight by distinct evidence, not by headcount. ## Step three: resolve, verify or disclose The lead has three legitimate outcomes and one illegitimate one. - **Resolve** — one side has decisive evidence; take it and keep the citation. - **Verify** — re-dispatch a narrow worker with the exact question and both candidate answers, pointed at the specific sources. This is cheap relative to the original fan-out and is the standard move for high-stakes claims. A verifier given a clean context is also, in practice, better at catching errors than the agent that produced the claim, because it is not anchored on its own reasoning. - **Disclose** — state the disagreement in the output with both claims and sources. For research and analysis products this is usually the *correct* answer, not a cop-out. - **Average** — write a vague sentence that is compatible with both. This destroys information, hides the conflict from the reader, and is the outcome to argue against explicitly in an interview. ## Making synthesis auditable Have the lead's final answer carry citations back to worker returns, and have the returns carry locators back to sources. That chain is what lets you debug a bad answer after the fact: without it, a wrong final claim cannot be attributed to a worker, a source, or the lead's own synthesis, and you are left re-running the whole fan-out to guess. Failure-attribution in multi-agent systems is genuinely hard — step-level attribution accuracy remains low even with full traces as of mid-2026 — so building the citation chain in by construction is worth more than any post-hoc analysis tool. ## Cost note Synthesis is not free either: the lead must read every return and reason across them, and that pass scales with fan-out width. Deep fan-outs often need a two-stage merge — condense batches of returns first, then synthesize the condensations — which trades another layer of information loss for a lead that can still reason well.

  • Why is a worker's self-reported confidence a weak tie-breaker?
    Because confidence tracks the worker's internal coherence, not the quality of what it read. A worker given a truncated transcript or a stale document can be entirely consistent and entirely wrong. Use confidence to decide what to verify and to break ties between otherwise equal evidence, but rank primarily on locatable evidence the lead can re-fetch.
  • Six workers agree and one dissents. Is the majority right?
    Not necessarily. If the six read overlapping sources or inherited the same framing from the lead's prompt, their agreement is one correlated observation, not six independent ones — and the dissenter may be the only worker that looked somewhere new. Weight by distinct evidence and check whether the dissent rests on a source the others never saw.
  • How do you keep a bad final answer traceable to its cause?
    Build the citation chain by construction: the lead's output cites worker returns, and each return cites sources with locators. Then a wrong claim can be attributed to a specific worker, a specific source, or the lead's own synthesis without re-running the fan-out. Attribution in multi-agent traces is otherwise unreliable, so design for it rather than hoping to reconstruct it.

saying these in an interview costs you the question

  • Averaging contradictory briefs into a vague hedge
  • Taking the more fluent or more confident brief
  • Treating worker agreement as independent confirmation
  • Concatenating returns and calling it synthesis
  • Dropping provenance so a wrong claim cannot be traced

context