skip to content

How do you attribute one failed multi-agent run to the specific agent that caused it?

level: seniorimportance: must knowfreq 50%

answer

  1. one bit of failure names no owner
  2. first divergence, not last error
  3. producer versus the agent that should catch it
  4. classify so failures aggregate across runs
  5. confirm by replaying with a corrected step

basics

~20 s

Replay the recorded trace to find the first step where state diverged from a correct run, not the last step that errored. Then separate the agent that produced the bad output from the one that should have caught it, and confirm by counterfactual replay.

solid answer

~60 s

Attribution is a credit-assignment problem, and the instinctive answer — blame whichever agent threw the error or produced the bad final answer — is usually wrong, because the damage was done several handoffs earlier. The method is: define the failure precisely as an observable difference from the expected outcome, then walk the trace forward and find the **first divergence** — the earliest step whose output no longer matches what a correct run would hold. That step's owner is the primary suspect. Then ask a second question: which agent had the information and the mandate to catch it? A reviewer or verifier that approved a wrong artifact is a distinct failure mode from the producer that emitted it, and mitigations differ. Classify against a taxonomy — MAST groups multi-agent failures into specification and system-design issues, inter-agent misalignment, and verification failures — so counts aggregate across runs. Confirm the diagnosis by counterfactual replay: pin the trace, substitute a corrected output at the suspect step, re-run downstream, and see whether the run now succeeds.

go deeper

for a junior

Know that a failed multi-agent run needs a trace to diagnose, and that the agent that visibly errored is often not the one that caused the failure.

for a middle

Explain first-divergence analysis over a recorded trace, and distinguish the agent that produced a bad artifact from the agent that failed to verify it, since the two imply different fixes.

for a senior

Show the full loop: precise failure definition, first-divergence walk, counterfactual replay to confirm, and classification against a shared taxonomy so a batch of failures ranks by mode rather than by anecdote.

for a principal

Own the finding that most multi-agent failures are specification and inter-agent misalignment defects rather than model quality, and steer investment toward decomposition, handoff contracts and verification design instead of per-agent prompt tuning.

## Why attribution is the hard part An end-to-end failure in a multi-agent system yields one bit: the task did not succeed. That bit does not name an owner, a mitigation, or a next action. Attribution — deciding which agent's behaviour caused this particular failure — is what converts a pile of failed runs into a ranked work list. It is also genuinely hard: the observable symptom is usually far downstream of the cause, several agents have touched the state in between, and the run is not reproducible by re-executing it. ## Step one: define the failure crisply Before attributing anything, write down what "failed" means for this task in checkable terms: the wrong applicant was rejected, the summary contradicted the source, the pipeline terminated without an answer, the run exceeded budget. Vague failure definitions produce vague attribution and make aggregation across runs meaningless. If a run fails several ways at once, attribute each failure separately rather than picking one. ## Step two: find the first divergence, not the last error Walk the trace from the start and compare each step's output against what a correct run would have held at that point — either from a reference run, from ground-truth labels on the task, or from a human reading the artifact. The first step whose output is already wrong is the primary suspect. This matters because the loudest signal is almost always the last one: an agent that crashes on a malformed input, or a final summariser that produces a confidently wrong answer, is frequently an innocent consumer of an artifact that went wrong three handoffs earlier. Blaming the last error produces a mitigation (better error handling in the summariser) that leaves the actual defect untouched. In practice the first divergence is often not a wrong fact at all but a **misalignment**: two agents holding different interpretations of what the task is, or a handoff payload that is well-formed but carries less than the receiving agent assumed. Those show up in the trace as an output that looks locally plausible and only becomes wrong in light of the receiver's expectations. ## Step three: separate producing from failing to catch Every failing run has at least a producer and, if the design includes any verification, one or more agents that had a chance to stop it. These are different defects with different fixes. If the reviewer agent approved a wrong artifact, the reviewer is not innocent — its rubric, its context, or its incentive to agree is the defect, and improving the producer alone will not raise the system's reliability. Record both: primary attribution (who produced the divergence) and secondary (who should have caught it). ## Step four: classify against a taxonomy One-off narratives do not aggregate. Score each failing run against a shared taxonomy so you can count and rank. MAST is the reference taxonomy for multi-agent failures: fourteen modes in three clusters — specification and system-design problems, inter-agent misalignment, and verification failures — with published distributions putting roughly 42%, 37% and 21% of observed failures in those clusters respectively. The exact percentages matter less than the shape of the finding: most multi-agent failures are not "the model was dumb", they are design and communication defects that a better prompt on one agent will not fix. ## Step five: confirm with counterfactual replay Attribution from reading a trace is a hypothesis. Test it: pin the recorded trace, replace the suspect step's output with a corrected one, and re-execute everything downstream. If the run now succeeds, the attribution holds. If it still fails, the real cause is elsewhere or there are two independent defects. This is the closest thing to a controlled experiment available in a nondeterministic system, and it is cheap because upstream steps are replayed from the record rather than re-executed. ## Automating it, and its limits Attribution can be automated with a judge model reading the trace, and there are dedicated benchmarks for exactly this task — Who&When, AgenTracer, MP-Bench and TraceElephant among them. Be honest about the current state: even with full traces available, step-level attribution accuracy for automated attributors remains around 30%. That is well above chance on a long trace and useful for triage and for ranking, but not reliable enough to close a bug on its own. The practical posture in 2026 is to use an automated attributor to cluster and prioritise, and to have a human adjudicate a sample — typically the highest-frequency clusters — before committing engineering effort. ## What you do with the result Aggregate attributions across a batch of failing runs — forty is enough to see structure — and act on the mode, not the anecdote. A cluster of inter-agent misalignment failures points at the handoff contract and the task specification, not at any single agent's prompt. A cluster of verification failures points at giving the reviewer a clean, independent context rather than the producer's own reasoning. A cluster of specification failures usually means the orchestrator's decomposition is wrong and no amount of per-agent tuning will help.

  • Why can't you just re-run the failing task to find the cause?
    Agentic runs are nondeterministic: sampling, tool responses, retrieval results and timing all vary, so the rerun frequently succeeds or fails somewhere else entirely. The recorded trace is the only artifact that pins what actually happened. Reruns are still useful for a different purpose — estimating how often a failure reproduces — but they cannot substitute for the trace when locating a specific run's cause.
  • How do you attribute a failure when two agents ran in parallel and both outputs look plausible?
    Attribute at the aggregation point first: check whether the merging step had enough information to detect the conflict. If it did and merged anyway, that is a verification failure at the aggregator. If neither branch output was individually wrong but their combination was inconsistent, the defect is in the specification — the decomposition allowed two branches to answer overlapping questions without a shared constraint.
  • What do you record at runtime so this analysis is possible after the fact?
    Verbatim inputs and outputs for every agent turn, the handoff payload crossing each boundary, tool calls with arguments and results, model identity and version, and a single correlation id linking all of it to one task. Without the verbatim payloads you can see that a step happened but not what it decided, which is exactly the information first-divergence analysis needs.

saying these in an interview costs you the question

  • Blaming the last agent that errored or produced the final answer
  • Treating a reviewer that approved a wrong artifact as blameless
  • Re-running the task and assuming the failure reproduces identically
  • Writing one-off failure narratives that never aggregate into counts
  • Trusting an automated attributor's step-level verdict without human adjudication

context