skip to content

Multi-Agent Evaluation

Scoring a system with many moving parts: end-to-end task success, per-agent quality, reading the full trace, and attributing a failure to the agent that caused it. Interviewers ask because an end-to-end pass rate tells you nothing about which agent to fix.

part ofMulti-agent LLM systemsoverview, primer and where to startread it →
on this pageshow

questions

4

In a multi-agent pipeline, why can every agent score high yet end-to-end success be low?

level: middleimportance: must knowfreq 58%

answer

  1. stages multiply, they do not average
  2. clean fixtures, messy real inputs
  3. downstream agents trust upstream verdicts
  4. compare predicted product to observed rate
  5. the gap is the interaction effect

basics

~20 s

Stage accuracies compound: four agents at 90% each leave roughly 66% end-to-end. Errors also propagate, because downstream agents trust upstream output, and per-agent scores measured on clean inputs never see the messy handoffs real predecessors emit.

solid answer

~50 s

Two effects stack. First, **compounding**: in a sequential pipeline the end-to-end success rate is roughly the product of the per-stage rates, so four 90% agents cap out near 66% even with no interaction at all. Second, **propagation and distribution shift**: per-agent scores are almost always measured on hand-made clean inputs, while in production each agent consumes whatever its predecessor actually emitted — a truncated summary, a hedged verdict, a field the schema allowed but the next agent misreads. A downstream agent has no way to know an upstream verdict was wrong, so it confidently builds on it. Concretely, on a grant-screening pipeline the eligibility agent scored 97% and the scoring agent 71%, which predicts about 69% end-to-end — but observed end-to-end success was 62%. That 7-point gap is the interaction effect, and it is the number that tells you the stages are not independent. Report both views, and always evaluate each agent on real upstream outputs, not curated ones.

code

python · 10 lines
python
stage_rates = {"intake": 0.99, "eligibility": 0.97, "scoring": 0.71, "summary": 0.95}

predicted = 1.0
for rate in stage_rates.values():
    predicted *= rate

observed_end_to_end = 0.62
print(f"predicted from stages: {predicted:.3f}")
print(f"observed end-to-end:   {observed_end_to_end:.3f}")
print(f"interaction gap:       {predicted - observed_end_to_end:.3f}")

go deeper

for a junior

Be able to say that a chain of agents multiplies its stage success rates, so several good stages can still make a poor system, and give the four-stages-at-90-percent example.

for a middle

Explain both mechanisms — arithmetic compounding and propagation of upstream errors into inputs the agent was never evaluated on — and show how comparing the predicted product against the observed end-to-end rate exposes coupling between stages.

for a senior

Demonstrate the operating habit: score every agent on real upstream outputs replayed from traces, add contract checks on handoff payloads, and prove a fix matters by substituting ground truth at one stage before investing in it.

for a principal

Own the measurement policy across teams: which numbers are published to stakeholders versus used for engineering, how the end-to-end rate is decomposed so each team sees its own contribution, and how you stop stage owners from optimising a local score that does not move the system.

## The question behind the question An interviewer asking this wants to know whether you have actually measured a pipeline rather than a model. A single end-to-end pass rate is the number a stakeholder asks for, but on its own it is nearly useless for engineering: it tells you the system fails 38% of the time and nothing about where to spend the next week. Per-agent scores tell you where quality is low but not whether fixing it moves the system. You need both, and you need to understand precisely why they disagree. ## Effect one: accuracies compound In a strictly sequential pipeline where each stage must be correct for the final answer to be correct, end-to-end success is bounded above by the product of the stage success rates. Four independent stages at 90% give 0.9^4 = 0.656. This is arithmetic, not a bug, and it has a blunt consequence: on a long chain, per-stage quality that sounds excellent produces a system that sounds broken. It also tells you where leverage is — the *lowest* stage rate dominates the product, so a stage at 71% is worth far more attention than one at 97%, even though both look like single-digit percentage improvements in isolation. A useful sanity check is to compute the predicted product and compare it to the observed end-to-end rate before drawing any conclusions. ## Effect two: propagation and distribution shift The observed rate is usually *below* the product, and the gap is the interesting part. Three mechanisms produce it. **Input distribution shift.** Per-agent evaluation is typically done against a fixture set someone wrote by hand: well-formed, unambiguous, complete. In production, agent N's input is agent N-1's real output — which may be a summary that dropped a caveat, a verdict phrased as "likely eligible" where the fixture said "eligible", or a JSON object that validates but whose free-text field carries the load. The agent's measured accuracy simply does not apply to the inputs it actually sees. **Blind trust downstream.** LLM agents rarely challenge upstream conclusions. If the eligibility agent wrongly rejects an applicant, the scoring agent does not re-derive eligibility; it scores what it was handed, or is never invoked at all. One early error therefore doesn't just fail its own stage, it silently determines the rest of the run. **Correlated errors.** The same underlying model, the same ambiguous source document, or the same missing retrieval hit can cause several agents to fail together. This breaks the independence assumption in either direction: it can make the observed rate worse than the product (a hard case fails everywhere) or, misleadingly, better (an easy case sails through). ## The grant-screening example A four-agent screening pipeline — intake, eligibility, scoring, summary — was evaluated on 300 past applications with known outcomes. Eligibility scored 97% against its own labelled set, scoring 71%. Multiplying the stage rates predicts roughly 69% end-to-end; the measured end-to-end success was 62%. Two readings follow immediately. The scorer is the bottleneck and deserves the effort budget. And the 7-point unexplained gap says the stages interact — in this case, because eligibility's borderline "conditionally eligible" verdicts, correct by its own rubric, were inputs the scorer had never been evaluated on. ## What to measure instead - **Both layers, always.** End-to-end task success is the number that matters to the product; per-agent scores are the number that tells you what to fix. Publish them side by side. - **Score each agent on real upstream output.** Take production or replayed traces and grade agent N's output given the *actual* input it received, not a fixture. This is the single highest-value change and it usually drops the flattering per-agent numbers. - **Check the handoff, not just the agent.** Add contract checks on the payload crossing each boundary: required fields present, verdict drawn from the expected vocabulary, confidence expressed where downstream logic depends on it. Many "agent quality" failures are interface failures. - **Report the predicted-vs-observed gap** as a first-class metric. It is a direct measure of how coupled your stages are, and it moves when you fix an interface. - **Segment end-to-end failures by which stage first diverged**, so the end-to-end rate decomposes into something actionable rather than one bit per run. ## The trap to avoid Do not optimise a stage because its number is low in absolute terms; optimise the stage whose improvement moves end-to-end success. A 71% scorer sitting behind an eligibility agent that already discards 40% of applications correctly has a smaller reachable population than the raw number suggests. Sensitivity — how much end-to-end success moves per point of stage improvement — is what should drive the roadmap, and you can estimate it cheaply by replaying traces with one stage's output replaced by ground truth.

  • How would you estimate whether fixing your weakest agent actually moves end-to-end success?
    Replay the recorded traces with that agent's output replaced by ground truth and re-run everything downstream. The resulting end-to-end rate is an upper bound on what perfecting that agent buys you. If it barely moves, the stage is not the bottleneck no matter how low its own score is — the loss lives elsewhere, usually in an interface or in a later stage that discards the improvement.
  • Does the compounding argument still apply when agents run in parallel rather than in sequence?
    The multiplication still applies to any agent whose output is required for the final answer, parallel or not — parallelism changes latency, not the correctness dependency. What parallelism does change is propagation: parallel branches cannot corrupt each other's inputs, so the observed rate sits closer to the naive product. The aggregation step that merges the branches then becomes the stage that carries the interaction risk.
  • If per-agent scores look excellent on fixtures, how do you build a more honest per-agent eval set?
    Mine it from real traces: take the actual input each agent received in production or replay runs, and label the output it should have produced. That set carries the true input distribution — truncated summaries, hedged verdicts, unusual field values — so the score reflects the job the agent really does. Keep the fixture set as a regression guard, but never quote it as the agent's accuracy.

saying these in an interview costs you the question

  • Assuming end-to-end success is the average of per-agent scores
  • Claiming high per-agent accuracy proves the pipeline works
  • Evaluating each agent only on hand-written clean fixtures
  • Optimising the lowest-scoring agent without checking end-to-end sensitivity
  • Treating stage errors as independent when they share a model or source document

context

open as a page

How do you attribute one failed multi-agent run to the specific agent that caused it?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Replay the recorded trace to find the first step where state diverged from a correct run, not the last step that errored. Then separate the agent that produced the bad output from the one that should have caught it, and confirm by counterfactual replay.

open as a page

What must a multi-agent trace record for failure analysis to be possible later?

level: middleimportance: should knowfreq 45%

basics

~20 s

One correlation id spanning the whole task, one span per agent turn nested under the orchestrator, and verbatim inputs, outputs, handoff payloads, tool calls, model identity, token counts and latency. Runs are nondeterministic, so the trace is the only faithful record.

open as a page

How would you design a fair evaluation comparing a multi-agent pipeline to one agent?

level: principalimportance: should knowfreq 36%

basics

~20 s

Hold the token and cost budget equal, use the same task set and the same grading, repeat each task many times because both systems are nondeterministic, and report cost and latency per solved task with variance rather than one pass rate from one run.

open as a page