How would you design a fair evaluation comparing a multi-agent pipeline to one agent?
answer
- uncapped spend confounds every result
- equal token budget, several budget levels
- distributions, not one run per task
- cost per solved task decides, not raw quality
- segment by parallel-read versus sequential-write
basics
~20 sHold the token and cost budget equal, use the same task set and the same grading, repeat each task many times because both systems are nondeterministic, and report cost and latency per solved task with variance rather than one pass rate from one run.
solid answer
~50 sThe comparison that convinces nobody is one run of each on a handful of tasks with an uncapped budget, because the multi-agent system will usually win simply by spending more. Design it as a controlled experiment. **Equalise the budget.** Orchestrated multi-agent runs can consume on the order of fifteen times the tokens of a single-agent run, so quality per run and quality per dollar are different questions. Fix a token or cost ceiling per task and let each system spend it however it wants. Several 2026 results report single-agent systems matching or beating multi-agent designs at equal token budget, which is exactly the finding an uncapped comparison hides. **Repeat and report variance.** Both systems are nondeterministic; run-to-run spread often exceeds the difference you are trying to measure, so a single run per task is not evidence. **Strengthen the baseline.** A weak single-agent baseline with a thrown-together prompt inflates the multi-agent win. And segment results by task type — read-heavy parallel work and write-heavy sequential work give opposite answers.
go deeper
Know that comparing an agentic architecture needs the same tasks, the same grading and more than one run, because these systems give different answers each time.
Explain why an uncapped comparison is confounded by spend, and be able to name the metrics that matter beyond success rate — cost per solved task, latency, and run-to-run variance.
Design the experiment: budget sweep with equal ceilings per arm, repeated runs with reported intervals, a genuinely strong single-agent baseline, and segmentation by whether the work is parallel-read or sequential-write.
Own the decision process, not just the numbers: pre-register the adoption rule, separate ownership of the baseline from the team that built the system, and be willing to report a crossover that argues against the architecture already under construction.
## Why this comparison is so easy to get wrong Multi-agent architectures are expensive, complicated and fashionable, which is a combination that produces motivated evaluation. The default comparison — build the multi-agent system, run it against a single-agent version somebody wrote in an afternoon, on twenty tasks, once each, with no budget limit — will nearly always favour the multi-agent design, and it will not predict production behaviour. Designing this evaluation honestly is a leadership task because the result determines whether an organisation takes on a large operational burden. ## Control one: equal budget The dominant confound is spend. An orchestrator that fans out to subagents and reads their summaries consumes far more tokens than a single agent on the same task — the widely cited figure for research-style orchestration is roughly fifteen times a chat interaction. If the multi-agent system is allowed to spend fifteen times as much, a quality win tells you almost nothing: you have not learned that the architecture is better, only that more compute helps. The fix is a per-task budget, enforced identically on both arms, with each system free to allocate it as it likes — the single agent may use it on longer reasoning or more tool calls, the multi-agent system on more subagents. Run the sweep at several budget levels rather than one, because the interesting result is usually a crossover: single-agent leads at low budget, multi-agent catches up or overtakes at high budget, or never does. That curve is the actual deliverable. ## Control two: repeated runs and variance Both systems sample, so both are distributions, not points. Multi-agent systems typically carry *more* variance because there are more sampled decisions and more places for a run to diverge. Report the distribution: run each task multiple times, give the mean with an interval, and state how many runs per task the number rests on. A five-point difference in mean success is not a finding when the per-system spread is fifteen points. Reporting a single run per task is the most common way a multi-agent comparison manufactures a result that does not replicate. ## Control three: an honest baseline The single-agent arm must be the best single-agent system you know how to build, not a strawman. That means the same tool access, comparable prompt effort, and any single-agent techniques that apply — a longer effort budget, self-critique within one context, or a fixed workflow if a fixed workflow would genuinely do the job. If the multi-agent design only wins against a baseline nobody tried on, the evaluation is measuring prompt effort rather than architecture. Symmetrically, hold everything else constant across arms: the same task set, the same grading procedure, the same underlying models, the same retrieval corpus. Change one thing. ## Control four: measure what the decision needs Quality alone does not decide anything. Report at minimum: - **Success rate** at each budget level, with variance. - **Cost per solved task** — this is often the number that flips the decision, because a system that succeeds slightly more often at four times the cost is worse for most products. - **Latency**, including tail latency. Fan-out can improve wall-clock time when branches are genuinely parallel and can destroy it when they are sequential. - **Failure profile**, not just failure count. A system that fails loudly and recoverably is not equivalent to one that fails silently with a confident wrong answer, even at identical rates. ## Control five: segment by workload Aggregate results across a mixed task set hide the finding. The consistent 2026 picture is that gains come from parallel, read-and-verify shaped work — breadth-first research, independent checks, review passes given a deliberately clean context — while sequential work with a single writer sees little benefit and pays the full coordination cost. Report per-segment, because the honest answer is usually "multi-agent for this class of task, single agent for that one", and an aggregate number forces a decision that neither segment supports. ## Be honest that the ground is contested This is one of the genuinely unsettled questions in the field as of 2026. Public positions have moved in both directions within a year, results are sensitive to budget normalisation and task mix, and reported wins frequently shrink under controlled comparison. An interviewer at this level is not looking for a verdict; they are looking for whether you can design a comparison whose result you would still trust if it went against the architecture you had already started building — and whether you would define the decision criteria before running it, rather than after seeing the numbers. ## Reporting it Pre-register the decision rule: state before the run what result would make you adopt, reject or defer. Publish the budget levels, run counts, task segments and the strengthened baseline's configuration alongside the numbers. Without that, the evaluation is a narrative and every reader is entitled to discount it.
- Your multi-agent system wins at high budget and loses at low budget. What do you report?Report the crossover point and the budget at which the product actually operates. If real traffic runs below the crossover, the multi-agent design loses for this product regardless of the headline win. The curve is also a planning input: it tells you what per-task spend would be required to make the architecture pay, which is a concrete number to weigh against the operational burden.
- How do you keep the evaluation honest when the team has already invested months in the multi-agent build?Pre-register the decision rule and the metrics before the run, have someone who did not build the system own the single-agent baseline, and fix the task set in advance so it cannot be trimmed after seeing results. Sunk cost shows up as post-hoc segment selection and quietly weakened baselines, so remove both degrees of freedom before any number exists.
- Why might the multi-agent system show higher variance, and does that matter commercially?More sampled decisions and more handoffs mean more places a run can diverge, and a divergence early in the graph changes everything downstream. It matters commercially because users experience the tail, not the mean: a system with the same average quality but a fatter bad tail generates more support load and erodes trust faster. Report tail behaviour explicitly, not just the average.
saying these in an interview costs you the question
- Comparing quality with no cap on tokens or cost
- Reporting one run per task as if it were a stable measurement
- Benchmarking against a weak single-agent baseline nobody tuned
- Aggregating across mixed task types instead of segmenting
- Choosing the winning metric after seeing the results