skip to content

Your five-persona agent crew performs no better than one agent — how do you diagnose it?

level: principalimportance: should knowfreq 44%

answer

  1. what differs between the agents, really
  2. prose is the cheapest lever, and the weakest
  3. smaller window, smaller action space
  4. ablate one variable at a time
  5. always keep a single-agent baseline

basics

~20 s

Check whether the roles differ in anything except wording. If all five share the same context and the same tools, the personas are decoration — the measurable gains in multi-agent systems come from context isolation and tool scoping, not from character descriptions.

solid answer

~50 s

First I check what actually differs between the five agents. If they read the same history and hold the same tool set, five personas are five rewordings of the same distribution, and paying five times the tokens for that is the expected outcome — no gain. Then I ablate deliberately. Hold context and tools fixed and vary only persona text; you will usually see small effects. Then hold persona text fixed and vary tool scope and what each agent sees; that is where the movement is. Frameworks with `role`, `goal` and `backstory` fields invite the opposite emphasis, which is why crews get tuned in the least effective place. Finally I ask whether the roles partition the work at all. If every agent could do every subtask, there is no specialization — just duplication. And I am honest that at equal token budget a well-built single agent is a genuinely competitive baseline in 2026.

go deeper

for a junior

Know that writing different personalities for agents is not what makes a crew work, and that what each agent sees and can call matters more than how it is described.

for a middle

Explain the two mechanisms — a smaller context window and a smaller tool set — and why long shared contexts degrade reasoning. Be able to say what you would change first in a crew that shows no gain.

for a senior

Run the diagnosis: an eval set, a single-agent baseline, one-variable-at-a-time ablations, repeated runs, and cost reported alongside score. Expect to defend the conclusion that a crew should be cut back to one agent.

for a principal

Own why teams land here — prose is the cheapest lever to pull and the hardest to disprove without an eval harness — and set the rule that no additional role ships without evidence that the previous one earned its place.

## The symptom and what it usually means A crew of five carefully written personas that scores the same as one agent is not an unlucky result — it is the default result when the roles differ only in prose. This is worth stating plainly because the intuition runs the other way: it feels as though giving a model a distinct identity, goal and backstory should sharpen its behaviour, and frameworks that expose exactly those fields reinforce the feeling. The empirical picture is less flattering. When tool access and visible context are held constant, changes to persona wording move outcomes by small amounts on most tasks. The levers that move outcomes substantially are the ones that change *what the agent sees* and *what the agent can do*. ## Lever one: context isolation Model quality degrades as the context window fills with material — the standard term is context rot. An agent carrying twenty tool results, three abandoned hypotheses and a long transcript reasons worse about the next step than the same model given only what that step requires. That is the real mechanism behind subagent specialization. A role that receives a narrow brief, works in its own window, and returns a short summary is not smarter because it was told it is a specialist; it is more accurate because its window contains a high ratio of relevant to irrelevant tokens. The same effect explains why a reviewer role with a deliberately clean context catches defects the producing agent missed: the producer's window is saturated with the reasoning that led to the defect, and that reasoning is exactly what makes the defect look correct. So the diagnostic question is: does each of my five agents see a *different, smaller* slice of the problem? If they all see the same full history, I have five copies of one degraded context. ## Lever two: tool scoping The second lever is the action space. Tool-selection accuracy falls as the number of available tools rises, so an agent holding four relevant tools makes better choices than the same model holding thirty. Scoping tools per role therefore improves behaviour directly, independently of anything written in the prompt — and it does so in a way the model cannot override, because it is enforced outside the model. The combination is what specialization actually is: a smaller window and a smaller action space, with prose serving mainly to set priorities and output shape within those bounds. ## How to run the diagnosis Ablate, one variable at a time, against a fixed eval set. 1. **Baseline.** One agent, full tools, doing the whole task. Record score, tokens and wall-clock time. People skip this step and then cannot tell whether the crew helped. 2. **Persona ablation.** Keep the five agents, their identical tools and identical context, and replace all five personas with one neutral instruction. If the score barely moves, the personas were decoration. 3. **Scope ablation.** Restore the personas but give each role only the tools its job needs. Re-measure. Movement here tells you the action space was the problem. 4. **Isolation ablation.** Give each role only its own slice of context, returning a compressed result to the orchestrator. Re-measure. This is usually where the largest change appears. 5. **Partition check.** Ask whether any two roles could swap tasks without noticing. If yes, they are not specialized; they are duplicated. Run each configuration several times, because agent runs are nondeterministic and a single sample will happily confirm whichever hypothesis you started with. ## What to do with the answer If the ablations show the gain lives in isolation and scoping, redesign around those: fewer roles, each with a genuinely narrow window and a genuinely narrow tool set, with prose reduced to a job statement and a return contract. If the ablations show no configuration beats the single-agent baseline at equal token budget, that is a real and publishable result for your workload — 2026 papers report exactly this on a range of tasks — and the right move is to keep the single agent and spend the budget on better tools, better retrieval or a verification pass. A verification role is the one addition that tends to survive this analysis, because its value comes from being structurally uncontaminated rather than from being described as critical. ## The organisational reason this happens Worth naming, because a principal-level answer should. Persona text is the cheapest thing to change and the easiest to feel good about; it requires no plumbing, no permission model and no evaluation. Context partitioning and tool scoping require both engineering and an eval harness to prove they helped. Teams optimise the visible, cheap lever and then conclude multi-agent does not work. The fix is process as much as design: no role ships without a baseline comparison, and no crew grows a sixth agent without an ablation showing the fifth earned its place.

  • Why does a reviewer agent with a clean context catch defects the producing agent missed?
    Because the producer's window is saturated with the reasoning that created the defect, and that reasoning is exactly what makes the defect look correct. A reviewer starting from the artifact alone has no such commitment and a far higher ratio of relevant to irrelevant tokens. The gain is structural — it comes from the clean window, not from calling the role a critic.
  • How would you design the ablation to avoid fooling yourself?
    Fix the eval set first, then change one variable at a time: personas neutralised, tools scoped, context isolated. Always keep a single-agent full-task baseline with matched token budget. Run every configuration multiple times, because agent runs are nondeterministic and one sample will confirm whatever you hoped. Report tokens and wall-clock alongside the score, since a gain bought at five times the cost may not be a gain.
  • If ablation shows the single agent wins at equal budget, what do you do?
    Keep it, and say so. That result is common enough in 2026 that it should not be treated as a failure of execution. Spend the freed budget on the things that did move the needle for the workload — better tools, better retrieval, a verification pass — and revisit multi-agent only where the work is genuinely independent and slow enough that isolation or parallelism pays.
  • Do personas ever matter?
    Yes, as a weak but real lever. They set vocabulary and priorities — a verifier told to weigh reachability rather than severity ranks evidence differently — and they shape output register, which matters when one agent's output is another's input. The error is treating prose as the primary design surface when window contents and tool scope dominate the measured outcome.

saying these in an interview costs you the question

  • Believing distinct personas alone make agents behave differently
  • Giving every agent in a crew the same context and tools
  • Tuning backstory prose instead of scoping tools and context
  • Adding roles without a single-agent baseline to compare against
  • Judging a crew change from a single nondeterministic run

context