skip to content

Multi-agent LLM systems

5 roadmaps41 questionsupdated

Getting several agents to work as one system: who plays which role, how they talk, which topology routes the work, what state they share, and what happens when they disagree or stall. Interviewers ask because multi-agent designs multiply capability and failure modes at the same rate.

on this pageshow

guide

overview

~1 min

Multi-agent systems questions test whether you can split one job across several model-driven agents and still reason about the whole. Interviewers care less about which framework you have used than about whether you see the bill: every extra agent adds tokens, latency, a boundary where meaning can drift, and one more place for a failure to hide. A good answer justifies the second agent before it designs the tenth, and at the senior end several questions are really asking when a single agent would have done better. The subject splits into six sections. [Agent roles and specialization](/topics/found-multi-agent-systems-agent-roles) covers what makes a role real rather than cosmetic. [Orchestration topologies](/topics/found-multi-agent-systems-orchestration-topologies) is the largest: orchestrator-worker fan-out, supervisor routing, peer handoffs, and debate or critique ensembles, with the question of when each fits. [Inter-agent communication](/topics/found-multi-agent-systems-communication-protocols) is about how work crosses from one agent to another and what happens when a message arrives broken. [Shared state and context](/topics/found-multi-agent-systems-shared-state) covers the durable store agents coordinate through and who may change it. [Coordination and conflict resolution](/topics/found-multi-agent-systems-coordination-conflict) treats a fleet of agents as the distributed system it is: races, stalls, contradictions, restarts. [Multi-agent evaluation](/topics/found-multi-agent-systems-evaluation) asks how to score the whole and find the part to fix. Start with roles, because every topology is a way of arranging them. Topologies come next, then communication and shared state, the two channels every topology uses. Coordination and evaluation come last, since both assume you can already draw the system you are defending.

primer

A few ideas run under every section. Once they are in place, most questions below become a matter of applying them to a specific design. - **An agent is what it can see and what it can do.** Its context window and its tool scope separate it from its neighbours. A new persona laid over identical context and identical tools changes the wording, not the capability. Most of the measurable benefit of splitting work comes from isolation: a worker explores in its own window and hands back something much smaller than what it read. - **Every boundary is a lossy, priced channel.** Moving work between agents costs input tokens on the receiving side and risks the receiver reading the task differently from the sender. A typed contract makes the boundary checkable; free prose leaves it to interpretation. - **Parallelism suits independent reading, not shared writing.** Fan-out works when slices do not overlap and the results can be combined afterwards. When several agents must change the same artifact, you have built a coordination problem, and adding agents makes it worse. - **Errors compound and correlate.** A chain of stages multiplies its failure chances, and downstream agents tend to trust what upstream handed them. Agents built on the same model and prompt also tend to fail in the same way, so agreement between them is weaker evidence than it looks. Redundancy helps only against errors that are independent. - **Durable state lives outside the context window.** Windows are small and disappear with the process; an external store and a log give restartability, inspection and one version of the truth. The design question that follows is who is allowed to write. - **Control needs an owner.** Something has to decide who acts next, when a step has run too long, when a loop must be broken and when the task is finished. Left to the agents themselves, those decisions tend to turn into ping-pong and silent stalls. - **You cannot fix what you cannot attribute.** An end-to-end score tells you whether the system works. Only a complete trace tells you which agent to change.

Orchestrator
The lead agent that splits a task into subtasks, dispatches them to worker agents in their own contexts, and combines what they return.
Supervisor
A routing agent that, on each turn, picks which specialist should act next or declares the task complete, keeping control centralized.
Handoff
A transfer of the live conversation from one agent to a peer, which then continues under its own instructions and tools.
Fan-out and gather
Dispatching independent subtasks to several agents in parallel, then collecting and merging their results in one place.
Tool allowlist
The explicit set of tools an agent may call. It bounds what a role can actually do, independent of what its prompt says.
Context isolation
Running each agent in a separate context window so its exploration, dead ends and raw material do not crowd other agents' windows.
Handoff contract
The agreed shape of what passes between two agents: required fields, types and limits, checked at the boundary instead of trusted.
Artifact store
Shared durable storage for the files and results agents produce, referenced by path or id rather than copied into every context.
Append-only event log
A record of every state change in order, never edited in place, used to rebuild state, audit decisions and resume after a crash.
Single-writer discipline
A rule that one owner performs all writes to a given piece of shared state, while other agents may only propose changes.
Hop budget
A cap on how many handoffs or routing steps one task may take before the system stops and escalates, used to break loops.
Correlated errors
Mistakes that several agents make together because they share a model, prompt or context; voting and debate cannot cancel them.

The sections describe one system from different angles, and they depend on each other in a fairly fixed order. **Roles are the parts; topologies are the wiring.** A topology only makes sense once you know what each node is allowed to see and do. Orchestrator-worker, supervisor routing, peer handoff and debate differ mainly in who holds control and where results come back to. The central designs keep one agent accountable at the cost of that agent's context filling up; the handoff designs spread control out at the cost of auditability. Questions in [orchestration topologies](/topics/found-multi-agent-systems-orchestration-topologies) usually ask you to pick one for a workload and defend how centralized it is. **Communication and shared state are the two channels.** Agents either pass information directly (a chained prompt, a handoff payload, a message on a queue, a call to a remote agent over a cross-party protocol such as A2A) or leave it somewhere others can read. Direct messages are simple and ephemeral; a shared store is durable but raises the question of write authority. Most real systems use both, with small messages carrying references into the store. **Coordination is what the channels need to stay sane.** Once agents run concurrently, the distributed-systems problems arrive: two writers, a worker that never returns, a crash halfway through, two agents that each believe they succeeded. The fixes in [coordination and conflict resolution](/topics/found-multi-agent-systems-coordination-conflict), such as single ownership, deadlines, checkpoints and independent verification, borrow from ordinary distributed systems, adjusted for participants that do not reliably follow a protocol. **Evaluation closes the loop.** [Multi-agent evaluation](/topics/found-multi-agent-systems-evaluation) depends on everything above being observable. If handoffs are typed and every turn is traced, a failed run can be replayed and blamed on one agent; if they are prose and untraced, all you have is a pass rate.

  1. Agent Roles and Specialization →

    Every design is an arrangement of roles, so first learn what makes a role more than a persona: its context, its tools and its contract.

  2. Orchestrator-Worker →

    The most common topology and the clearest case of context isolation paying off; the other topologies are easiest to read as variations on it.

  3. Inter-Agent Communication →

    Every topology moves work across boundaries; learn what each kind of message costs and how a malformed one is caught.

  4. Shared State and Context →

    The durable half of communication, and where questions about write authority, auditing and restartability begin.

  5. Coordination and Conflict Resolution →

    With the channels understood, study how concurrent agents race, stall and contradict each other, and the controls that contain it.

  6. Multi-Agent Evaluation →

    Last, because scoring and blaming a multi-agent run assumes you know every part that could have caused the failure.

  • Proposing several agents without first saying why one agent with good tools would fall short; interviewers often want that baseline defended.

  • Treating role prompts as the specialization while every agent shares the same context and tools; the split then adds cost and no capability.

  • Passing free-text summaries between agents and trusting them, so a missing or misread detail is silently filled in downstream.

  • Letting several agents write the same shared artifact directly, then trying to fix the resulting conflicts with retries.

  • Counting agreement between agents on one model and one prompt as independent confirmation; they tend to be wrong together.

  • Leaving termination to the workers: no deadline, no hop budget, no loop detection, so one stuck agent blocks the whole run.

  • Routing every specialist's full output back through the lead or supervisor until its own window becomes the bottleneck.

  • Reporting only an end-to-end pass rate, or comparing a multi-agent system to one agent on unequal token budgets.

The same few decisions come up in almost every design question here, and naming the one you are making usually earns more credit than the diagram. - **One agent versus several.** More agents buy parallel exploration and clean, separate contexts; they cost tokens, latency and debuggability. The honest default is the smallest number that the task's independent parts justify. - **Central control versus local handoff.** A lead or supervisor keeps decisions auditable and policy in one place, but its context becomes a bottleneck. Peer handoffs keep each conversation lean and suit sequential flows, but routing becomes harder to predict and to govern. - **Direct messages versus shared state.** Messages are simple and leave no trace unless you log them; a store survives restarts and can be inspected, but needs rules about who may write. - **Full history versus summary at a boundary.** Passing everything preserves detail and costs tokens and focus; passing a summary is cheap and risks dropping the one fact that mattered. - **Redundancy versus cost.** Debate, voting and independent critics catch some errors at a multiple of the price, and only when the agents can actually fail independently. A deterministic check, where one exists, is usually cheaper.

A handful of shapes recur across the sections under different names; spotting them is the quickest way to place a new question. - **Isolate, then compress.** A worker, a subagent or a private scratchpad does messy work out of sight and releases only a short, checked result. It appears in orchestrator-worker designs, in shared-state promotion rules and in context budgeting. - **Contract at every boundary.** Typed payloads, worker briefs with an explicit scope and output shape, and structured handoff summaries are one idea: make what crosses the boundary checkable. - **One owner per decision.** A single writer for shared state, an orchestrator that owns termination, a supervisor that owns routing: conflicts shrink when exactly one party decides. - **Budget and escalate.** Deadlines on subtasks, hop limits on handoffs, retry caps on failed workers and round limits on debate all bound the work and pass whatever is left over to a person or a fallback path. - **Verify independently.** Fresh-context critics, judges given the evidence rather than the arguments, and checks against the original request all avoid asking an agent to grade itself. - **Record everything, replay later.** Event logs, checkpoints and full traces serve restart, audit and failure attribution with the same data.

explore

report an issue with this guide →

questions

page 1 of 2

What does prompt chaining between two LLM agents cost in tokens and fidelity?

level: juniorimportance: must knowfreq 60%

answer

  1. the simplest possible agent-to-agent channel
  2. output text becomes the next prompt
  3. billed again on every turn
  4. errors arrive looking like facts
  5. condense to named fields instead

basics

~20 s

Prompt chaining pastes one agent's finished output into the next agent's prompt. It is the cheapest wiring to build, but the receiver pays input tokens for every word on every turn and inherits any error, hedge or ambiguity verbatim.

solid answer

~50 s

Prompt chaining is the simplest inter-agent channel: agent A produces text, and that text is dropped into agent B's prompt as context. Two costs follow. The **token cost** is structural — a 6k-token analysis pasted into B is 6k input tokens on B's *first* turn and on every subsequent turn of B's loop, because the transcript keeps growing; a multi-hop chain multiplies this. The **fidelity cost** is subtler: B has no way to distinguish A's verified findings from A's speculation, so a confident hallucination arrives looking exactly like a fact, and long pasted blocks degrade B's attention to its own instructions. The fix is not to stop chaining but to narrow the seam: have A emit a short, named set of fields or a compressed summary that B can check, rather than its whole working transcript.

code

python · 23 lines
python
def chain_raw(output_a: str) -> str:
    return (
        "Here is the analysis from the previous agent:\n"
        f"{output_a}\n\nNow draft the supplier reply."
    )


def chain_fields(result: dict) -> str:
    return (
        f"supplier={result['supplier']} "
        f"unit_price={result['unit_price']} "
        f"currency={result['currency']} "
        f"lead_days={result['lead_days']}\n\nNow draft the supplier reply."
    )


print(len(chain_raw("word " * 4800)))
print(len(chain_fields({
    "supplier": "Northwind Metals",
    "unit_price": 1240,
    "currency": "EUR",
    "lead_days": 14,
})))

go deeper

for a junior

Be able to say plainly that chaining means one agent's output text becomes part of the next agent's prompt, and that the receiver pays input tokens for all of it.

for a middle

Explain that the block is re-billed on every turn of the receiver's loop, and that free text carries no provenance — so the receiver cannot separate a verified finding from a guess.

for a senior

Show the production fix: subagents return a condensed result, artifacts move by reference, and fields carry where they came from. Be ready to estimate the token bill of a three-hop chain.

for a principal

Own the tradeoff between build speed and seam discipline. Argue when a chain should stay a string in a prompt and when it must become a versioned contract, and tie that decision to who releases each side.

## What prompt chaining is Prompt chaining is the most basic way two LLM agents talk: agent A runs, produces output text, and that text is inserted into agent B's prompt — usually as a block like *"Here is the analysis from the previous step: …"*. No protocol, no schema, no transport beyond a string concatenation in your own code. It is how almost every multi-agent system starts, and for short pipelines it is often the correct answer. It is worth naming what chaining actually transfers: **text, and nothing else**. The receiving agent does not get A's tool results, A's confidence, A's retries, or A's reasoning state. It gets the words A happened to emit, in a prompt slot that carries no more authority than any other sentence in the context. ## The token cost A language model is charged on input tokens for the whole context on **every** call. If agent A emits a 6,000-token report and you paste it into agent B, that 6,000 tokens is billed on B's first model call — and again on B's second call, and its third, because B's own transcript (its tool calls and results) is appended to a context that still contains the pasted block. An agent that takes ten turns to finish has paid for that block ten times. Chains compound this. In A → B → C, if each step forwards what it received plus what it produced, context grows super-linearly and the last agent in the chain is the most expensive one to run. This is a large part of why orchestrated multi-agent systems burn far more tokens than a single-agent chat for the same task — reported multiples of an order of magnitude are common for research-style pipelines. ## The fidelity cost The fidelity problem is the one candidates usually miss. - **Provenance is erased.** A's grounded citation and A's guess arrive in B's prompt as the same kind of text. B has no signal for which is which, so B will treat a fabricated supplier price with the same seriousness as a retrieved one. Errors do not get filtered by the hop; they get laundered by it. - **Attention is diluted.** Model quality degrades as the window fills — the common name for this is *context rot*. A large pasted block competes with B's own system prompt and task instructions, and the longer the block, the more likely B drifts from what it was actually asked to do. - **Ambiguity survives.** A's hedged sentence ("the lead time is probably around two weeks") becomes B's input. B must either re-derive the uncertainty or, more often, flatten it into a confident downstream claim. - **Format drift.** Because the channel is free text, A can change its output shape between runs — a heading disappears, a list becomes prose — and B's parsing or reasoning silently changes with it. There is no failing test at the boundary; there is only a worse answer. ## Narrowing the seam The standard remedy is to make the handoff *smaller and more structured*, not to abolish it: 1. **Have A emit a condensed result, not its transcript.** In production orchestrator/subagent designs, a subagent that did a great deal of work returns on the order of one to two thousand tokens of findings. The exploration stays in A's own context and dies with it. 2. **Give the payload named fields.** `supplier`, `unit_price`, `currency`, `lead_days` is both cheaper and checkable; the receiver can validate it before spending a model call on it. 3. **Pass references, not contents.** Hand over an identifier or a path to an artifact and let B fetch only what it needs. 4. **Mark provenance.** If a field came from a tool result rather than from the model's own inference, say so in the payload, so B can weight it. ## When plain chaining is right Do not over-engineer. If the pipeline is two steps, the intermediate output is a few hundred tokens, and both steps ship from the same codebase, a string in a prompt is the correct amount of machinery. The cost curve only bites when outputs are large, the chain is long, or the steps are owned by different teams — at which point the seam deserves a contract. ## What interviewers listen for A weak answer describes chaining as "just passing the output along" and stops. A strong answer prices it: input tokens billed per turn and re-billed as the loop runs, plus the loss of provenance that lets one agent's mistake become the next agent's premise — and then proposes the condensed, named-field handoff as the fix.

  • If the receiving agent runs a ten-turn tool loop, how many times is the pasted block paid for?
    Once per model call, so roughly ten times. Each turn re-sends the full context — the pasted block plus everything the agent has appended since — as input tokens. This is why a large handoff is far more expensive than the single paste suggests, and why condensing before the hop, rather than after, is what saves money.
  • How would you keep provenance across the hop without shipping the whole transcript?
    Emit a structured result where each field carries where it came from — a tool result, a retrieved document id, or the model's own inference. The receiver can then trust measured fields and treat inferred ones as claims needing verification. It costs a few tokens per field and prevents the receiver from promoting a guess into a premise.
  • When is chaining raw output actually the right choice?
    When the output is small, the chain is short, and both ends ship from the same codebase. Structuring a 200-token handoff between two steps you deploy together adds schema maintenance for no measurable gain. Reach for a contract when payloads grow, hops multiply, or the two ends stop being released at the same time.

Forwarding an entire email thread instead of writing the one-line summary the reader needs: everything is technically there, the reader pays to wade through it, and a wrong claim buried three replies down looks as authoritative as the rest.

saying these in an interview costs you the question

  • Assumes context is free because the model accepts it
  • Pastes the whole upstream transcript, then blames the model
  • Believes the receiver can tell fact from speculation in text
  • Treats an upstream hallucination as verified once it is quoted
  • Thinks chaining transfers the sender's reasoning state

context

open as a page

What defines an agent's role in a multi-agent system beyond its persona prompt?

level: middleimportance: must knowfreq 62%

basics

~20 s

A role is three things: a persona stating the job, a tool allowlist bounding what the agent can actually do, and behavioural constraints covering when it stops and what it returns. The allowlist does most of the real work.

open as a page

Why send schema-validated payloads between agents instead of free-text messages?

level: middleimportance: must knowfreq 62%

basics

~20 s

A schema turns an implicit agreement into a checkable one. Required fields, types and enums let the receiving agent reject a malformed message at the boundary with a precise error, instead of a model quietly inventing the missing value and shipping a wrong result downstream.

open as a page

Why do multi-agent systems use single-writer discipline instead of locking concurrent writers?

level: middleimportance: must knowfreq 68%

basics

~20 s

Locking assumes writers that wait, know their own blast radius, and roll back cleanly. LLM agents do none of that. Routing every write through one owner, while the others contribute proposals, removes the conflict rather than arbitrating it.

open as a page

In a multi-agent pipeline, why can every agent score high yet end-to-end success be low?

level: middleimportance: must knowfreq 58%

basics

~20 s

Stage accuracies compound: four agents at 90% each leave roughly 66% end-to-end. Errors also propagate, because downstream agents trust upstream output, and per-agent scores measured on clean inputs never see the messy handoffs real predecessors emit.

open as a page

In multi-agent debate, what do extra rounds buy beyond a single critique pass?

level: middleimportance: must knowfreq 62%

basics

~20 s

Extra debate rounds let each agent see the others' revised arguments and change its own, which corrects confident first-shot errors. Most of the lift lands in the first one or two exchanges; after that agents converge and mostly restate a consensus that may be wrong.

open as a page

In an agent handoff, what actually changes when one agent transfers control to a peer?

level: middleimportance: must knowfreq 68%

basics

~20 s

The active agent swaps. The transfer is exposed to the model as an ordinary tool call, and once it fires, the next turn runs under the peer's system prompt, tool set and model — on the same live conversation, with no third party in between.

open as a page

In an orchestrator-worker agent system, what stays in the lead agent's context?

level: middleimportance: must knowfreq 75%

basics

~20 s

The lead holds the goal, its decomposition into subtasks, and the short results workers hand back. Each worker's own reasoning, tool calls and raw source material stay inside that worker's separate window and never reach the lead.

open as a page

How does a supervisor agent decide which specialist runs next each turn?

level: middleimportance: must knowfreq 62%

basics

~20 s

A supervisor is an LLM prompted with the roster of specialists, a one-line scope for each, and the conversation so far. Every turn it emits one constrained choice - the next specialist's name, or a finish signal - and the runtime dispatches accordingly.

open as a page

Why do multi-agent systems keep shared state in an artifact store plus an append-only event log?

level: middleimportance: must knowfreq 66%

basics

~20 s

An artifact store plus an append-only event log gives every agent one durable, inspectable source of truth that lives outside any context window. State then survives restarts, and no agent's view depends on which messages it happened to see.

open as a page

When does splitting work across specialized agent roles actually pay off?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Specialization pays when the subtasks are independent and mostly read-and-verify, so each role investigates its own slice and returns a separate finding. It stops paying when several roles must converge on the same mutable artifact.

open as a page

How do you catch semantic misalignment when two agents both report success?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Self-reports cannot catch it, because each agent met its own reading of the task. Catch it by verifying the produced artifact against the original request with an independent checker, and by making handoffs typed rather than prose.

open as a page

How do you attribute one failed multi-agent run to the specific agent that caused it?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Replay the recorded trace to find the first step where state diverged from a correct run, not the last step that errored. Then separate the agent that produced the bad output from the one that should have caught it, and confirm by counterfactual replay.

open as a page

How do you detect and break a handoff loop between two agents in a swarm?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Log every transfer as a structured event, then enforce a hop budget and a repeated-pair check on that log. When the budget is spent or the same two agents trade the conversation twice, stop transferring and escalate to a human or a generalist rather than letting the cycle continue.

open as a page

Eight parallel research workers return overlapping, uneven briefs — what do you fix first?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Fix the worker contract before anything else. Each worker needs an explicit objective, an exclusive slice of the sources, an allowed tool set, and a required output shape with a length cap. Overlap and unevenness are almost always a lead that copied the goal instead of partitioning it.

open as a page

Why should agents keep private scratchpads and promote only verified results to shared state?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Anything written to shared state is read by other agents as established fact, so an unverified guess propagates as truth. Keeping exploration in a private scratchpad and promoting only checked, compressed results puts a verification gate between reasoning and everyone else's input.

open as a page

What role does a triage agent play in a handoff-based support swarm?

level: juniorimportance: should knowfreq 48%

basics

~20 s

A triage agent is the swarm's front door. It holds the opening turns, works out what the customer actually needs, and then transfers the conversation to the specialist peer that owns that need instead of answering itself.

open as a page

In a multi-agent system, why treat the context window as RAM and shared storage as disk?

level: juniorimportance: should knowfreq 52%

basics

~20 s

The context window is small, volatile working memory; shared storage is large and durable. Agents keep references — paths, ids, ranges — in context and pull only the bytes a step needs, so total state size stops being limited by the window.

open as a page

Why package an agent role as a skill file that loads only when the role activates?

level: middleimportance: should knowfreq 42%

basics

~20 s

Packaging a role as a file — short metadata always visible, full instructions loaded only on activation — keeps the baseline context small while making the role portable, versionable and reviewable like code instead of buried in one giant system prompt.

open as a page

In A2A, what is an Agent Card and what does it let a calling agent do?

level: middleimportance: should knowfreq 45%

basics

~20 s

An Agent Card is a machine-readable description a remote agent publishes about itself: identity, the skills it offers, its endpoint, and how to authenticate. A calling agent fetches cards to discover peers and decide who can do a job, before sending any work.

open as a page

What must a multi-agent trace record for failure analysis to be possible later?

level: middleimportance: should knowfreq 45%

basics

~20 s

One correlation id spanning the whole task, one span per agent turn nested under the orchestrator, and verbatim inputs, outputs, handoff payloads, tool calls, model identity, token counts and latency. Runs are nondeterministic, so the trace is the only faithful record.

open as a page

Why give a critic agent a fresh context instead of the author agent's thread?

level: middleimportance: should knowfreq 45%

basics

~20 s

A critic started in a fresh context judges the artifact rather than the story that produced it. Continuing the author's thread anchors the critic on the author's assumptions and inherits a long, degraded window, so it tends to ratify work instead of finding defects.

open as a page

Two of eight parallel worker agents fail — should the orchestrator synthesize, retry, or abort?

level: middleimportance: should knowfreq 44%

basics

~20 s

It depends on whether the missing slices are load-bearing. Retry the two once with a bounded budget; if they still fail, synthesize from the six and state the gap explicitly, and abort only when the task is invalid without complete coverage. Never silently answer from partial data.

open as a page

In a supervisor system, how does agent-as-tool differ from agent-as-graph-node?

level: middleimportance: should knowfreq 44%

basics

~20 s

Agent-as-tool exposes each specialist as a callable the supervisor's model invokes, so control always returns and the reply lands in the supervisor's context as a tool result. Agent-as-graph-node makes each specialist a node in an explicit state machine, with edges deciding where control goes and shared state carrying the work.

open as a page

In A2A, how does a caller track a remote task that runs for minutes?

level: seniorimportance: should knowfreq 38%

basics

~20 s

A2A makes the unit of work a task with an id and a lifecycle, not a single request/response. The caller subscribes to streamed status and artifact updates, or registers a webhook for push notifications, and can re-attach to the task by id after a disconnect.

open as a page

A subagent dies mid-task; how does the orchestrator resume without redoing work?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Persist completed steps and their outputs to a durable store outside the agent's process, keyed by step. On restart, skip steps already recorded and re-run only the rest — which requires every step's side effects to be idempotent, since a step may have applied before it died.

open as a page

How should an orchestrator handle a subagent that stalls and never returns?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Termination belongs to the orchestrator, not the worker. Dispatch every subtask with a deadline, cancel on expiry, and degrade gracefully: return the completed results with the stalled scope explicitly marked unresolved rather than blocking or failing the whole run.

open as a page

When does ensembling independent LLM samples stop improving accuracy?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Aggregation only cancels errors that are independent. Samples drawn from one model on one prompt fail in the same direction, so five runs agreeing means the model is consistent, not correct — and more samples then only sharpen a biased estimate.

open as a page

How do you design a judge agent that picks between two agents' answers?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Give the judge the rubric and the underlying evidence, not just the two arguments, so it checks claims instead of rating rhetoric. Constrain its output to a verdict plus cited support, blind and shuffle the candidates, and route low-confidence cases to a human.

open as a page

At an agent handoff, should the receiving agent get the whole conversation history?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Usually not the raw whole. Pass a structured handoff summary — reason, verified identity, established facts, what was already tried — plus the last few verbatim turns. Full transcripts cost tokens and invite the peer to re-litigate work the previous agent finished.

open as a page

showing 1–30 of 41