In an agent handoff, what actually changes when one agent transfers control to a peer?
answer
- routing reduced to a tool choice
- the baton, not a subroutine
- no return value, no supervisor turn
- new prompt, new tools, same conversation
- exactly one holder at a time
basics
~20 sThe active agent swaps. The transfer is exposed to the model as an ordinary tool call, and once it fires, the next turn runs under the peer's system prompt, tool set and model — on the same live conversation, with no third party in between.
solid answer
~50 sA handoff is modelled as a tool the agent can call, typically named after its destination (`transfer_to_billing`). Calling it does not return data to the caller the way a normal tool does; it swaps which agent is *active*. From the next turn on, the runtime uses the receiving agent's instructions, its allowed tools and possibly a different model, while the conversation and its state continue uninterrupted. Three consequences matter. First, control is **peer-to-peer**: no supervisor turn sits between the two agents, so there is no extra model call and no central place that sees the decision. Second, the transferring agent is **out** — it does not get control back unless someone transfers to it. Third, because the destination is decided by the transferring agent's own prompt, routing logic is distributed across every agent's instructions rather than centralized in one router prompt.
code
python · 11 linesdef apply_transfer(session, target_agent, reason):
"""Swap the active agent; the conversation itself continues."""
session["handoff_log"].append((session["active_agent"], target_agent, reason))
session["active_agent"] = target_agent # prompt + tools now come from target
session["hops"] += 1
return session
session = {"active_agent": "triage", "hops": 0, "handoff_log": []}
apply_transfer(session, "billing", "customer disputes a charge")
print(session["active_agent"], session["hops"])go deeper
Know that a handoff is a tool call that switches which agent is talking, and that the new agent brings its own instructions and tools while the conversation continues.
Explain the semantics precisely: no return value, no supervisor turn in between, exactly one holder at a time, and transfer-tool descriptions acting as the routing prompt. Contrast it with call-and-return delegation.
Talk about operating it — structured transfer events, hop budgets, bridging messages to the user, and the fact that routing logic scattered across prompts is hard to change and to test.
Own the consequence: distributed routing trades a central bottleneck for a governance problem. Be ready to argue for generating transfer descriptions from one routing spec, or for accepting a router where auditability of routing decisions is a compliance requirement.
## The mechanism In a handoff (or swarm) topology, agents are peers and each one carries transfer tools for the peers it may hand to. To the model, a transfer looks like any other tool: it has a name, a description explaining when to use it, and often a small argument payload. The difference is what the runtime does with the call. A normal tool call runs code and feeds a result back to the *same* agent. A transfer call instead marks a different agent as the holder of the conversation, and the loop continues under that agent's configuration. Concretely, after a transfer: - the **system prompt** in play is the receiving agent's; - the **tool list** offered to the model is the receiving agent's; - the **model and its settings** may change (a cheap triage model handing to a frontier specialist); - the **conversation** continues — the user is not restarted, and whatever state the runtime tracks (session id, verified identity, tickets opened) persists; - the **transferring agent is gone** from the loop until something transfers back to it. OpenAI's Swarm made this pattern well known as an experimental framework; it was deprecated in 2025 in favour of the Agents SDK, where handoffs are a first-class primitive. Microsoft's Agent Framework ships an explicit handoff orchestration. The primitive is the same everywhere, and the exam-worthy part is the semantics, not any one library's call. ## Why exposing it as a tool matters This is the elegant part of the design: routing needs no new plumbing. The model already knows how to choose among tools, so "decide who should handle this next" reduces to "choose the right tool", and the *description* of each transfer tool becomes the routing prompt. Write "use when the customer disputes a charge, asks about their invoice, or wants a refund" and you have specified routing without a router. It also means the usual tool-use failure modes apply to routing: with too many peers, selection accuracy degrades; overlapping tool descriptions cause misroutes; and a badly described destination is simply never chosen. ## Peer-to-peer versus deciding centrally The contrast that interviewers want is about *who decides* and *how often*. In a swarm, the agent currently holding the conversation decides, once, at the moment it recognizes the request is not its own. Between billing and retention there is no third model call and no arbiter — the cost of a routing decision is folded into a turn the agent was taking anyway. That buys latency and token efficiency, and it costs you a single place to reason about. Routing policy is now scattered across N agent prompts; changing "escalate cancellations to retention" means editing every prompt that could see a cancellation. Observability suffers similarly: there is no one component whose log answers "why did this conversation end up here?", so you must emit a structured event on every transfer — from, to, reason, hop number — or you will be reconstructing routes from raw traces. ## Ownership: the baton rule Exactly one agent owns the conversation at a time. This single-holder property is what makes swarms tractable, and it lines up with the discipline that has held across 2025–26 practice: the acting, writing path stays single-threaded, and extra agents contribute intelligence rather than concurrent actions. Two agents replying to the same user at once is not a swarm; it is a race condition with a chat interface. Handoff is therefore *not* delegation. Delegation implies the caller waits for a result and resumes — that is the orchestrator-worker shape, where a lead keeps ownership and subagents return summaries. A handoff has no return. If your design needs the first agent to come back and use the second's output, you want a call-and-return structure, not a transfer, and confusing the two is a common design error. ## What to carry across the boundary Because the receiving agent must continue the conversation cold, the transfer payload matters. At minimum it should carry the reason for the transfer; usually also the facts already established (identity verified, account located, steps already tried). Whether the peer sees the whole transcript or a filtered summary is a real design decision with real costs, and it is the first thing a good interviewer probes after you describe the mechanism. ## Failure modes to name - **Silent transfer.** The user sees no acknowledgement and the new agent opens with a non sequitur. Emit a short bridging message. - **Transfer as a black hole.** The transferring agent assumes it will regain control and its prompt promises "I'll confirm once billing is done". It never will. - **Unbounded onward transfers.** Every agent can transfer, so conversations can chain or cycle; hop budgets are required. - **Routing drift.** Two peers' transfer-tool descriptions overlap and the same request lands in different places on different runs.
- How is a handoff different from calling another agent as a tool?Agent-as-tool is call-and-return: the caller keeps ownership, blocks on a result, and continues its own turn with that result in context. A handoff has no return — control moves permanently to the peer and the caller drops out. Choose agent-as-tool when the first agent must synthesize the answer; choose a handoff when the peer should own the rest of the conversation.
- Where does routing policy live in a swarm, and why is that awkward to change?It lives in the descriptions of each agent's transfer tools, so it is distributed across every prompt. A policy change like "cancellation intents go to retention first" must be applied to every agent that could see a cancellation, and there is no single artifact to review or test. Teams usually respond by generating transfer-tool descriptions from one shared routing spec.
- What should you log at each transfer boundary?A structured event with source agent, destination, the model-stated reason, the hop index, elapsed time, and a session id. Without it you cannot answer "why did this conversation end up in retention?" after the fact, because no central router exists to have recorded the decision. These events are also what your loop detection and first-hop accuracy metrics are computed from.
- Can two agents hold the conversation at once in a swarm?No — the topology assumes exactly one holder, and that is what keeps it tractable. Concurrent writers on the same conversation produce interleaved replies and conflicting tool actions. Current practice keeps the acting path single-threaded and uses parallelism only for read-only or advisory work that funnels back to one holder.
saying these in an interview costs you the question
- Thinks the transferring agent waits for a result and then resumes
- Describes a supervisor model call happening between the two agents
- Assumes the receiving agent keeps the sender's system prompt and tools
- Believes handoff and agent-as-tool delegation are the same thing
- Says routing policy lives in one central place in a swarm