How do you detect and break a handoff loop between two agents in a swarm?
answer
- nobody holds the route
- each hop is locally defensible
- count hops, not intentions
- withdraw the tool, do not ask nicely
- a trip is a bug report about scope
basics
~20 sLog every transfer as a structured event, then enforce a hop budget and a repeated-pair check on that log. When the budget is spent or the same two agents trade the conversation twice, stop transferring and escalate to a human or a generalist rather than letting the cycle continue.
solid answer
~50 sPing-pong is the signature failure of decentralized handoff: billing decides the case is a retention matter, retention decides it is a billing matter, and the conversation bounces four times while the customer waits. No agent is wrong locally — each correctly judged the request outside its own scope — and because there is no supervisor, nobody sees the cycle. The fix has three parts. **Detect**: emit a transfer event with source, destination, reason and hop index, and check it for a repeated agent pair or a repeated (agent, intent) state. **Bound**: cap hops per conversation, typically three to five, and make the cap enforced by the runtime rather than requested in a prompt. **Escape**: on trip, route to a human queue or a generalist that is not allowed to transfer, so the terminal state is an answer rather than another hop. Then treat every trip as a bug report: repeated ping-pong between the same pair means their scope boundaries overlap or a category has no owner at all.
code
python · 17 linesdef detect_ping_pong(handoffs, window=4):
"""handoffs: ordered list of (from_agent, to_agent) transfer events."""
if len(handoffs) < window:
return False
recent = handoffs[-window:]
pair = frozenset(recent[0])
return all(frozenset(hop) == pair for hop in recent)
log = [
("triage", "billing"),
("billing", "retention"),
("retention", "billing"),
("billing", "retention"),
("retention", "billing"),
]
print(detect_ping_pong(log)) # True: billing and retention are trading itgo deeper
Know what handoff ping-pong is — two agents passing the same conversation back and forth — and that a simple hop counter is the basic defence.
Explain why the topology invites it: each agent decides with only local information and no component sees the whole route. Describe repeated-pair detection and a runtime-enforced hop budget with an escape destination.
Demonstrate operating depth — structured transfer events, cycle and no-progress detection, capability-level enforcement rather than prompt pleading, and dashboards that surface the worst ping-pong pair as a prompt-boundary bug to fix.
Own the systemic view: recurring cycles mean the agent scope decomposition is wrong or a request category has no owner. Decide the hop budget from measured distributions, and judge when the cycle rate justifies moving that flow to a centralized router.
## What the loop looks like In a mobile-carrier support swarm, a customer says their bill jumped after they threatened to cancel. Billing reads "threatened to cancel" and transfers to retention. Retention reads "bill jumped" and transfers to billing. Each transfer is locally defensible; together they are a cycle. Left unbounded it repeats until the user leaves, and it burns a full model call plus the accumulated context on every hop. The deeper variant is a longer cycle — triage to billing to retention to triage — which a naive "same pair twice" check misses. And the subtlest variant is a loop with *progress-shaped noise*: each hop adds a sentence to the summary, so the state is never byte-identical even though nothing is actually advancing. ## Why the topology invites it A supervisor topology has one component that sees every routing decision in sequence, so a cycle is visible to whoever makes the decisions. A swarm deliberately removes that component. Each agent decides with only local information: its own scope, its own prompt, and the conversation as it received it. Nobody holds the route. Loop safety therefore cannot be an emergent property of good prompts — it has to be a runtime property. This is worth saying plainly in an interview, because it reframes the answer from "write better instructions" to "the topology gave up global knowledge, so I add back a global counter". ## Detection signals, cheapest first **Hop count.** A monotonic counter per conversation. Crude, universal, and it catches everything eventually. Set the budget from data: measure the hop distribution of successful conversations and cap a little above the tail. **Repeated pair.** Look at the last N transfers; if they all involve the same unordered pair of agents, it is ping-pong. Cheap, exact, catches the common two-agent case immediately, and does not need to wait for the hop budget to drain. **Cycle in the hop sequence.** Treat the transfer log as a walk over agents and look for a repeated agent within the conversation, or a repeated (agent, open-question) pair. Catches three- and four-node cycles. **No-progress check.** Compare the handoff record across hops: if verified facts, actions taken and the open question are unchanged, the conversation is not advancing regardless of how the summary prose was reworded. This is the one that catches disguised loops. **Cost and wall clock.** A per-conversation token and time budget is a backstop that catches loop shapes you did not anticipate. ## Breaking the loop Enforcement belongs in the runtime, not the prompt. Once the budget is exceeded, the transfer tool should stop being offered to the model at all, or should return an error result telling the agent it must resolve or escalate. An agent that is merely *asked* not to transfer will transfer under pressure. The escape route matters as much as the trip. Options, roughly in order of preference: - **human escalation** — the honest terminal state for a conversation the swarm could not place; - **a generalist agent with no transfer tools** — it must answer with what it has; - **return to the original triage agent with the loop history attached** and a strict instruction to choose a destination it has not tried; - **a supervisor invoked as a one-off tiebreaker** — a hybrid that keeps the swarm's cheap common path and pays for arbitration only in the rare cycle. Do not silently drop the conversation, and do not just re-prompt the same agent with "try harder" — that spends tokens and usually reproduces the same decision. ## Treat every trip as a defect A hop cap is a seatbelt, not a fix. When two agents ping-pong repeatedly across many conversations, one of three things is true: their scope descriptions overlap, so both plausibly own the request; a request category has *no* owner, so everyone forwards it; or the handoff record drops the information the receiving agent needed to see the request as its own. Dashboards should show the ping-pong pair with the highest frequency, because that pair names the prompt boundary to fix this week. This is also where the field's failure taxonomy is useful: published multi-agent failure analyses attribute the largest share of failures to specification and system-design issues, with inter-agent misalignment close behind, and only a minority to weak verification. Handoff cycles sit squarely in those first two clusters — they are design defects that show up at runtime, not model stupidity. ## Bounding is not the same as solving A capped loop still cost the user four hops of latency and you four model calls. The goal is a hop distribution where the great majority of conversations finish within one or two transfers, with the cap firing rarely enough that each trip is worth reading individually.
- Why is a hop cap enforced in the runtime better than an instruction in the prompt?Because an instruction is a request the model can rationalize away, especially under a persuasive user turn or an ambiguous request. Withdrawing the transfer tools once the budget is spent makes the constraint unarguable: the model physically cannot route again and must answer or escalate. Capability control beats prompt control for every hard limit in agent systems.
- How would you catch a loop that never repeats an exact state?Compare the semantic content of the handoff record across hops rather than raw text: verified identity, confirmed facts, actions taken and the open question. If none of those changed while the summary prose did, no progress was made. A per-conversation token and wall-clock budget is a useful backstop for loop shapes you did not model.
- Repeated ping-pong between the same two agents keeps appearing in your dashboards. What is the underlying cause likely to be?Either overlapping scope — both agents' transfer descriptions plausibly claim the request — or an orphan category that neither owns, so both forward it. Occasionally it is a lossy handoff record hiding the fact that makes the request clearly one agent's. Fix the boundary in the prompts or add an owner for the orphan category; the cap only bounds the damage.
- What should happen to the user when the hop budget trips?They should get an explicit, honest transition — a short message that the request is going to a person or to a generalist who will handle it directly — not another silent transfer and not a dead end. The terminal state must produce an answer or a queued human, because at that point the swarm has already demonstrated it cannot place the request itself.
saying these in an interview costs you the question
- Relies on prompt wording alone to tell agents not to transfer back
- Only checks for an immediate two-agent bounce and misses longer cycles
- Treats hitting the hop cap as success rather than as a design defect to fix
- Escapes the loop by handing back to the same agent that just transferred
- Assumes each agent can see the whole route and will notice the cycle itself