At an agent handoff, should the receiving agent get the whole conversation history?
answer
- neither everything nor nothing
- the customer should not repeat themselves
- discarded hypotheses look like facts
- typed record beats improvised prose
- auth state from session, not from summary text
basics
~20 sUsually not the raw whole. Pass a structured handoff summary — reason, verified identity, established facts, what was already tried — plus the last few verbatim turns. Full transcripts cost tokens and invite the peer to re-litigate work the previous agent finished.
solid answer
~50 sBoth extremes fail. Passing nothing makes the customer repeat their account number and their problem, which is the single most visible defect of a badly built swarm. Passing a raw 40-turn transcript costs tokens on every subsequent turn, degrades attention over a long window, and lets the receiving agent pick up the previous agent's abandoned hypotheses as if they were facts. The workable middle is a **structured handoff record**: why the transfer happened, identity and authorization state already established, confirmed facts (account, plan, device), actions already taken and their outcomes, and a short verbatim tail of the last few turns so tone and the user's exact last request survive. Everything else stays in the session log, retrievable if the peer asks for it. The hard part is that summarization is lossy and the loss is silent. Whatever the summary omits, the receiving agent cannot know it omitted — so treat the handoff schema as a contract you test, not as prose the model improvises.
code
python · 20 linesdef build_handoff_record(session, target, reason, open_question):
"""Deterministic fields from session state; only prose comes from the model."""
return {
"target": target,
"reason": reason,
"open_question": open_question,
"identity_verified": session["auth"]["verified"], # runtime truth
"confirmed_facts": session["facts"], # account, plan, device
"actions_taken": session["actions"], # includes negatives
"recent_turns": session["turns"][-4:], # verbatim tail only
}
session = {
"auth": {"verified": True},
"facts": {"account": "A-1187", "plan": "unlimited"},
"actions": ["credit applied", "roaming charge ruled out"],
"turns": ["..."] * 40,
}
print(build_handoff_record(session, "device_repair", "looks like SIM fault", "is the SIM faulty?"))go deeper
Know that the receiving agent does not automatically know what happened earlier, and that a handoff needs to carry something — at minimum why the transfer happened and what was already established.
Explain both failure modes concretely: too little causes re-asks and repeated work, too much costs tokens, degrades long-context reliability and lets abandoned hypotheses look like facts. Describe the structured-record middle ground.
Show operating judgment — a typed handoff schema, deterministic fields from session state, a short verbatim tail, on-demand retrieval of older turns, and metrics like re-ask rate and repeat-action rate that tell you which field is missing.
Own the trade-off as a policy: decide where curation is mandatory (regulated, multi-hop, long conversations) versus where full history is acceptable, and treat model-authored handoff text as untrusted input in the same way tool output is treated.
## Why the boundary is a real decision When control passes from one peer to another, the receiving agent starts a turn with whatever context the runtime hands it. There is no shared working memory by default and no supervisor holding the thread. So "what does the next agent see?" is a design choice with user-visible consequences on both sides. Take a mobile-carrier support swarm. A customer spends 40 turns with billing: identity verified, three invoices examined, one credit applied, two hypotheses about a roaming charge tried and discarded. Billing now transfers to device repair, because the real cause looks like a faulty SIM. ## Failure mode one: too little If the repair agent starts with only the user's last message, it asks for the account number again. To the customer, the "single support experience" has just revealed itself as three bots in a trench coat. Worse, the repair agent repeats the diagnostics billing already ran, because it has no record that they were run. This is the failure users complain about, and it is the reason handoffs need an explicit payload rather than relying on the destination agent to ask. ## Failure mode two: too much Dumping the raw 40 turns has four distinct costs. **Token cost compounds.** The transcript is not paid once — it rides along on every subsequent turn of the conversation, and if a third transfer happens it grows again. Multi-agent orchestration already carries a large token premium over plain chat; unfiltered history multiplies it. **Context degradation.** Long windows measurably reduce a model's reliability at using what is in them — the effect commonly called context rot. The repair agent's actual job, diagnosing a SIM fault, competes for attention with 35 turns of invoice arithmetic. **Contamination.** The transcript contains billing's *discarded* reasoning. Models do not reliably distinguish "hypothesis the previous agent abandoned" from "established fact", so the repair agent may confidently continue chasing a roaming charge that was already ruled out. **Role bleed.** A verbatim history full of billing's persona, tone and policy statements pulls the repair agent toward behaving like a billing agent, undermining the specialization the swarm exists to provide. ## The structured handoff record Define a schema and populate it deterministically wherever you can: - **reason** — why this transfer, in the transferring agent's own words; - **identity/auth state** — verified or not, and to what level. Carry this as runtime state, never as a claim the summary makes, so the peer never re-verifies and never *skips* verification on a forged claim; - **confirmed facts** — account id, plan, device, ticket ids. Prefer values already in session state over values the model restates; - **actions taken and outcomes** — including negative results ("roaming charge ruled out"), which are the most valuable and most frequently dropped items; - **open question** — what the receiving agent is being asked to determine; - **verbatim tail** — the last two to four turns, so the user's exact phrasing and emotional register survive. Anything not in the record stays in the durable session log. Give the receiving agent a tool to fetch older turns on demand — progressive disclosure — so nothing is truly lost while the working context stays small. ## Who writes the summary Three options, in ascending trust. Let the transferring model write it as free text: cheapest, most lossy, and it will flatter its own work. Have it fill a **typed schema** as the transfer tool's arguments: better, because required fields force the omissions to be conscious ones. Or assemble it **programmatically** from session state, with the model contributing only the reason and the open question: most reliable, and the right choice for anything security- or money-adjacent. A related caution: the transfer reason is model-authored text that another model will read as instruction-shaped context. If any of it derives from user input, it is untrusted, and the same prompt-injection discipline you apply to tool results applies here — a delegated agent is a favourite injection target precisely because nobody is watching it as closely as the user-facing one. ## How you would know it is wrong Instrument it. Count re-asks — how often the receiving agent requests a fact that was already confirmed. Count repeat actions — the same diagnostic run twice in one conversation. Track post-handoff context size. Sample transcripts where the customer used the word "again". Each of these maps to a specific field missing from the record, which makes the schema improvable instead of a matter of taste. ## The honest summary There is no settled default. Systems with short conversations and cheap tokens often pass everything and are fine. Long, multi-hop, regulated conversations almost always need a curated record. State the trade-off, pick a default of "structured record plus short verbatim tail plus on-demand retrieval", and say which measurements would move you off it.
- What is the single most valuable field teams forget in a handoff record?Negative results — what was tried and ruled out. Summaries naturally record findings, not eliminations, so the receiving agent re-runs discarded hypotheses and the customer watches the same diagnostic twice. Making "actions taken and their outcomes" a required field, including failures, removes most repeated work in multi-hop conversations.
- How do you keep identity verification from being repeated or, worse, spoofed across a handoff?Carry auth state as runtime session state that the receiving agent reads directly, not as a sentence in a model-written summary. A summary claiming "identity verified" is text another model will trust, which is exactly the confused-deputy shape. The runtime should also re-check that the receiving agent's own permissions cover the action it is about to take.
- If you filter history, how does the receiving agent recover something the summary dropped?Give it a retrieval tool over the durable session log — fetch turns by index, by time range, or by keyword. That is progressive disclosure: the working context stays small while nothing is actually lost. It also converts a silent omission into a visible miss you can measure, because you can count how often the peer has to go looking.
- Does the user need to be told a transfer happened?Almost always yes, in one short bridging line before the new agent speaks. It sets expectations for a change in tone and scope, and it makes a re-ask feel like a small hiccup rather than evidence that nobody is listening. Silent transfers are a common source of "the bot forgot everything" complaints even when the record was passed correctly.
saying these in an interview costs you the question
- Always passes the full transcript because more context is safer
- Passes nothing and expects the receiving agent to just ask again
- Lets the model assert identity verification in free-text summary prose
- Records only findings and drops what was tried and ruled out
- Treats the summary as lossless once a model has written it