skip to content

What does prompt chaining between two LLM agents cost in tokens and fidelity?

level: juniorimportance: must knowfreq 60%

answer

  1. the simplest possible agent-to-agent channel
  2. output text becomes the next prompt
  3. billed again on every turn
  4. errors arrive looking like facts
  5. condense to named fields instead

basics

~20 s

Prompt chaining pastes one agent's finished output into the next agent's prompt. It is the cheapest wiring to build, but the receiver pays input tokens for every word on every turn and inherits any error, hedge or ambiguity verbatim.

solid answer

~50 s

Prompt chaining is the simplest inter-agent channel: agent A produces text, and that text is dropped into agent B's prompt as context. Two costs follow. The **token cost** is structural — a 6k-token analysis pasted into B is 6k input tokens on B's *first* turn and on every subsequent turn of B's loop, because the transcript keeps growing; a multi-hop chain multiplies this. The **fidelity cost** is subtler: B has no way to distinguish A's verified findings from A's speculation, so a confident hallucination arrives looking exactly like a fact, and long pasted blocks degrade B's attention to its own instructions. The fix is not to stop chaining but to narrow the seam: have A emit a short, named set of fields or a compressed summary that B can check, rather than its whole working transcript.

code

python · 23 lines
python
def chain_raw(output_a: str) -> str:
    return (
        "Here is the analysis from the previous agent:\n"
        f"{output_a}\n\nNow draft the supplier reply."
    )


def chain_fields(result: dict) -> str:
    return (
        f"supplier={result['supplier']} "
        f"unit_price={result['unit_price']} "
        f"currency={result['currency']} "
        f"lead_days={result['lead_days']}\n\nNow draft the supplier reply."
    )


print(len(chain_raw("word " * 4800)))
print(len(chain_fields({
    "supplier": "Northwind Metals",
    "unit_price": 1240,
    "currency": "EUR",
    "lead_days": 14,
})))

go deeper

for a junior

Be able to say plainly that chaining means one agent's output text becomes part of the next agent's prompt, and that the receiver pays input tokens for all of it.

for a middle

Explain that the block is re-billed on every turn of the receiver's loop, and that free text carries no provenance — so the receiver cannot separate a verified finding from a guess.

for a senior

Show the production fix: subagents return a condensed result, artifacts move by reference, and fields carry where they came from. Be ready to estimate the token bill of a three-hop chain.

for a principal

Own the tradeoff between build speed and seam discipline. Argue when a chain should stay a string in a prompt and when it must become a versioned contract, and tie that decision to who releases each side.

## What prompt chaining is Prompt chaining is the most basic way two LLM agents talk: agent A runs, produces output text, and that text is inserted into agent B's prompt — usually as a block like *"Here is the analysis from the previous step: …"*. No protocol, no schema, no transport beyond a string concatenation in your own code. It is how almost every multi-agent system starts, and for short pipelines it is often the correct answer. It is worth naming what chaining actually transfers: **text, and nothing else**. The receiving agent does not get A's tool results, A's confidence, A's retries, or A's reasoning state. It gets the words A happened to emit, in a prompt slot that carries no more authority than any other sentence in the context. ## The token cost A language model is charged on input tokens for the whole context on **every** call. If agent A emits a 6,000-token report and you paste it into agent B, that 6,000 tokens is billed on B's first model call — and again on B's second call, and its third, because B's own transcript (its tool calls and results) is appended to a context that still contains the pasted block. An agent that takes ten turns to finish has paid for that block ten times. Chains compound this. In A → B → C, if each step forwards what it received plus what it produced, context grows super-linearly and the last agent in the chain is the most expensive one to run. This is a large part of why orchestrated multi-agent systems burn far more tokens than a single-agent chat for the same task — reported multiples of an order of magnitude are common for research-style pipelines. ## The fidelity cost The fidelity problem is the one candidates usually miss. - **Provenance is erased.** A's grounded citation and A's guess arrive in B's prompt as the same kind of text. B has no signal for which is which, so B will treat a fabricated supplier price with the same seriousness as a retrieved one. Errors do not get filtered by the hop; they get laundered by it. - **Attention is diluted.** Model quality degrades as the window fills — the common name for this is *context rot*. A large pasted block competes with B's own system prompt and task instructions, and the longer the block, the more likely B drifts from what it was actually asked to do. - **Ambiguity survives.** A's hedged sentence ("the lead time is probably around two weeks") becomes B's input. B must either re-derive the uncertainty or, more often, flatten it into a confident downstream claim. - **Format drift.** Because the channel is free text, A can change its output shape between runs — a heading disappears, a list becomes prose — and B's parsing or reasoning silently changes with it. There is no failing test at the boundary; there is only a worse answer. ## Narrowing the seam The standard remedy is to make the handoff *smaller and more structured*, not to abolish it: 1. **Have A emit a condensed result, not its transcript.** In production orchestrator/subagent designs, a subagent that did a great deal of work returns on the order of one to two thousand tokens of findings. The exploration stays in A's own context and dies with it. 2. **Give the payload named fields.** `supplier`, `unit_price`, `currency`, `lead_days` is both cheaper and checkable; the receiver can validate it before spending a model call on it. 3. **Pass references, not contents.** Hand over an identifier or a path to an artifact and let B fetch only what it needs. 4. **Mark provenance.** If a field came from a tool result rather than from the model's own inference, say so in the payload, so B can weight it. ## When plain chaining is right Do not over-engineer. If the pipeline is two steps, the intermediate output is a few hundred tokens, and both steps ship from the same codebase, a string in a prompt is the correct amount of machinery. The cost curve only bites when outputs are large, the chain is long, or the steps are owned by different teams — at which point the seam deserves a contract. ## What interviewers listen for A weak answer describes chaining as "just passing the output along" and stops. A strong answer prices it: input tokens billed per turn and re-billed as the loop runs, plus the loss of provenance that lets one agent's mistake become the next agent's premise — and then proposes the condensed, named-field handoff as the fix.

  • If the receiving agent runs a ten-turn tool loop, how many times is the pasted block paid for?
    Once per model call, so roughly ten times. Each turn re-sends the full context — the pasted block plus everything the agent has appended since — as input tokens. This is why a large handoff is far more expensive than the single paste suggests, and why condensing before the hop, rather than after, is what saves money.
  • How would you keep provenance across the hop without shipping the whole transcript?
    Emit a structured result where each field carries where it came from — a tool result, a retrieved document id, or the model's own inference. The receiver can then trust measured fields and treat inferred ones as claims needing verification. It costs a few tokens per field and prevents the receiver from promoting a guess into a premise.
  • When is chaining raw output actually the right choice?
    When the output is small, the chain is short, and both ends ship from the same codebase. Structuring a 200-token handoff between two steps you deploy together adds schema maintenance for no measurable gain. Reach for a contract when payloads grow, hops multiply, or the two ends stop being released at the same time.

Forwarding an entire email thread instead of writing the one-line summary the reader needs: everything is technically there, the reader pays to wade through it, and a wrong claim buried three replies down looks as authoritative as the rest.

saying these in an interview costs you the question

  • Assumes context is free because the model accepts it
  • Pastes the whole upstream transcript, then blames the model
  • Believes the receiver can tell fact from speculation in text
  • Treats an upstream hallucination as verified once it is quoted
  • Thinks chaining transfers the sender's reasoning state

context