skip to content

In an agent red-team harness, why is the adversarial instruction placed in the result a tool returns rather than in the user's message, and what does a hit at each of those two boundaries actually prove?

level: juniorimportance: must knowfreq 72%

answer

  1. user turn = request, tool result = data
  2. the agent trusts what it fetched
  3. which boundary did you inject at
  4. fix differs: model vs scaffold
  5. log what the model received

basics

~20 s

Put it in the tool result, not the user message. The agent is already trained to be wary of user requests, but treats tool output as data it fetched. The harness intercepts the call and returns attacker-controlled content in its place. A hit there shows the agent obeys untrusted retrieved content.

solid answer

~50 s

The two boundaries carry different trust. Content in the user turn is a **direct** request: a violation there says the model's refusal behaviour is weak, which is a model-safety result. Content inside a tool result is content the agent did not ask a human for -- the output of a search, a file read, a page load -- and the agent treats it as ground truth to act on. A violation there says the agent takes instructions from data, which is the failure that matters once it has real tools. Suites of the InjecAgent class exist to deliver at that second boundary, so the harness must sit between the agent and the tool and substitute the returned content. Your finding must record *which* boundary you injected at: the fix differs -- a refusal gap versus a data/instruction separation gap in the scaffold.

go deeper

for a junior

Should say the payload goes in the tool result because the agent treats fetched data as trustworthy, and that the harness has to intercept the call to put it there.

for a middle

Should contrast what each hit proves -- refusal weakness versus data/instruction confusion -- and name where the fix would live for each.

for a senior

Should insist the report names the injection boundary and the interception point, and should run the user-turn variant as a control so the tool-boundary result is not just restated refusal weakness.

for a principal

Frames it as which class of defence the organisation is buying: model-level refusal versus scaffold-level trust separation, and which findings actually change the agent's architecture.

### The loop, and the two slots an attacker can reach An agent run is an alternating sequence of messages. The model emits a turn; if that turn contains a tool call, the scaffold (the framework code that drives the loop: an agent-executor class, a graph runner, or a function-calling loop you wrote yourself) executes the call and appends the returned value to the message list as a **tool-result message**; then the model runs again with that message in its context. Two slots in that list can carry text an attacker chose. The **user turn** is what a human types. The **tool-result turn** is whatever the called function returned -- a search snippet, a file's contents, an email body, an HTTP response. Post-training teaches a model to treat the user turn as a *request*, something that can be evaluated and refused. Nothing comparable exists for the tool-result turn. In a plain chat template it is just another message with a role label; it carries no provenance, no author, and no trust marker, and it arrived because the model itself asked for it. That asymmetry is the entire reason a tool-boundary test exists. ### Mechanism: how content gets into a tool result You cannot type into that slot, so the harness must sit in the loop. Three placements do it: wrap the callable the framework's tool dispatch invokes and return your object instead of the real one; run a proxy in front of the tool's HTTP endpoint and answer the request yourself; or plant the content in the data a genuine tool will fetch (a document, a page, a record, a mailbox). Benchmark suites of the InjecAgent class ship the task, the tool inventory and the attacker content as fixtures, plus a label saying which behaviour counts as a hit -- but the substitution itself is still your harness's job, and the suite's numbers are only as honest as that plumbing. ### What a hit at each boundary licenses you to claim | injected at | what a violation shows | where the fix lives | |---|---|---| | user turn | the model complied with a direct request | model choice, system prompt, an input classifier | | tool result | the agent cannot separate fetched data from instructions | the scaffold: marking tool output untrusted, gating which calls are reachable after a fetch, confirmation before consequential actions | The second is the failure that matters once the agent has real tools, because content in a tool result can be planted by a third party who never speaks to the user at all. ### What it costs One attempt is at least two model calls -- the turn that emits the call and the turn that consumes the result -- and a realistic multi-step task runs five to fifteen. A 200-case suite at ~8 turns is roughly 1,600 model calls per pass; agent runs are high variance, so you want three to five samples per case, which puts a defensible pass in the thousands-to-ten-thousand call range, tens of dollars on a frontier model with long contexts, and tens of minutes to hours of wall clock. The dominant cost is neither: building a working interception point for one framework at one version is typically one to two engineer-days, and it breaks whenever the tool-dispatch path changes. ### Where the number misleads The headline "N% attack success at the tool boundary" fails in three specific ways. **One:** the payload never actually went through the tool boundary. If it was typed into the user turn and written up as indirect exposure, the rate describes refusal behaviour under a different label, and the scaffold fix the team ships will not close it. **Two:** no user-turn control was run. If the same text also succeeds as a direct request, the tool-boundary number is contaminated by cases that succeed through any channel; the interesting quantity is the *excess* over the direct-request rate, not the raw rate. **Three:** the harness manufactured it -- a shim that always returns success, returns instantly, or delivers on the first call puts the agent in a state the real environment gates, and the transcript of that looks identical to a real breach. ### What you would check Dump the rendered context for the turn after the intercepted call and confirm your text is present and intact -- the shim's own log proves only that you handed the text over, not that the model read it. Run the same payload in the user turn as a control and report both rates. Run the shim with benign content of the same shape and length to see what the instrumentation alone produces. And record the placement in the finding, because the first question any owning team asks about a tool-boundary hit is how much of their stack you replaced.

  • The same payload succeeds from the user turn and from the tool result. What do you report?
    Two findings, or one finding noting the agent refuses nothing on this behaviour. The tool-boundary result is not independently interesting until the direct request is refused -- otherwise you have measured general refusal weakness, not channel trust.
  • Your harness intercepts the tool call and returns the payload. Name one thing the report must state about the harness itself.
    Where the substitution happened -- inside the framework's tool dispatch, at a proxy in front of the endpoint, or in the data a real tool read -- because that determines whether the agent's own result post-processing was exercised.
  • Why can a tool-boundary hit be more serious than a user-turn hit even when the same text triggers both?
    Because content in a tool result can be planted by a third party who never talks to the user -- a web page, a shared document, an inbound email -- so it does not require a malicious operator.

A user request is a stranger at the door asking you to do something; a tool result is the answer to a phone call you placed yourself. People screen the first and trust the second, which is exactly why an attacker wants to be on the other end of the call you made.

saying these in an interview costs you the question

  • Describing a payload typed into the user message as an indirect or tool-boundary finding.
  • Cannot say what the harness had to intercept to get content into a tool result.
  • Claims a hit at either boundary proves the same weakness.
  • Never verifies that the injected text actually appeared in the model's context.

context