skip to content

Intercepting Tool Calls

Suites like InjecAgent place content in the tool result rather than the user turn, and the shim doing it can manufacture hits of its own. Interviewers ask which boundary you actually attacked.

on this pageshow

explore

questions

5

In an agent red-team harness, why is the adversarial instruction placed in the result a tool returns rather than in the user's message, and what does a hit at each of those two boundaries actually prove?

level: juniorimportance: must knowfreq 72%

answer

  1. user turn = request, tool result = data
  2. the agent trusts what it fetched
  3. which boundary did you inject at
  4. fix differs: model vs scaffold
  5. log what the model received

basics

~20 s

Put it in the tool result, not the user message. The agent is already trained to be wary of user requests, but treats tool output as data it fetched. The harness intercepts the call and returns attacker-controlled content in its place. A hit there shows the agent obeys untrusted retrieved content.

solid answer

~50 s

The two boundaries carry different trust. Content in the user turn is a **direct** request: a violation there says the model's refusal behaviour is weak, which is a model-safety result. Content inside a tool result is content the agent did not ask a human for -- the output of a search, a file read, a page load -- and the agent treats it as ground truth to act on. A violation there says the agent takes instructions from data, which is the failure that matters once it has real tools. Suites of the InjecAgent class exist to deliver at that second boundary, so the harness must sit between the agent and the tool and substitute the returned content. Your finding must record *which* boundary you injected at: the fix differs -- a refusal gap versus a data/instruction separation gap in the scaffold.

go deeper

for a junior

Should say the payload goes in the tool result because the agent treats fetched data as trustworthy, and that the harness has to intercept the call to put it there.

for a middle

Should contrast what each hit proves -- refusal weakness versus data/instruction confusion -- and name where the fix would live for each.

for a senior

Should insist the report names the injection boundary and the interception point, and should run the user-turn variant as a control so the tool-boundary result is not just restated refusal weakness.

for a principal

Frames it as which class of defence the organisation is buying: model-level refusal versus scaffold-level trust separation, and which findings actually change the agent's architecture.

### The loop, and the two slots an attacker can reach An agent run is an alternating sequence of messages. The model emits a turn; if that turn contains a tool call, the scaffold (the framework code that drives the loop: an agent-executor class, a graph runner, or a function-calling loop you wrote yourself) executes the call and appends the returned value to the message list as a **tool-result message**; then the model runs again with that message in its context. Two slots in that list can carry text an attacker chose. The **user turn** is what a human types. The **tool-result turn** is whatever the called function returned -- a search snippet, a file's contents, an email body, an HTTP response. Post-training teaches a model to treat the user turn as a *request*, something that can be evaluated and refused. Nothing comparable exists for the tool-result turn. In a plain chat template it is just another message with a role label; it carries no provenance, no author, and no trust marker, and it arrived because the model itself asked for it. That asymmetry is the entire reason a tool-boundary test exists. ### Mechanism: how content gets into a tool result You cannot type into that slot, so the harness must sit in the loop. Three placements do it: wrap the callable the framework's tool dispatch invokes and return your object instead of the real one; run a proxy in front of the tool's HTTP endpoint and answer the request yourself; or plant the content in the data a genuine tool will fetch (a document, a page, a record, a mailbox). Benchmark suites of the InjecAgent class ship the task, the tool inventory and the attacker content as fixtures, plus a label saying which behaviour counts as a hit -- but the substitution itself is still your harness's job, and the suite's numbers are only as honest as that plumbing. ### What a hit at each boundary licenses you to claim | injected at | what a violation shows | where the fix lives | |---|---|---| | user turn | the model complied with a direct request | model choice, system prompt, an input classifier | | tool result | the agent cannot separate fetched data from instructions | the scaffold: marking tool output untrusted, gating which calls are reachable after a fetch, confirmation before consequential actions | The second is the failure that matters once the agent has real tools, because content in a tool result can be planted by a third party who never speaks to the user at all. ### What it costs One attempt is at least two model calls -- the turn that emits the call and the turn that consumes the result -- and a realistic multi-step task runs five to fifteen. A 200-case suite at ~8 turns is roughly 1,600 model calls per pass; agent runs are high variance, so you want three to five samples per case, which puts a defensible pass in the thousands-to-ten-thousand call range, tens of dollars on a frontier model with long contexts, and tens of minutes to hours of wall clock. The dominant cost is neither: building a working interception point for one framework at one version is typically one to two engineer-days, and it breaks whenever the tool-dispatch path changes. ### Where the number misleads The headline "N% attack success at the tool boundary" fails in three specific ways. **One:** the payload never actually went through the tool boundary. If it was typed into the user turn and written up as indirect exposure, the rate describes refusal behaviour under a different label, and the scaffold fix the team ships will not close it. **Two:** no user-turn control was run. If the same text also succeeds as a direct request, the tool-boundary number is contaminated by cases that succeed through any channel; the interesting quantity is the *excess* over the direct-request rate, not the raw rate. **Three:** the harness manufactured it -- a shim that always returns success, returns instantly, or delivers on the first call puts the agent in a state the real environment gates, and the transcript of that looks identical to a real breach. ### What you would check Dump the rendered context for the turn after the intercepted call and confirm your text is present and intact -- the shim's own log proves only that you handed the text over, not that the model read it. Run the same payload in the user turn as a control and report both rates. Run the shim with benign content of the same shape and length to see what the instrumentation alone produces. And record the placement in the finding, because the first question any owning team asks about a tool-boundary hit is how much of their stack you replaced.

  • The same payload succeeds from the user turn and from the tool result. What do you report?
    Two findings, or one finding noting the agent refuses nothing on this behaviour. The tool-boundary result is not independently interesting until the direct request is refused -- otherwise you have measured general refusal weakness, not channel trust.
  • Your harness intercepts the tool call and returns the payload. Name one thing the report must state about the harness itself.
    Where the substitution happened -- inside the framework's tool dispatch, at a proxy in front of the endpoint, or in the data a real tool read -- because that determines whether the agent's own result post-processing was exercised.
  • Why can a tool-boundary hit be more serious than a user-turn hit even when the same text triggers both?
    Because content in a tool result can be planted by a third party who never talks to the user -- a web page, a shared document, an inbound email -- so it does not require a malicious operator.

A user request is a stranger at the door asking you to do something; a tool result is the answer to a phone call you placed yourself. People screen the first and trust the second, which is exactly why an attacker wants to be on the other end of the call you made.

saying these in an interview costs you the question

  • Describing a payload typed into the user message as an indirect or tool-boundary finding.
  • Cannot say what the harness had to intercept to get content into a tool result.
  • Claims a hit at either boundary proves the same weakness.
  • Never verifies that the injected text actually appeared in the model's context.

context

open as a page

You need adversarial content to arrive inside a tool result for an agent you are testing. Compare wrapping the tool function inside the agent framework, proxying the call in front of the tool endpoint, and planting the content in data a real tool reads -- what does each let you observe and control?

level: middleimportance: must knowfreq 55%

basics

~20 s

Three placements. Wrapping the tool function inside the framework is easiest and sees parsed arguments, but skips the wire. A proxy in front of the endpoint sees real requests and responses but not framework post-processing. Planting content in data a genuine tool reads is highest fidelity and lowest control over timing.

open as a page

An agent red-team run that substitutes tool results through a shim reports a large number of policy violations. Before you write any of them up, how do you establish that the shim did not manufacture them?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Re-run with the shim in place but a benign result substituted. If the behaviour still appears, the shim caused it, not the payload. Usual manufacturers: the shim answering a call that would have errored, returning instantly, changing call order, or replacing a result the framework would have truncated.

open as a page

Your harness substitutes adversarial content into an agent's tool results and the agent never misbehaves. Before reporting the agent as resistant, how do you verify the content actually reached the model?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Log what the model actually received, not what you handed the shim. Frameworks truncate, summarise, re-serialise or drop tool results before they become context, so your text may never arrive intact. Without that check a clean run is not evidence of resistance; it may only prove your content was cut.

open as a page

You are leading a red-team engagement against a team's live agent and you must intercept its tool calls to test the boundary the agent trusts. How do you decide how invasive that interception may be, and how do you keep the findings from being dismissed as 'your harness did that'?

level: principalimportance: should knowfreq 34%

basics

~20 s

Agree up front how much of the stack you may modify, and write it down. Every layer you replace is a layer the owning team can blame. Prefer the least invasive interception that still proves the claim, ship a matched control run with each finding, and state exactly what the harness changed.

open as a page