In a packaged agent tool world such as AgentDojo — a benchmark that ships mock tools plus tasks — where does the attacker's text actually enter the episode, and why is that different from typing it into the user's prompt?
answer
- payload rides in tool output
- benign user task, hostile fetched data
- data vs instruction boundary
- clean and poisoned replay of one fixture
- channels bounded by the mocks
basics
~20 sIt enters through tool output. The mock tools return data — an email body, a document, a transaction note — and the attacker text sits inside that returned content. The user's request stays benign, so the episode tests whether the agent obeys instructions it read while working, not instructions its user typed.
solid answer
~50 sA packaged tool world is a fixture: mock services (a mailbox, a payments app, a workspace) plus a set of benign user tasks. The attacker content is planted **in what a tool returns**, not in the user turn. That placement is the whole point of the instrument. A prompt-typed instruction tests whether the model refuses a hostile *user*; a tool-returned instruction tests whether the model can tell *data it fetched* from *instructions it was given* — the trust boundary an agent crosses on every tool call. The practical consequence is that the same fixture can be replayed twice: once with the clean mock content, once with the same content carrying a planted instruction. The user task, the tool set and the world state are identical, so a behavioural difference between the two replays has one candidate cause. That is what you buy from a packaged world, and it is why the mocks' realism — what they return and what actions they expose — bounds everything you can learn.
go deeper
Should say the attacker text comes back inside a tool result rather than from the user, and that the user's task is a normal one.
Adds why the placement matters — the model cannot separate fetched data from given instructions — and that the same fixture replays clean and poisoned.
Points out that channel count is a property of the mocks, insists reports name the channel exercised, and flags public payloads being pre-known to filters.
Frames tool-output trust as an architectural property of the product, not a benchmark artifact, and decides what the organisation reports from fixtures like this.
### What a packaged tool world is made of Three separable parts, and the vocabulary matters because interviewers probe it. 1. **An environment** — a small serialisable state object (a mailbox of messages, an account with a transaction list, a file drive, a travel-booking store) plus *mock tools*: ordinary functions that read and write that state. No network, no real service, no side effects outside the object. 2. **User tasks** — benign goals ("summarise my unread mail", "pay the invoice my landlord sent"), each paired with a programmatic checker that inspects the end state to decide whether the agent did the job. 3. **Injection tasks** — a *separate* goal the attacker wants ("send the balance to an outside address"), each with its own checker over that same end state. AgentDojo's contribution is the wiring between them. Its seeded environment data contains named placeholder slots sitting inside ordinary-looking content — a message body, a file, a memo line on a record — and the harness substitutes the attacker's string into those slots before the episode starts. Then it runs the *benign* user task. The agent calls a mock tool because its own plan needs the data; the mock returns the seeded record; and the planted instruction arrives as the **return value of a call the agent chose to make**. ### Why that channel, and not the user turn By the time the tool result reaches the model it is just more tokens in the same context window as the system prompt and the user's request. There is no field, no colour, no privilege bit that says "this is data, not instruction". So the fixture measures a different property from hostile-user prompting: not *will the model refuse a person who asks for something bad*, but *can the model keep fetched content on the data side of the boundary while acting on the user's behalf*. Refusal training governs the first; almost nothing in a typical stack governs the second. A product can score cleanly on one and badly on the other, which is why the two must never be collapsed into a single "safety" percentage. The second thing the placement buys is the paired replay: run the same user task twice, once against clean seed data and once against the same data carrying the planted slot. Same task text, same tools, same world state, one variable different. ### What a sweep costs Nothing per tool call — the mocks are local functions — but everything per *episode*. An episode is a multi-turn agent loop: plan, call, read result, call again, answer. Five to fifteen model calls is typical, each carrying the accumulated transcript, so tokens per episode grow super-linearly with steps. A suite in the hundreds of user-task/injection-task pairings is therefore thousands of model calls and hours of wall clock even at healthy concurrency, and hosted rate limits usually bind before the budget does. The mocks are free; the agent loop is not. ### Where the number misleads The headline output is an attack success rate. Three readings of it are wrong. - **Zero does not mean immune.** The denominator counts pairings the *fixture* can express. Coverage of this channel is a property of the mocks: if only three tools return free text, you have three entry points no matter how many payload variants you write. Payload variety and channel variety are different axes and only one of them is under your control. - **Zero can mean nothing was delivered.** In many pairings the user task never calls the tool holding the planted content, so the agent never saw it. Those episodes score as attack failures and drag the average down. Report success conditioned on delivery, and log per episode whether the poisoned bytes ever entered the context. - **Zero can mean recognition.** Public fixtures circulate; their strings reach vendor filters and training corpora. A clean pass may record familiarity with a well-known suite rather than robustness of your product. ### What to check before believing it Dump the fully rendered tool result for one episode and confirm the payload is where you think it is, in the container you think it is in. Run a deliberately obedient stub agent and confirm every injection checker actually fires — a checker that can never fire produces exactly the 0% you were hoping to see. Confirm the clean pass still satisfies the user-task checker, so you know the agent was capable of the benign job at all. And when you report, name the channel: "instruction planted in returned tool content, benign user task", never a bare percentage.
- Why does the user task in these fixtures stay benign?So the only hostile element is the planted tool content. If the user request were also hostile you could not tell which one moved the agent, and you would be measuring refusal of a user rather than confusion of data with instructions.
- If the fixture's mocks expose only two tools that return free text, what does that cap?The number of distinct entry channels you can exercise, regardless of how many payloads you write. Payload variety and channel variety are separate axes, and the mocks fix the second one.
- Name a production tool surface a packaged mailbox mock does not represent.Anything with paginated, truncated or multi-format results — long threads, attachments, quoted HTML, other tenants' shared documents — plus tools that fail with auth or permission errors partway through.
saying these in an interview costs you the question
- Describes the fixture as sending jailbreak prompts as the user — that is a different instrument.
- Assumes the harness marks tool output as untrusted for the model; nothing in the context does.
- Reports one success number without saying whether the instruction arrived via user turn or tool result.
- Treats a pass as evidence about hostile users, or vice versa.