skip to content

Benchmark Environments

A packaged tool world replays one task with and without a payload and scores utility beside the attack, which a live system cannot do. Interviewers ask what its fixtures leave untested.

on this pageshow

explore

questions

13

In a packaged agent tool world such as AgentDojo — a benchmark that ships mock tools plus tasks — where does the attacker's text actually enter the episode, and why is that different from typing it into the user's prompt?

level: juniorimportance: must knowfreq 68%

answer

  1. payload rides in tool output
  2. benign user task, hostile fetched data
  3. data vs instruction boundary
  4. clean and poisoned replay of one fixture
  5. channels bounded by the mocks

basics

~20 s

It enters through tool output. The mock tools return data — an email body, a document, a transaction note — and the attacker text sits inside that returned content. The user's request stays benign, so the episode tests whether the agent obeys instructions it read while working, not instructions its user typed.

solid answer

~50 s

A packaged tool world is a fixture: mock services (a mailbox, a payments app, a workspace) plus a set of benign user tasks. The attacker content is planted **in what a tool returns**, not in the user turn. That placement is the whole point of the instrument. A prompt-typed instruction tests whether the model refuses a hostile *user*; a tool-returned instruction tests whether the model can tell *data it fetched* from *instructions it was given* — the trust boundary an agent crosses on every tool call. The practical consequence is that the same fixture can be replayed twice: once with the clean mock content, once with the same content carrying a planted instruction. The user task, the tool set and the world state are identical, so a behavioural difference between the two replays has one candidate cause. That is what you buy from a packaged world, and it is why the mocks' realism — what they return and what actions they expose — bounds everything you can learn.

go deeper

for a junior

Should say the attacker text comes back inside a tool result rather than from the user, and that the user's task is a normal one.

for a middle

Adds why the placement matters — the model cannot separate fetched data from given instructions — and that the same fixture replays clean and poisoned.

for a senior

Points out that channel count is a property of the mocks, insists reports name the channel exercised, and flags public payloads being pre-known to filters.

for a principal

Frames tool-output trust as an architectural property of the product, not a benchmark artifact, and decides what the organisation reports from fixtures like this.

### What a packaged tool world is made of Three separable parts, and the vocabulary matters because interviewers probe it. 1. **An environment** — a small serialisable state object (a mailbox of messages, an account with a transaction list, a file drive, a travel-booking store) plus *mock tools*: ordinary functions that read and write that state. No network, no real service, no side effects outside the object. 2. **User tasks** — benign goals ("summarise my unread mail", "pay the invoice my landlord sent"), each paired with a programmatic checker that inspects the end state to decide whether the agent did the job. 3. **Injection tasks** — a *separate* goal the attacker wants ("send the balance to an outside address"), each with its own checker over that same end state. AgentDojo's contribution is the wiring between them. Its seeded environment data contains named placeholder slots sitting inside ordinary-looking content — a message body, a file, a memo line on a record — and the harness substitutes the attacker's string into those slots before the episode starts. Then it runs the *benign* user task. The agent calls a mock tool because its own plan needs the data; the mock returns the seeded record; and the planted instruction arrives as the **return value of a call the agent chose to make**. ### Why that channel, and not the user turn By the time the tool result reaches the model it is just more tokens in the same context window as the system prompt and the user's request. There is no field, no colour, no privilege bit that says "this is data, not instruction". So the fixture measures a different property from hostile-user prompting: not *will the model refuse a person who asks for something bad*, but *can the model keep fetched content on the data side of the boundary while acting on the user's behalf*. Refusal training governs the first; almost nothing in a typical stack governs the second. A product can score cleanly on one and badly on the other, which is why the two must never be collapsed into a single "safety" percentage. The second thing the placement buys is the paired replay: run the same user task twice, once against clean seed data and once against the same data carrying the planted slot. Same task text, same tools, same world state, one variable different. ### What a sweep costs Nothing per tool call — the mocks are local functions — but everything per *episode*. An episode is a multi-turn agent loop: plan, call, read result, call again, answer. Five to fifteen model calls is typical, each carrying the accumulated transcript, so tokens per episode grow super-linearly with steps. A suite in the hundreds of user-task/injection-task pairings is therefore thousands of model calls and hours of wall clock even at healthy concurrency, and hosted rate limits usually bind before the budget does. The mocks are free; the agent loop is not. ### Where the number misleads The headline output is an attack success rate. Three readings of it are wrong. - **Zero does not mean immune.** The denominator counts pairings the *fixture* can express. Coverage of this channel is a property of the mocks: if only three tools return free text, you have three entry points no matter how many payload variants you write. Payload variety and channel variety are different axes and only one of them is under your control. - **Zero can mean nothing was delivered.** In many pairings the user task never calls the tool holding the planted content, so the agent never saw it. Those episodes score as attack failures and drag the average down. Report success conditioned on delivery, and log per episode whether the poisoned bytes ever entered the context. - **Zero can mean recognition.** Public fixtures circulate; their strings reach vendor filters and training corpora. A clean pass may record familiarity with a well-known suite rather than robustness of your product. ### What to check before believing it Dump the fully rendered tool result for one episode and confirm the payload is where you think it is, in the container you think it is in. Run a deliberately obedient stub agent and confirm every injection checker actually fires — a checker that can never fire produces exactly the 0% you were hoping to see. Confirm the clean pass still satisfies the user-task checker, so you know the agent was capable of the benign job at all. And when you report, name the channel: "instruction planted in returned tool content, benign user task", never a bare percentage.

  • Why does the user task in these fixtures stay benign?
    So the only hostile element is the planted tool content. If the user request were also hostile you could not tell which one moved the agent, and you would be measuring refusal of a user rather than confusion of data with instructions.
  • If the fixture's mocks expose only two tools that return free text, what does that cap?
    The number of distinct entry channels you can exercise, regardless of how many payloads you write. Payload variety and channel variety are separate axes, and the mocks fix the second one.
  • Name a production tool surface a packaged mailbox mock does not represent.
    Anything with paginated, truncated or multi-format results — long threads, attachments, quoted HTML, other tenants' shared documents — plus tools that fail with auth or permission errors partway through.

saying these in an interview costs you the question

  • Describes the fixture as sending jailbreak prompts as the user — that is a different instrument.
  • Assumes the harness marks tool output as untrusted for the model; nothing in the context does.
  • Reports one success number without saying whether the instruction arrived via user turn or tool result.
  • Treats a pass as evidence about hostile users, or vice versa.

context

open as a page

An agent red-team harness reports a single number: the share of injected runs in which the agent carried out the attacker's instruction. A proposed defence drives that number to nearly zero by making the agent refuse almost every request. Why does it score so well, and what else must the harness measure?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Attack success rate only counts attacker-instruction executions. An agent that refuses everything never executes one, so it scores near zero: perfectly secure and completely useless. The harness must also run each task without any injection and score whether the agent finished the user's real work, then report both numbers together.

open as a page

A multi-step harmful-task set such as AgentHarm gives an agent a goal it can only reach through a chain of tool calls, while a jailbreak prompt list gives one prompt per item and reads the reply. What does the task set measure that the prompt list cannot?

level: middleimportance: must knowfreq 55%

basics

~20 s

A prompt list only records whether the model said something harmful once. A task set records how far the agent got along a chain of tool calls before it stopped. That makes a refusal at step one and a stop at step eight different outcomes, and it exposes actions the words alone never produce.

open as a page

A packaged agent tool world replays each benign user task once per injection task, plus one payload-free pass. For a suite of 40 user tasks and 15 injection tasks, how many task episodes is a full sweep, and what do you cut first when the model-call bill is too large?

level: middleimportance: must knowfreq 58%

basics

~20 s

Forty clean episodes plus 40 x 15 poisoned pairings is 640 episodes, each a multi-step agent loop costing many model calls. Cut the cross product first: sample a few injection tasks per user task instead of all of them, keep the full sweep for release candidates, and always report which pairings you actually ran.

open as a page

In an agent injection benchmark, "task completion under attack" is the share of tasks the agent still finishes while an injection is present. Should that share be computed over every task in the suite, or only over the tasks the agent finishes when no injection is present — and why does the choice change what the number means?

level: middleimportance: must knowfreq 55%

basics

~20 s

Compute it over the same fixed task set as the no-injection run, and read the drop between the two. Conditioning only on tasks the agent already passes cleanly isolates the defence's cost but shrinks the sample. Either way, report the benign number beside it; a bare under-attack percentage mixes agent capability with attack damage.

open as a page

In a multi-step harmful-task suite where most attempts end half-finished, why does scoring each task as a binary completed / not-completed throw away most of what the run cost you, and what would you report instead?

level: middleimportance: should knowfreq 42%

basics

~20 s

Binary scoring maps every unfinished attempt onto one bucket, so an agent that stopped at the first call and one that stopped one call short look identical. You paid for the whole chain, so report chain depth: which required steps were reached per task, plus the reason each attempt stopped.

open as a page

In a packaged agent tool world (mock tools plus tasks, e.g. AgentDojo), you compare a clean replay of a task with a poisoned replay of the same task. What must be identical between the two replays for the behavioural difference to be attributable to the planted content?

level: middleimportance: should knowfreq 44%

basics

~20 s

Everything except the planted bytes: same task, same starting world contents, same tool list and descriptions, same model, sampling settings and step limit. Hold the payload's container and position fixed too, so only its text differs. Drift in timestamps, result ordering or sampling shows up as a fake attack effect.

open as a page

You are about to run a multi-step agentic harmful-task suite against a metered hosted endpoint. How do you estimate the model and tool calls it will consume before you start, and what does capping the steps per task do to the results you get out?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Multiply tasks by required steps by model turns per step, add retries and repeats, then pilot ten tasks and measure the real per-task call count before extrapolating. A step cap bounds the bill but truncates long chains, so attempts that hit it look like non-completions unless you record the cap as their stop reason.

open as a page

Your agent passes every task in a packaged tool world built on mailbox and banking-style mocks, then leaks customer data in production against your real mail and payments APIs. Which properties of that mock environment make the outcome unsurprising, and what do you add?

level: seniorimportance: should knowfreq 52%

basics

~20 s

The mocks are not your tools. They return small, clean, deterministic data with no pagination, partial failures or auth errors, and they expose a fixed action set, so harms outside that set cannot be observed at all. A pass says the fixture found nothing, not that your production tool surface is safe.

open as a page

In an agent security benchmark, some episodes end in a context-length error, an exception from a simulated tool, or a provider timeout. The harness records no attacker action for those, so they sit in the same bucket as runs the defence genuinely stopped. What does that do to your reported security number, and how would you change the instrumentation?

level: seniorimportance: should knowfreq 42%

basics

~20 s

They inflate it. Give every episode an explicit termination reason and split three ways: attacker action succeeded, agent completed the task safely, or the episode aborted. Report the abort rate as its own figure and compare it between the defended and undefended arms; a defence that lengthens prompts can manufacture aborts.

open as a page

You lead red teaming for an agent product. When is adopting a packaged tool world — a public fixture of mock tools and paired benign/injection tasks — the wrong instrument, and what do you stand up instead?

level: principalimportance: should knowfreq 38%

basics

~20 s

It is the wrong instrument when you need a ship decision. Its mock tools, data shapes and action space are not yours, so its number describes the fixture. Keep it as a cheap, comparable regression check, and stand up a bespoke world replaying your own tool schemas with your own harm predicates.

open as a page

Two candidate defences are measured on the same agent injection suite. With defence A the attacker-action rate falls from 24% to 3% and task completion under attack falls from 71% to 38%. With defence B they are 11% and 66%. A platform team asks you to fold these into one "security score" so the two can be ranked. How do you respond?

level: principalimportance: should knowfreq 34%

basics

~20 s

Give them the pair, not a scalar. The exchange rate between one blocked injection and one failed user task is a product decision, not a benchmark output. B keeps most of the utility; A buys a lower attacker-action rate by breaking a third of the work. Ask what a residual success actually costs.

open as a page

You have one week of compute and a fixed call budget for a multi-step agentic harmful-task suite. How do you split it between breadth (many tasks, one attempt each, tight step caps) and depth (few tasks, repeated attempts, generous caps), and what is each split's resulting number unable to say?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Split by what the run has to answer. Breadth finds which task families the agent will engage with at all, but one attempt per task has no repeatability and tight caps truncate long chains. Depth gives a defensible per-task number with an interval, but says nothing about families you did not run.

open as a page