skip to content

Agent Red-Team Harnesses

An agent attack counts only when the world changed - a transfer sent, a row deleted - so the environment, the interception and the state check are the work. Interviewers ask what a harness missed.

on this pageshow

explore

questions

page 1 of 2

In a packaged agent tool world such as AgentDojo — a benchmark that ships mock tools plus tasks — where does the attacker's text actually enter the episode, and why is that different from typing it into the user's prompt?

level: juniorimportance: must knowfreq 68%

answer

  1. payload rides in tool output
  2. benign user task, hostile fetched data
  3. data vs instruction boundary
  4. clean and poisoned replay of one fixture
  5. channels bounded by the mocks

basics

~20 s

It enters through tool output. The mock tools return data — an email body, a document, a transaction note — and the attacker text sits inside that returned content. The user's request stays benign, so the episode tests whether the agent obeys instructions it read while working, not instructions its user typed.

solid answer

~50 s

A packaged tool world is a fixture: mock services (a mailbox, a payments app, a workspace) plus a set of benign user tasks. The attacker content is planted **in what a tool returns**, not in the user turn. That placement is the whole point of the instrument. A prompt-typed instruction tests whether the model refuses a hostile *user*; a tool-returned instruction tests whether the model can tell *data it fetched* from *instructions it was given* — the trust boundary an agent crosses on every tool call. The practical consequence is that the same fixture can be replayed twice: once with the clean mock content, once with the same content carrying a planted instruction. The user task, the tool set and the world state are identical, so a behavioural difference between the two replays has one candidate cause. That is what you buy from a packaged world, and it is why the mocks' realism — what they return and what actions they expose — bounds everything you can learn.

go deeper

for a junior

Should say the attacker text comes back inside a tool result rather than from the user, and that the user's task is a normal one.

for a middle

Adds why the placement matters — the model cannot separate fetched data from given instructions — and that the same fixture replays clean and poisoned.

for a senior

Points out that channel count is a property of the mocks, insists reports name the channel exercised, and flags public payloads being pre-known to filters.

for a principal

Frames tool-output trust as an architectural property of the product, not a benchmark artifact, and decides what the organisation reports from fixtures like this.

### What a packaged tool world is made of Three separable parts, and the vocabulary matters because interviewers probe it. 1. **An environment** — a small serialisable state object (a mailbox of messages, an account with a transaction list, a file drive, a travel-booking store) plus *mock tools*: ordinary functions that read and write that state. No network, no real service, no side effects outside the object. 2. **User tasks** — benign goals ("summarise my unread mail", "pay the invoice my landlord sent"), each paired with a programmatic checker that inspects the end state to decide whether the agent did the job. 3. **Injection tasks** — a *separate* goal the attacker wants ("send the balance to an outside address"), each with its own checker over that same end state. AgentDojo's contribution is the wiring between them. Its seeded environment data contains named placeholder slots sitting inside ordinary-looking content — a message body, a file, a memo line on a record — and the harness substitutes the attacker's string into those slots before the episode starts. Then it runs the *benign* user task. The agent calls a mock tool because its own plan needs the data; the mock returns the seeded record; and the planted instruction arrives as the **return value of a call the agent chose to make**. ### Why that channel, and not the user turn By the time the tool result reaches the model it is just more tokens in the same context window as the system prompt and the user's request. There is no field, no colour, no privilege bit that says "this is data, not instruction". So the fixture measures a different property from hostile-user prompting: not *will the model refuse a person who asks for something bad*, but *can the model keep fetched content on the data side of the boundary while acting on the user's behalf*. Refusal training governs the first; almost nothing in a typical stack governs the second. A product can score cleanly on one and badly on the other, which is why the two must never be collapsed into a single "safety" percentage. The second thing the placement buys is the paired replay: run the same user task twice, once against clean seed data and once against the same data carrying the planted slot. Same task text, same tools, same world state, one variable different. ### What a sweep costs Nothing per tool call — the mocks are local functions — but everything per *episode*. An episode is a multi-turn agent loop: plan, call, read result, call again, answer. Five to fifteen model calls is typical, each carrying the accumulated transcript, so tokens per episode grow super-linearly with steps. A suite in the hundreds of user-task/injection-task pairings is therefore thousands of model calls and hours of wall clock even at healthy concurrency, and hosted rate limits usually bind before the budget does. The mocks are free; the agent loop is not. ### Where the number misleads The headline output is an attack success rate. Three readings of it are wrong. - **Zero does not mean immune.** The denominator counts pairings the *fixture* can express. Coverage of this channel is a property of the mocks: if only three tools return free text, you have three entry points no matter how many payload variants you write. Payload variety and channel variety are different axes and only one of them is under your control. - **Zero can mean nothing was delivered.** In many pairings the user task never calls the tool holding the planted content, so the agent never saw it. Those episodes score as attack failures and drag the average down. Report success conditioned on delivery, and log per episode whether the poisoned bytes ever entered the context. - **Zero can mean recognition.** Public fixtures circulate; their strings reach vendor filters and training corpora. A clean pass may record familiarity with a well-known suite rather than robustness of your product. ### What to check before believing it Dump the fully rendered tool result for one episode and confirm the payload is where you think it is, in the container you think it is in. Run a deliberately obedient stub agent and confirm every injection checker actually fires — a checker that can never fire produces exactly the 0% you were hoping to see. Confirm the clean pass still satisfies the user-task checker, so you know the agent was capable of the benign job at all. And when you report, name the channel: "instruction planted in returned tool content, benign user task", never a bare percentage.

  • Why does the user task in these fixtures stay benign?
    So the only hostile element is the planted tool content. If the user request were also hostile you could not tell which one moved the agent, and you would be measuring refusal of a user rather than confusion of data with instructions.
  • If the fixture's mocks expose only two tools that return free text, what does that cap?
    The number of distinct entry channels you can exercise, regardless of how many payloads you write. Payload variety and channel variety are separate axes, and the mocks fix the second one.
  • Name a production tool surface a packaged mailbox mock does not represent.
    Anything with paginated, truncated or multi-format results — long threads, attachments, quoted HTML, other tenants' shared documents — plus tools that fail with auth or permission errors partway through.

saying these in an interview costs you the question

  • Describes the fixture as sending jailbreak prompts as the user — that is a different instrument.
  • Assumes the harness marks tool output as untrusted for the model; nothing in the context does.
  • Reports one success number without saying whether the instruction arrived via user turn or tool result.
  • Treats a pass as evidence about hostile users, or vice versa.

context

open as a page

An agent red-team harness reports a single number: the share of injected runs in which the agent carried out the attacker's instruction. A proposed defence drives that number to nearly zero by making the agent refuse almost every request. Why does it score so well, and what else must the harness measure?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Attack success rate only counts attacker-instruction executions. An agent that refuses everything never executes one, so it scores near zero: perfectly secure and completely useless. The harness must also run each task without any injection and score whether the agent finished the user's real work, then report both numbers together.

open as a page

When red-teaming an agent, why do teams stand up their own tool server for the agent to call instead of pointing it at the vendor's live tool server?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Because a server you host is repeatable. You can change a tool's description or schema between runs and replay the same attempt to see what moved. A live vendor server gives one shot: it changes under you, its calls touch real data, and a result there cannot be reproduced or varied.

open as a page

Before red-teaming a deployed assistant that can send email, file tickets and write CRM rows, what do you provision so that a successful attack does not touch real people or real records?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Provision throwaway identities the agent acts as: a scratch mailbox, a test tenant or sandbox project, and disposable CRM records. Point outbound tools at dry-run or test endpoints. Mark everything with a unique per-run tag so you can find it later, and agree who owns cleanup.

open as a page

In an agent red-team harness, why is the adversarial instruction placed in the result a tool returns rather than in the user's message, and what does a hit at each of those two boundaries actually prove?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Put it in the tool result, not the user message. The agent is already trained to be wary of user requests, but treats tool output as data it fetched. The harness intercepts the call and returns attacker-controlled content in its place. A hit there shows the agent obeys untrusted retrieved content.

open as a page

You are running an agent red-team harness against a tool-using agent. Beyond the assistant's text messages, what must the harness write to the run trace for another engineer to rerun your finding without you?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Capture every tool call the agent made: the tool name, the exact arguments, and the result returned to it, in order, with timestamps. Also record the prompts, the model settings, and which harness and environment version ran. Text-only transcripts leave a finding unreproducible.

open as a page

In an agent red-team harness, why is the agent's own transcript claim that it sent nothing not enough to score the run as a pass?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Because the transcript is the agent's own account, not evidence. Score against world state: snapshot the sandbox before the run, run it, then diff. If the outbound queue gained a message or a row is gone, the run failed no matter what the agent narrated in its final turn.

open as a page

A multi-step harmful-task set such as AgentHarm gives an agent a goal it can only reach through a chain of tool calls, while a jailbreak prompt list gives one prompt per item and reads the reply. What does the task set measure that the prompt list cannot?

level: middleimportance: must knowfreq 55%

basics

~20 s

A prompt list only records whether the model said something harmful once. A task set records how far the agent got along a chain of tool calls before it stopped. That makes a refusal at step one and a stop at step eight different outcomes, and it exposes actions the words alone never produce.

open as a page

A packaged agent tool world replays each benign user task once per injection task, plus one payload-free pass. For a suite of 40 user tasks and 15 injection tasks, how many task episodes is a full sweep, and what do you cut first when the model-call bill is too large?

level: middleimportance: must knowfreq 58%

basics

~20 s

Forty clean episodes plus 40 x 15 poisoned pairings is 640 episodes, each a multi-step agent loop costing many model calls. Cut the cross product first: sample a few injection tasks per user task instead of all of them, keep the full sweep for release candidates, and always report which pairings you actually ran.

open as a page

In an agent injection benchmark, "task completion under attack" is the share of tasks the agent still finishes while an injection is present. Should that share be computed over every task in the suite, or only over the tasks the agent finishes when no injection is present — and why does the choice change what the number means?

level: middleimportance: must knowfreq 55%

basics

~20 s

Compute it over the same fixed task set as the no-injection run, and read the drop between the two. Conditioning only on tasks the agent already passes cleanly isolates the defence's cost but shrinks the sample. Either way, report the benign number beside it; a bare under-attack percentage mixes agent capability with attack damage.

open as a page

In an agent red-team harness you replace the agent's real email-sending tool with a stub that records the call and returns success. An injection attempt now shows the agent calling that stub with attacker-chosen recipients and body. What does that result prove, and what does it not?

level: middleimportance: must knowfreq 55%

basics

~20 s

It proves the model was hijacked into deciding to send, and shows the exact arguments it chose. It does not prove delivery: the real endpoint might reject the recipient, demand a confirmation, or fail an authorisation check. The stub tests the model's decision, not the system's outcome.

open as a page

You need adversarial content to arrive inside a tool result for an agent you are testing. Compare wrapping the tool function inside the agent framework, proxying the call in front of the tool endpoint, and planting the content in data a real tool reads -- what does each let you observe and control?

level: middleimportance: must knowfreq 55%

basics

~20 s

Three placements. Wrapping the tool function inside the framework is easiest and sees parsed arguments, but skips the wire. A proxy in front of the endpoint sees real requests and responses but not framework post-processing. Planting content in data a genuine tool reads is highest fidelity and lowest control over timing.

open as a page

Your agent red-team harness replays a recorded attempt with the same prompt, the same sampling settings and the same seed, and the agent takes a different action. Why is recording the seed not enough to make an agent run replayable, and what does the trace need instead?

level: middleimportance: must knowfreq 48%

basics

~20 s

Because the run is not a pure function of the seed. The environment answers differently each time: tool results, clocks, IDs, retries, and the endpoint's own sampling all vary. Record the actual request and response bytes at every step so a replay can fall back to the recorded exchange.

open as a page

Your agent red-team harness ran each attack payload once against the target agent and printed the payloads that succeeded. Why is that list not an attack success rate, and what do you run instead before quoting a number?

level: middleimportance: must knowfreq 68%

basics

~20 s

One trial per payload gives a yes/no from a sampled, stateful system, not a rate. The same payload can land one run in five, so a single pass misses real attacks and promotes flukes. Re-run each payload a fixed number of trials and report hits over trials attempted.

open as a page

An agent red-team harness grades a run as "partial": the harmful chain reached step three of six before the agent was stopped. What does that grade license you to claim about the target system, and what does it not?

level: middleimportance: must knowfreq 60%

basics

~20 s

It licenses a claim about that run only: the chain reached step three under this prompt, seed and environment state. It does not show the target is safe at step four; a retry, a different phrasing or a fuller mailbox may carry it further. Partial is evidence of reach, not of a boundary.

open as a page

What does it mean to write the must-not-happen assertion for an agent red-team task before the run, and what goes wrong if you write it after reading the transcripts?

level: middleimportance: must knowfreq 62%

basics

~20 s

It means naming the forbidden side effect as a concrete, machine-checkable condition over environment state, and committing to it before you run anything. Written afterwards, the behaviour you observed sets the bar: you quietly accept whatever happened, and a near-miss that emptied one row becomes a pass.

open as a page

Your stand-in tool server reproduces a hit in the harness every time, but the same attempt against the partner's real tool server produces nothing. Which differences between your stand-in and the real server would you check first?

level: seniorimportance: must knowfreq 48%

basics

~20 s

Check where your copy is more permissive than the real service: looser schema validation, unsanitised and untruncated result bodies, stale description text, different field names, and errors that echo input back. Then diff your served tool definitions against the partner's current contract and a capture of real responses.

open as a page

An agent red-team run that substitutes tool results through a shim reports a large number of policy violations. Before you write any of them up, how do you establish that the shim did not manufacture them?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Re-run with the shim in place but a benign result substituted. If the behaviour still appears, the shim caused it, not the payload. Usual manufacturers: the shim answering a call that would have errored, returning instantly, changing call order, or replacing a result the framework would have truncated.

open as a page

An agent red-team harness run ends mid-chain because the tool endpoint the agent was calling returned rate-limit errors, not because the agent refused or a policy control fired. How do you score that run, and what do you change in the harness?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Score it as an invalid run, not a defence. Rate limiting is infrastructure back-pressure from the endpoint, not a model refusal or a policy control; it would not stop a patient attacker. Mark the run inconclusive, exclude it from the denominator, back off and retry, and log the cause.

open as a page

An agent red-team harness drives an LLM agent through a multi-step harmful task in a sandboxed environment and scores the outcome. Why record how far along the chain the agent got, instead of only a pass/fail on the final harmful outcome?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Because a pass/fail hides how close the run came. Recording the last step the agent reached shows whether it was stopped early at planning or only at the final harmful action, tells you which control fired, and gives severity a distance input. Two zero-score runs can be very different.

open as a page

In a multi-step harmful-task suite where most attempts end half-finished, why does scoring each task as a binary completed / not-completed throw away most of what the run cost you, and what would you report instead?

level: middleimportance: should knowfreq 42%

basics

~20 s

Binary scoring maps every unfinished attempt onto one bucket, so an agent that stopped at the first call and one that stopped one call short look identical. You paid for the whole chain, so report chain depth: which required steps were reached per task, plus the reason each attempt stopped.

open as a page

In a packaged agent tool world (mock tools plus tasks, e.g. AgentDojo), you compare a clean replay of a task with a poisoned replay of the same task. What must be identical between the two replays for the behavioural difference to be attributable to the planted content?

level: middleimportance: should knowfreq 44%

basics

~20 s

Everything except the planted bytes: same task, same starting world contents, same tool list and descriptions, same model, sampling settings and step limit. Hold the payload's container and position fixed too, so only its text differs. Drift in timestamps, result ordering or sampling shows up as a fake attack effect.

open as a page

You host the tool server your agent under test calls. Between two runs you edit both the tool's description text and its parameter schema, and the attempt starts succeeding. What is wrong with that experiment, and how would you rerun it?

level: middleimportance: should knowfreq 42%

basics

~20 s

Two variables moved, so you cannot say which caused the change — and with a stochastic model, neither may have. Rerun with one edit at a time from a pinned baseline, repeat each configuration enough times to compare success rates, and record which fixture version produced each run.

open as a page

A red-team run against a mailbox-and-documents assistant uses a freshly created empty test mailbox and a brand-new account with no history or shared drives. Why can this understate real risk, and what do you put in that account before running?

level: middleimportance: should knowfreq 42%

basics

~20 s

An empty account gives an exfiltration attempt nothing to steal and no permissions to abuse, so the run scores a false negative. Seed it with realistic decoy documents, contacts and message history, grant the same roles a real user holds, and mark every seeded item with a canary string you can search for.

open as a page

In an agent red-team finding you attach both a rate — successes over trials for one attack payload — and a saved recording of one successful run that a reader can replay. What does the replay show that the rate cannot, and what does the rate show that the replay cannot?

level: middleimportance: should knowfreq 48%

basics

~20 s

The replay shows the mechanism: the exact turns, tool calls and world change that made it a real break, so a reader can watch it happen rather than trust a number. The rate shows how often it happens, which sets priority. Neither substitutes for the other; a finding needs both.

open as a page

You are about to run a multi-step agentic harmful-task suite against a metered hosted endpoint. How do you estimate the model and tool calls it will consume before you start, and what does capping the steps per task do to the results you get out?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Multiply tasks by required steps by model turns per step, add retries and repeats, then pilot ten tasks and measure the real per-task call count before extrapolating. A step cap bounds the bill but truncates long chains, so attempts that hit it look like non-completions unless you record the cap as their stop reason.

open as a page

Your agent passes every task in a packaged tool world built on mailbox and banking-style mocks, then leaks customer data in production against your real mail and payments APIs. Which properties of that mock environment make the outcome unsurprising, and what do you add?

level: seniorimportance: should knowfreq 52%

basics

~20 s

The mocks are not your tools. They return small, clean, deterministic data with no pagination, partial failures or auth errors, and they expose a fixed action set, so harms outside that set cannot be observed at all. A pass says the fixture found nothing, not that your production tool surface is safe.

open as a page

In an agent security benchmark, some episodes end in a context-length error, an exception from a simulated tool, or a provider timeout. The harness records no attacker action for those, so they sit in the same bucket as runs the defence genuinely stopped. What does that do to your reported security number, and how would you change the instrumentation?

level: seniorimportance: should knowfreq 42%

basics

~20 s

They inflate it. Give every episode an explicit termination reason and split three ways: attacker action succeeded, agent completed the task safely, or the episode aborted. Report the abort rate as its own figure and compare it between the defended and undefended arms; a defence that lengthens prompts can manufacture aborts.

open as a page

After a live-fire agent red-team run in which the agent really sent mail and created records, how do you establish exactly what the test left behind and remove it — and what residue can you not remove?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Do not reconstruct it afterwards. Have the harness log every outbound tool call with the identifier the system returned, so the run produces a delete list. Tag artefacts with a per-run marker, delete, then re-search until the marker returns nothing. Delivered mail, webhook fan-out and audit entries stay.

open as a page

Your harness substitutes adversarial content into an agent's tool results and the agent never misbehaves. Before reporting the agent as resistant, how do you verify the content actually reached the model?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Log what the model actually received, not what you handed the shim. Frameworks truncate, summarise, re-serialise or drop tool results before they become context, so your text may never arrive intact. Without that check a clean run is not evidence of resistance; it may only prove your content was cut.

open as a page

showing 1–30 of 44