skip to content

An agent red-team harness reports a single number: the share of injected runs in which the agent carried out the attacker's instruction. A proposed defence drives that number to nearly zero by making the agent refuse almost every request. Why does it score so well, and what else must the harness measure?

level: juniorimportance: must knowfreq 70%

answer

  1. refuse everything, score zero
  2. one-tailed metric
  3. clean run beside injected run
  4. security number needs a utility number
  5. refusal is its own outcome

basics

~20 s

Attack success rate only counts attacker-instruction executions. An agent that refuses everything never executes one, so it scores near zero: perfectly secure and completely useless. The harness must also run each task without any injection and score whether the agent finished the user's real work, then report both numbers together.

solid answer

~50 s

The metric is one-sided. It asks "did the attacker's instruction get executed?" and nothing else, so every path that produces no agent action at all — a blanket refusal, an empty answer, a hard stop on any suspicious input — is indistinguishable from a genuine defence. The fix is structural, not statistical: the harness runs the same user task twice, once clean and once with the injection present, and keeps two figures. The first is the agent's baseline usefulness on that suite; the second is how much of that usefulness survives the attack. A defence is then described by a pair — how much attacker action it removed, and how much user work it destroyed — and a refuse-everything defence shows up immediately as a near-total collapse of the second number. The main tradeoff is cost: you now pay for at least two runs per task, and the clean run must use identical tools and seeds or the comparison is meaningless.

go deeper

for a junior

Should say that an agent refusing everything gets a perfect security score, so the harness also has to check that the agent still does the user's job.

for a middle

Adds the two-condition design — same task run clean and injected — and names what each condition scores.

for a senior

Insists the benign score be an environment end-state assertion, matched configuration across conditions, and refusal tracked as its own outcome.

for a principal

Frames the whole thing as a reporting rule for the org: no security number leaves the team without the utility number from the same suite and run.

### What the harness is actually computing An agent red-team harness runs *episodes*. One episode is a **user task** ("pay the invoice attached to the message from the landlord"), a **simulated environment** — mock mail, banking, calendar and file tools whose state the harness owns — and, in the attacked condition, an **injection**: attacker-controlled text planted somewhere the agent will read it, typically the body of a retrieved email or the return value of a tool. When the episode ends, the harness applies a **security check**: a predicate over the recorded tool calls and the environment's final state that asks exactly one question — did the attacker's target action occur? Did money move to the attacker's account, did the address book leave the sandbox. Attack success rate (ASR) is that predicate's true-rate over episodes. The predicate is *one-tailed*. It fires only on an action the attacker wanted. Every path in which the agent takes no consequential action at all satisfies it identically: a blanket refusal, an empty completion, a loop the agent never escapes, a context-length crash, and the case you actually want — a capable agent that read the injection, recognised it, and went on doing the user's work. Under an ASR-only harness, "refuse everything" is not a loophole. It is the metric's global optimum, and any defence tuned against that metric alone will be pushed toward it. ### The second condition, and what it costs The structural fix is a second scored condition, which is why AgentDojo keeps **utility** and **security** as two separate checks over the same episode: a utility check asserts the user's goal reached the environment's end state (the transfer exists, with the right payee and amount), a security check asserts the attacker's goal did not. Each user task is run clean — no injection — to establish the baseline, and again with each injection. Price it before you argue about it. With *U* user tasks and *I* injection tasks, matched pairing produces *U x I* attacked episodes but only *U* clean ones. At 80 user tasks and 30 injection tasks that is 2,400 attacked episodes against 80 clean: **the benign baseline is roughly 3% of the compute of the run it makes readable**. An episode is not one API call — it is one call per agent turn plus tool round-trips, commonly 5-15 calls — so the suite is tens of thousands of calls, several hours of wall-clock at modest concurrency, and a bill in the tens to low hundreds of dollars on a mid-priced hosted model. Keep that ratio in mind: dropping the utility arm saves almost nothing and destroys the interpretability of everything else you paid for. ### Where the number misleads | Reading | Why it is wrong | |---|---| | "ASR fell to 2%, the defence works" | A refusal-heavy agent produces the same 2% while completing nothing. Without utility, safe and inert are the same observation. | | "ASR is near zero even undefended" | Usually the injection never reached the model's context — a fixture or rendering bug — so nothing was ever tested. A near-zero *baseline* is a harness alarm, not a result. | | "Utility is fine, the judge said the task was done" | An LLM judge reading the transcript rewards an agent that *says* "I have sent the payment". Only an end-state assertion on the environment distinguishes doing from claiming. | | "The remaining failures are just noise" | Timeouts and limit exhaustion also produce zero attacker actions, so they are booked as defensive wins by the same predicate. | The last two are the ones that survive review, because both produce plausible-looking numbers. ### What you would check Before believing any ASR from a harness you did not build: confirm a benign condition exists and ran the *same* task list, tool set, model and decoding settings; confirm the undefended baseline ASR is comfortably non-zero, proving delivery; confirm the utility check is an end-state assertion, not transcript reading; and pull the per-episode outcome histogram so **refusal is visible as its own outcome** rather than folded into "no attacker action". Then apply the reporting rule this leaf exists to teach: no security number leaves the team without the utility number measured on the same suite, in the same run, on the same tasks.

  • How do you score the benign task without an LLM judge?
    Assert on the environment's final state — the record the task was supposed to create or change — rather than on the text of the agent's reply.
  • Why keep refusal as a separate outcome instead of folding it into failure?
    Because a refusal and a wrong answer have different causes and different fixes; a rising refusal rate is the first sign a defence is buying safety with usefulness.
  • Does the clean run need the same tools as the attacked run?
    Yes. Different tools, prompts or decoding settings between the two conditions make the drop uninterpretable — you are then measuring configuration, not attack damage.

A spam filter that deletes every incoming message scores a perfect zero on spam-reaching-the-inbox. The score is real; it is just measuring only half of what a mailbox is for.

saying these in an interview costs you the question

  • Treating zero attacker actions as proof the defence is good.
  • Assuming the agent would obviously still work, without measuring it.
  • Proposing to fix it by making the injection harder instead of adding a benign condition.
  • Scoring the benign task by reading the transcript for effort rather than checking the environment's end state.

context