skip to content

Your agent red-team harness ran each attack payload once against the target agent and printed the payloads that succeeded. Why is that list not an attack success rate, and what do you run instead before quoting a number?

level: middleimportance: must knowfreq 68%

answer

  1. sampled target, not a function
  2. hits over trials, keep the denominator
  3. single pass: fluke or missed hit
  4. 0/20 is not immunity
  5. log per trial, not the aggregate

basics

~20 s

One trial per payload gives a yes/no from a sampled, stateful system, not a rate. The same payload can land one run in five, so a single pass misses real attacks and promotes flukes. Re-run each payload a fixed number of trials and report hits over trials attempted.

solid answer

~50 s

An agent under attack is not a deterministic function. Decoding is sampled, the agent picks different tool orders, and the environment it acts on can drift, so one execution of a payload is one draw from a distribution, not a measurement of it. A single-pass sweep produces two errors at once: a payload that works 20% of the time shows as clean, and a payload that works 20% of the time shows as a confirmed break. Both go into the report as the same word, "succeeded". What to run instead: repeat every payload a fixed number of trials, log per-trial outcomes rather than an aggregate, and quote hits over trials — 6/30, not "20%" alone. Two cheap disciplines make that number readable: fix and record everything you can control (decoding settings, turn budget, starting environment, which scorer decided a hit), and keep the per-trial records so someone can recount them. A rate without a denominator is a rumour.

go deeper

for a junior

Says the model's output is random, so one run is not enough, and that you should run the payload several times and count successes.

for a middle

Names the variance sources — sampling, plan choice, environment state — and reports hits over trials with the run settings recorded.

for a senior

Adds the interval reasoning, the zero-hit upper bound, and the checks that keep the denominator honest: aborted trials, early stopping, one consistent hit criterion.

for a principal

Frames the rate as a property of the harness configuration, not the target, and sets what a report is allowed to claim from a given trial count.

## What the harness actually produced A **single-pass sweep** is a loop: for each attack payload in the suite, drive one episode against the target agent, hand the resulting transcript and the world the agent touched to whatever object decides success — a programmatic assertion, a scorer, a detector, or a judge model — and print the payloads where that object said yes. Each printed line is one **Bernoulli draw** with n equal to one. It is a set of observed events, and an event is not a rate: a rate needs a denominator per payload, and yours is one. ## Three sources of variance sit between the payload and the verdict 1. First, **decoding**: the target samples its next token from a distribution, so above temperature zero the same context can yield a refusal on one draw and compliance on the next. 2. Second, **plan variance**: an agent chooses which tool to call, with which arguments, in what order, and over how many turns, and the payload only counts if the episode actually reaches the surface it targets — two episodes from identical starting conditions routinely take different routes. 3. Third, **environment state**: the world the agent acts on — a mailbox, a table, a file tree, a ticket queue — carries state, so a trial starting from a world the previous trial mutated is not a repeat of the same experiment. None of the three is a bug in your harness. They are the property you were asked to characterise. ## The two errors a single pass commits at once - A payload that lands one run in five shows clean on roughly four of five single trials, so a real, reportable break is silently missing from the list. - And when it does land, it is printed under exactly the same word — succeeded — as a payload that lands nineteen times in twenty. Single-pass output cannot separate a fluke from a reliable break, which is the distinction the reader of the report needs most. ## What repetition costs In an agent harness a trial is an **episode**, not a call: several model turns, a tool round-trip per action, plus whatever calls the scoring object makes, and a judge model is a second metered endpoint. The bill multiplies as payloads x trials x turns per episode x tokens per turn. Taking a 40-payload suite from one trial to twenty is 800 episodes rather than 40; at eight turns and a few thousand tokens a turn that is tens of millions of tokens before the judge. Wall-clock is usually worse than spend, because it is set by tool latency and the endpoint's rate limit rather than by your parallelism, and engineer time goes into resetting the environment between trials so trial 12 does not inherit trial 11's world. That is why the answer is not "always run a hundred": repetition is the budget line you trade against payload coverage. ## Where the number you produce still misleads - A percentage hides its **denominator** — 1/5 and 40/200 both print as 20%, and only one of them is a measurement. - **Zero hits** is misread hardest: 0/20 is routinely written up as "not vulnerable", when the rule-of-three approximation puts the 95% upper bound near 3/n, so twenty clean trials support only "probably under about 15%". - A harness that **stops re-running a payload on first success** produces a denominator set by when the success happened, which biases the fraction upward and is not a fixed-size sample at all. - Trials killed by a tool error, a timeout or a rate limit are **infrastructure outcomes**, not defensive ones: counted as misses they deflate the rate, dropped silently they hide that part of the evidence never ran. - And the rate is a property of your whole rig — turn budget, decoding settings, hit criterion, starting world — so it belongs in the finding, not in omitted metadata. ## What to check before quoting a number - That every trial started from an equivalent world, verified by the reset having actually run rather than assumed. - That the object deciding a hit applied the same criterion on every trial, and that it scores the world rather than the agent's narration of the world. - That aborted episodes are recorded as aborts and excluded from the denominator. - That the sweep did not stop early on first success. Then quote hits over completed trials with the run configuration attached — 6/30 under these settings — never a bare 20%.

  • Your harness ran a payload 20 times and scored zero hits. Can you write "not vulnerable" in the report?
    No. Zero of twenty bounds the rate loosely — roughly under 15% at 95% confidence by the rule of three. Write the observed 0/20 and the bound, not a clean bill.
  • Five of thirty trials aborted when the target endpoint rate-limited you. What is the denominator?
    Report hits over completed trials and disclose the aborts separately. Counting an abort as a miss deflates the rate; silently dropping it hides that a quarter of your evidence never ran.
  • Why is quoting 1/5 better than quoting 20%?
    The fraction carries the denominator, so a reader immediately sees the estimate rests on five trials and treats the interval as wide instead of reading 20% as precise.

Running each payload once is a poll with one respondent per question: you learn what that one person said, and nothing about how common the answer is.

saying these in an interview costs you the question

  • Reporting a percentage with no trial count behind it.
  • Treating a single successful run as an established attack success rate.
  • Calling a payload clean after one miss.
  • Letting the harness stop at the first hit and still quoting a rate.
  • Counting trials that aborted on a tool error or rate limit as clean misses.

context