You lead an AI red team whose engagement reports quote attack success rates from an agent harness. What standard do you set for trial counts and for how variance appears in a report, and what do you do with the payload that landed once in fifty?
answer
- floor by claim type, not one number
- no bare percentages: n and interval
- run manifest gates comparison
- rare hit: file on impact, mark uncertainty
- commit verification budget at file time
basics
~20 sTie the trial floor to the claim: an existence finding needs one replayable success with evidence; a quoted rate needs a denominator and an interval in the text. Ban bare percentages. The one-in-fifty still gets filed on impact, marked low-rate and not verifiable by re-measurement at that budget.
solid answer
~50 sSet the standard by claim type rather than by one magic number. **Existence findings** — "this agent can be driven to do X" — require one archived, replayable run with a world-state diff, plus the trial count it came from. No rate required. **Rate claims** require a stated denominator and an interval, and are compared only against a rate produced under a matching run configuration. A percentage with no n does not pass review. **Coverage statements** must separate tested-and-clean from screened-at-low-depth; "no hits in three trials" is written as a bound. The rare hit is where teams break discipline in both directions. File it on impact, because severity comes from what the agent did, not how often; label it a low-rate observation with wide uncertainty; and note that confirming a fix would cost a large re-measurement, so judge the fix by the mechanism changed.
go deeper
Says every reported rate should include how many runs it came from and that a rare success still gets reported.
Distinguishes existence findings from rate claims and requires a denominator plus an artefact on each filed item.
Adds the run manifest, disclosed screening depth, abort policy, and the verification cost that must be budgeted when the finding is filed.
Owns the trade — fewer payloads, better characterised — makes low-rate findings a named review category so the standard does not suppress them, and defines structural verification for rates too low to re-measure.
**Why a single trial floor fails.** "Twenty trials for everything" is the rule teams reach for, and it is wrong in both directions at once. It spends most of an engagement's budget deepening payloads that never land, and it is still nowhere near enough power for the rare-but-catastrophic ones, where the number you care about is a fraction of a percent. The floor has to be a function of what the report claims and what decision the number feeds, not a constant. **A house standard by claim type.** - *Existence findings* — "this agent can be driven to do X" — require one archived, re-drivable run with a world-state diff, scrubbed of credentials and customer data, plus the trial count it came out of. No rate required. - *Rate claims* require a stated denominator of completed trials and an interval in the text. Bare percentages fail review, and a rate is compared only against a rate produced under a matching run manifest. - *Coverage statements* must separate tested-and-clean from screened-at-low-depth. "No hits in three trials" is written as a bound, never as clean. - *Every sweep emits a manifest* — trials, turn budget, tool set, decoding settings, hit-criterion identifier, environment snapshot, abort policy — because that is what makes any later comparison meaningful. - *Verification budget is committed at file time*, not discovered later; the re-measurement that shows a fix worked costs roughly what the original measurement cost. **What the standard costs.** Depth is paid for in breadth. Against a fixed engagement budget, insisting on denominators and intervals means fewer payloads, better characterised, and that shows up to whoever funds the engagement as "we tested less this quarter" unless you say it out loud first. There is review overhead — templates, a manifest emitter in the harness, someone confirming a filed rate carries its n — and storage and scrubbing cost for the mandatory artefacts. Price all of it once, in the standard, rather than letting each engineer absorb it invisibly and quietly stop complying. **The payload that landed once in fifty.** Two opposite failures show up here. Dropping it as noise is the worse one: a once-in-fifty action that empties a production table is a shipping blocker regardless of frequency, because your trial count is your sampling cost, not the attacker's — they retry until it lands and pay almost nothing per attempt. Over-claiming is the other: 1/50 is not "2%" in any useful sense, since the interval runs from well under one percent to nearly a tenth, so the number cannot support an ordering against another finding quoted at 3%. File it, rank it on impact, and label it explicitly as a low-rate observation with wide uncertainty. **Where the standard's own numbers mislead.** The trap is post-fix verification. Engineering will ask for a re-run showing the rate went to zero, and a sweep at any affordable budget cannot distinguish "fixed" from "rarer than we can see": fifty clean trials after a 1/50 finding is exactly the outcome you would expect if nothing had changed at all. Committing to that verification is committing to a result that cannot be earned, and worse, delivering it manufactures false confidence with the standard's own authority behind it. For low-rate findings the honest verification path is structural — did the mechanism that permitted the action actually change, is the tool now gated, is the untrusted content no longer reaching the planner — with re-measurement kept as a supporting signal only. The second trap is organisational: a rate standard creates pressure to under-report rare findings, because they are expensive to defend under a rule that demands denominators. Counter it deliberately by making low-rate findings a first-class, named report category with its own template and review path, so filing one is routine rather than something an engineer must argue for against the standard they are being held to. **What you check as the owner of the standard.** Sample filed findings for a manifest and an artefact, not merely a number. Check that screened misses are written as bounds. Check that at least one low-rate finding appears in engagements where you would expect one — their absence across a quarter is a symptom, not a clean result. And check that the verification budget committed at file time was actually spent, because an uncashed commitment is how a standard degrades into paperwork.
- Engineering asks you to re-run the sweep after the mitigation and confirm the once-in-fifty finding is gone. What do you tell them?That a sweep affordable at this budget cannot distinguish fixed from rarer. Verify structurally — show the mechanism that permitted it is closed — and use re-measurement only as a supporting signal.
- Why not simply mandate 100 trials per payload for consistency?It spends most of the budget on payloads that never land while still being underpowered for the rare-catastrophic ones, and it collapses engagement breadth. Tie depth to the claim and the decision instead.
- What stops a rate standard from quietly suppressing rare findings?Make low-rate findings a first-class report category with its own template and review path, so filing one is routine rather than something an engineer has to argue for against the standard.
Frequency is your sampling cost, not the attacker's: you pay fifty full episodes to see the behaviour once, while they can pull the handle all day for the price of one request.
saying these in an interview costs you the question
- One universal trial count mandated for every payload.
- Dropping a rare but high-impact hit as noise.
- Quoting 1/50 as a precise 2% with no uncertainty stated.
- Promising a post-fix rate as verification when the budget cannot resolve a low rate.
- Reporting screened-at-low-depth payloads as tested and clean.