skip to content

In an agent red-team finding you attach both a rate — successes over trials for one attack payload — and a saved recording of one successful run that a reader can replay. What does the replay show that the rate cannot, and what does the rate show that the replay cannot?

level: middleimportance: should knowfreq 48%

answer

  1. existence claim vs frequency claim
  2. replay = mechanism, rate = priority
  3. world-state diff, not narration
  4. record turn budget and hit criterion
  5. replay may not reproduce — say so

basics

~20 s

The replay shows the mechanism: the exact turns, tool calls and world change that made it a real break, so a reader can watch it happen rather than trust a number. The rate shows how often it happens, which sets priority. Neither substitutes for the other; a finding needs both.

solid answer

~50 s

They answer different questions. A rate answers "how often", and that drives severity, remediation ordering, and whether a fix can be verified by re-measuring. It says nothing about how the break worked — a reader cannot tell from 6/30 whether the agent leaked data, called a destructive tool, or merely emitted words. A replay answers "what happened". A stored transcript with the tool calls, their arguments, and the before-and-after world state lets an engineer see the causal path and argue about the fix. It is also the artefact that survives disagreement: when someone calls the harness's hit criterion too loose, you point at the recording, not the percentage. What a replay cannot do is establish frequency, and it cannot promise to reproduce — a saved success is one draw. That is why the two travel together.

go deeper

for a junior

Says the recording lets someone see the attack actually work, while the rate says how often it works.

for a middle

Separates the existence claim from the frequency claim and names what the recording must contain to be replayable.

for a senior

Adds that a replay may not reproduce against a sampled target, reports replay-reproduction honestly, and uses the pre-fix rate to size the fix-verification run.

for a principal

Sets what evidence a filed finding must carry as a house rule, including scrubbing captured tool arguments and pinning the artefact to the harness configuration.

## Two claims, two artefacts - "This agent can be driven to do X" is an **existence claim**, and one well-evidenced run settles it. - "This agent does X about a fifth of the time under these conditions" is a **frequency claim**, and only repeated trials settle it. Reports carrying only one of the two draw predictable pushback: a bare rate draws "show me", and a bare recording draws "is this a fluke?". The artefacts are not redundant, and neither degrades gracefully into the other. ## What the rate does Successes over completed trials, with the trial count and run configuration attached, is a severity and ordering input. It tells a defender how often an attacker starting from the same position gets the outcome, which drives priority, and it is the baseline you re-measure against after a mitigation. What it cannot do is say what happened. From 6/30 a reader cannot tell whether the agent leaked a customer record, called a destructive tool, or merely emitted disallowed words — nor whether the object that scored those six hits was scoring the right event at all. ## What the replay does A stored, re-drivable recording of one successful episode shows the **causal path**. To be worth attaching it must carry: - the payload and where it was placed relative to what the agent reads; - the full turn sequence with tool names and arguments; - the harness settings that shaped the episode — turn budget, decoding settings and any seed, which object decided the run was a hit and on what criterion; - and the environment snapshot the episode started from plus the diff it produced. The **world-state diff** is what turns a transcript into evidence. An agent's narration is not testimony about its own actions: an agent can write "I did not send the message" in the turn after the send tool returned success, so a miss scored off transcript wording and a miss scored off the outbound queue are different measurements wearing one label. ## What it costs Capture is not free. - Full-fidelity recording of tool arguments and environment snapshots multiplies log volume per episode, which is why teams sample it — and then discover the one successful episode was the one not recorded. The rule that avoids that: sample freely, but always retain the complete record for any episode the scorer marked a hit. - There is engineer time in building the snapshot-and-diff machinery at all, - and review time in scrubbing what the recording caught, because captured tool arguments routinely contain bearer tokens, session cookies, customer records, or the contents of the mailbox the agent read. An unscrubbed replay artefact pasted into a ticket is its own incident. ## Where each one misleads The **replay's failure mode** is the implied promise "run this and watch it break". Against a sampled agent a replay may simply not land, and a reader who tries twice and fails concludes the finding was fabricated. State the reproduction fraction — "replayed 3 of 10 attempts" — because intermittency is the property under measurement, not an embarrassment. The replay also drifts silently: re-driven under a different turn budget, a different tool set, or a mutated environment it is a different experiment, so pin the artefact to the run manifest it came from. The **rate's failure mode** is the opposite. It looks objective while inheriting every assumption of the scorer that produced it; a rate computed by a judge that rewards refusal-shaped wording will report a comfortingly low number on an agent that is quietly acting. A number cannot flag its own scorer, and the replay is the only artefact that lets a sceptic audit one. ## Where the pair pays off: fix verification After a mitigation ships, the recording tells you which exact scenario to re-drive, and the pre-fix rate tells you how many trials the post-fix run needs before "we saw no successes" carries weight. - Without the rate you cannot size the verification; - without the recording you cannot be sure you re-tested the same thing rather than a rephrasing the mitigation happens to catch. And when the pre-fix rate is very low, the pair tells you honestly that re-measurement cannot settle it and the verification has to be structural. ## What to check - That the recording holds a world diff, not only a transcript. - That credentials and customer data are scrubbed before it leaves the harness. - That a manifest is attached so a replay is comparable to the original. - And that the replay's reproduction fraction is stated rather than implied.

  • Your saved successful run does not reproduce on the first two replays. Do you withdraw the finding?
    No. Report the replay-reproduction fraction alongside the original evidence. Intermittency is the measured property, not a reason to unfile a behaviour you have a recorded world-state diff for.
  • Why is the agent's own transcript statement 'I did not send the message' insufficient evidence for a miss?
    Narration is not observation. Score the world — the outbound queue, the row, the file — because an agent can report inaction after acting.

saying these in an interview costs you the question

  • Filing a rate with no artefact anyone can inspect.
  • Filing a single screenshot and calling it an attack success rate.
  • Attaching a transcript with no world-state diff and trusting the agent's own account of what it did.
  • Claiming a saved run will reproduce on demand against a sampled target.
  • Archiving a replay with live credentials or tokens still in the captured tool arguments.

context