skip to content

In an AI red-team report, what is the difference between writing "we exercised this behaviour and observed no failures" and "we never exercised this behaviour", and why must the deliverable keep the two visibly apart?

level: juniorimportance: must knowfreq 65%

answer

  1. three states: failed, clean, untested
  2. silence reads as passed
  3. remainder section, referenced from summary
  4. not exercised vs out of scope
  5. clean rows need their basis

basics

~20 s

The first is evidence; the second is a gap. A reader who sees neither statement assumes the area passed. So the report must list untested areas explicitly, in the same place as the results, rather than leaving them out. Absence of a finding is only meaningful where something was actually attempted.

solid answer

~50 s

Two sentences that look alike on a slide mean opposite things. **"Exercised, no failures observed"** is bounded evidence: something was attempted, some scoring step decided it did not succeed, and a reader can ask how many attempts and by whose judgement. **"Never exercised"** is an admission of a gap, and it carries no evidence at all. Readers default to charity. An area a report does not mention is assumed to have been looked at and found fine — nobody reads silence as a hole. That is why a deliverable needs a remainder section in the body, referenced from the summary, naming the attack categories, entry points and harm classes that were planned but never run, and why (hours, authorisation, environment, a target that was unavailable). The cost is honest: a stated remainder makes the headline look weaker than a silent one. That is the trade you are being paid to make.

go deeper

for a junior

Says untested is not the same as clean, and that the report should list what was not tested.

for a middle

Adds the third state (out of scope, agreed at kickoff) and attaches a basis — attempt counts, entry point — to each clean result.

for a senior

Structures the deliverable so the summary cannot outrun the evidence table, and knows which untested areas came from a decision versus an operational failure.

for a principal

Sets the house template and the rule that no assurance sentence may reference an area without a row, and negotiates with clients who want the remainder removed.

“Exercised and clean” and “never exercised” are two different amounts of evidence — one bounded, one zero — and once either one becomes white space on a page, a reader cannot tell them apart. Keeping them apart is a structural problem, not a wording one. **The four states** | State | What happened | What the evidence supports | |---|---|---| | Exercised, failed | attempts ran; a scorer called some of them hits | a defect exists here, sized by the attempts behind it | | Exercised, clean | attempts ran; no attempt scored a hit | nothing was found *at this depth, on this surface, by this scorer* | | Not exercised | nobody ran it — hours, environment, credential, provisioning | unknown | | Out of scope | a named person agreed at kickoff it would not be run | unknown, and someone accepted that | Most report templates only carry the first row well. The other three collapse into absence, and absence is read as the second one. **Why the tooling pushes you towards silence** Scanners report what they ran, never what they did not. `garak` writes, per run, a `*.report.jsonl` holding one record per attempt plus a `*.hitlog.jsonl` holding only the attempts a detector scored as a hit; a probe module you never named in `garak --probes` contributes no records at all, and a behaviour that no shipped probe module targets contributes none either. `promptfoo`'s red-team summary is assembled from the plugins named in `redteam.plugins` — anything absent from that list is absent from the table. So in the raw artefacts, “we drove it and it held” is a set of rows scoring zero, while “the idea never entered the config” is no rows; by the time a human pastes a summary into a deck, both are blank. The written plan is the only place the distinction survives, and only if somebody wrote it down before the run. **What a clean row is actually worth** Zero hits is a sampling statement, and its value comes from two numbers the row rarely carries. Depth: `garak`'s `-g` / `--generations` sets how many outputs are drawn per prompt, and a probe left at the tool's default is a handful of samples per prompt, not a campaign. Judgement: `garak` pairs each probe with detectors, and a detector that never fires on the phrasing your target happens to use produces a clean row meaning “nothing matched”, which is not “nothing happened”. So an exercised/clean row needs its basis attached — entry point driven, attempts made, and what decided a response did not count. **What it costs** Producing the ledger is the cheapest line in the engagement: enumerate the planned rows at kickoff with the plan open (about an hour), keep the state current as the run proceeds (minutes a day), reconcile it against the harness output at the end (about an hour). Under a day across a two-week job, and most of it is work you need anyway to know when you are finished. The expensive part is not editorial. A summary with a stated remainder always looks worse than a silent one produced by the identical run, and whoever commissioned the engagement may push back on exactly that. That is the trade this leaf exists to teach: you pay in headline strength to buy a claim that survives being questioned. **Where the number misleads** Two readings to head off. First, an unmentioned area reads as passed — readers default to charity, and nobody interprets silence as a hole. Second, absence quietly inflates every per-category statistic in the document: “3 findings across 9 categories” implies nine categories were driven, and if six were never touched, the honest statement is three findings across three driven categories — the same numerator against a denominator three times smaller, which is a far worse result wearing a better number. **What I check before sending** - Does every planned area appear exactly once, in exactly one state? - Does each exercised/clean row carry attempts, entry point and scorer? - Does any assurance sentence reference an area that has no row behind it? - Could a reader who reads only the summary page name the untested areas? If not, the summary misleads even though every sentence in it is true. - Is “not exercised” separated from “out of scope”? The second was signed by someone at kickoff; the first is an operational loss with no owner yet, and it is the one that should generate a follow-up.

  • Why separate "out of scope" from "not exercised" when both mean no evidence?
    Out of scope was a decision someone accepted and signed at kickoff; not exercised is usually an operational loss — hours, a missing environment, a credential. The first has an owner, the second needs one.
  • A client asks you to delete the untested-areas list because it 'undermines the report'. What do you do?
    Keep it and offer to move it: retitle it as agreed limitations, put the reasons alongside, and note which items a follow-up would cover. Removing it converts an honest report into an assurance claim you cannot support.

saying these in an interview costs you the question

  • Treating an unmentioned area as implicitly out of scope
  • An executive summary that says 'no issues found' with no bound on what was looked at
  • Listing untested areas only in an appendix nobody reading the summary reaches
  • Presenting clean results with no attempt counts or entry point

context