An AI red-team report is read both by the engineers who will fix the issues and by a governance function that maintains an AI risk register. What does each of those readers need from the same finding, and what happens to the attack narrative in the register version?
answer
- two renderings, one evidence base
- engineering = reproduction, register = control
- narrative moves to the annex
- stable finding ID links both
- generate the row, don't retype it
basics
~20 sEngineers need reproduction: the prompt pattern, the target configuration, what the model returned, and the fix. The register needs a risk statement, which control it is evidence about, how bad it is, and an owner and date. The attack narrative usually shrinks to one line, so keep the detail in a linked annex.
solid answer
~40 sThey are two renderings of one evidence base, not two reports. The engineering rendering is reproduction-shaped: which attack pattern, against which target configuration, what the response was, which detector or judge called it a hit, and what the remediation is. The governance rendering is control-shaped: a named risk, the control it bears on, the evidence date, a severity, and an accountable owner. Going from the first to the second, three things fall out: the multi-turn narrative, the specific probe or seed prompt that worked, and usually any numeric success rate. That is acceptable *if* the register row carries a stable finding identifier back into the technical annex. The failure mode is writing the two independently — six weeks later they disagree about what was tested, and nobody can tell which is right.
go deeper
Names the two audiences and says the register needs risk and control language while engineers need reproduction detail.
Adds that the unit differs — tool hits versus triaged findings versus register rows — and insists on a stable identifier linking the two documents.
Treats the register row as a coverage claim, so insists the scope sentence and assessment date travel with it, and builds the row from the finding record rather than by hand.
Sets the organisational contract: which team owns which rendering, how re-tests refresh evidence, and what a row is allowed to imply about untested surface.
### Three vocabularies, three units A **hit** is one row of tool output: one attempt against one target where a detector or judge decided the response counted as a success. A **finding** is what triage produces from those rows: one distinct weakness, evidenced by one or more hits. A **risk register row** is a governance object — a risk statement or a control statement carrying an owner, a severity, a treatment decision and a review date. A **control** is a stated safeguard somebody is accountable for; a **risk register** is the maintained list of those statements; the **technical annex** is the restricted companion document that holds the evidence. Nothing converts these units automatically. Each boundary is a human judgement, and each is where meaning leaks. ### Why the volume collapses so far Scanner output is multiplied by design. garak's `--generations` flag (short form `-g`) sets how many completions the harness requests per prompt, and each garak probe class contributes many prompts, so one `garak --model_type ... --probes ...` invocation can emit thousands of attempt rows against a single model. promptfoo's `redteam` run multiplies its `plugins` axis by its `strategies` axis and emits a row per generated test case. Four thousand attempt rows triaging down to nine findings and landing as three register rows is an entirely ordinary shape. All three numbers are true; none is a translation of another. ### What each rendering owes its reader | | engineering rendering | register rendering | |---|---|---| | unit | one distinct weakness | one risk or control statement | | target | system prompt version, tool and function permissions, retrieval corpus, model identifier, any input or output guard in front | named system, in business terms | | evidence | attack family, a described example, which mechanism called it a hit | a stable finding identifier pointing at the annex | | verdict | reproducible, re-testable | severity, treatment decision, owner, date, scope | The engineering rendering must name the mechanism that decided a response counted — which garak detector, which PyRIT scorer, which promptfoo grader — because that mechanism has its own false positives, and an engineer who cannot see it cannot argue with the hit. The register rendering must carry a scope sentence, because a row without one silently claims the whole system. ### What the conversion costs This is real work, not formatting. Triage is usually the largest engineer-time line in an AI engagement: a human has to read response text, and deduplicating a few hundred surviving hits into findings is on the order of a working day. Writing the register rows on top adds an hour or more per finding once you chase down which control it bears on and who owns it. That cost is exactly why teams retype the summary instead of generating it — and retyping is what makes the two documents disagree six weeks later. ### Where the number misleads Three specific misreadings, in order of how often they bite. *The count.* "Nine AI risks open" is an artefact of the grouping rule triage used. A team that groups by attack family produces a smaller number than a team that groups by affected component, from byte-identical evidence. Counts are therefore not comparable across teams, tools, or quarters unless the grouping rule is written down beside them — and a quarter-on-quarter drop from nine to five may be a re-grouping, not a remediation. *The implied coverage.* A populated row reads as an assessed row. If the engagement exercised one chat interface of a platform with four, a register that is fully populated tells a committee the platform was assessed. Nothing in the row is false; the artefact as a whole is. *The severity.* Tool output has no notion of business impact. A detector firing says a pattern matched, not that a customer was harmed. Severity is assigned at triage by a human, and a register row that carries a severity without recording who assigned it and against what impact scale is an opinion wearing a number. ### What you would check before either document ships Every register row traces to at least one finding identifier. Every finding either appears in the register or is explicitly marked out of governance scope — silence is not a decision. The scope sentence names what was *not* exercised, and it travels attached to the view rather than buried in an annex. The register layer was generated from the finding records rather than retyped, so the two cannot drift. And the widely circulated document contains no copyable attack string: describe the pattern, point at the restricted annex, and let distribution control the difference.
- A register row says the control passed. What minimum fields make that statement checkable a year later?The date, the scope of what was exercised, the target configuration at the time, the evidence identifier pointing at the finding record, and the accountable owner. Without the scope and date the row is an opinion.
- Should the technical annex include the exact working prompt?In a controlled artefact with restricted distribution, yes — it is the reproduction evidence. In a board-circulated document, no; describe the pattern and point at the restricted annex.
saying these in an interview costs you the question
- Treating the governance version as a shortened copy-paste of the technical report.
- Dropping the assessment scope, so a register row silently implies the whole system was tested.
- No identifier linking a register row back to the evidence it came from.
- Putting a copyable working attack string into a widely circulated executive document.