An AI red-team deliverable states "78% coverage" with no other qualifier. What are at least three different denominators that percentage could be measuring, and what would you write in the report instead?
answer
- catalogue vs surface vs behaviour
- denominator names the thing counted
- surface list comes from architecture walkthrough
- attempts budgeted is throughput, not coverage
- breadth needs a depth figure beside it
basics
~20 sCoverage has no meaning without its denominator. It can mean test families executed out of those the harness offers, attack surfaces reached out of the application's entry points, or harm categories exercised out of a taxonomy. Write the fraction with both numerator and denominator named, not a bare percentage.
solid answer
~50 sThe same engagement can honestly print three very different percentages, because three unrelated denominators all get called coverage: 1. **Catalogue coverage** — test families or attack techniques executed, out of those your tooling ships. Easy to compute, and the least meaningful to a customer: it measures your tool, not their system. 2. **Surface coverage** — entry points reached, out of the application's real ones: the chat UI, the API, retrieval ingestion, uploaded files, tool and function calls, any agent-to-agent hop. 3. **Behaviour coverage** — harm categories or behaviours exercised, out of a named taxonomy or the customer's own risk register. A run that fires every shipped test family at one chat endpoint is 100% on the first and can be under 30% on the second. So the report should never print a bare percentage: write each fraction with its denominator spelled out, and put the surface one first, because that is the one the reader's intuition is reaching for.
code
text · 12 linesCOVERAGE DECLARATION
Surface 5 of 9 entry points driven
driven chat UI, public API, document ingestion, file upload, tool results
not driven email intake (no test mailbox), scheduled agent runs (no auth),
image input (out of scope), inter-agent channel (not provisioned)
Behaviour 11 of 14 risks in the customer register exercised
depth min 30 attempts per risk; 3 risks exercised single-turn only (flagged)
not exercised R-07 training-data concerns (out of scope), R-12, R-13 (hours)
Throughput 4,180 of 6,000 budgeted attempts completed; 612 unscored (see 4.3)go deeper
Recognises that a percentage needs a denominator and asks which one was meant before quoting the number.
Names catalogue, surface and behaviour, and explains why the catalogue number is the flattering one.
Enumerates the surface denominator from an architecture walkthrough rather than from the tool, and pairs breadth with depth per item.
Standardises which denominators the firm reports, so numbers are comparable between engagements and cannot be gamed by swapping tooling.
“Coverage” is the word in a red-team report most likely to be read as an assurance number and least likely to be defined. Three unrelated denominators all answer to it, they move independently, and only one of them is free to compute — which is why that one ends up in the summary. **Denominator 1 — catalogue.** Numerator: the probe sets, plugins or attack techniques you actually ran. Denominator: those your harness ships. `garak --list_probes` prints the module list, and `garak --probes` selects the subset you drive; `promptfoo`'s `redteam.plugins` array selects from that tool's plugin catalogue. Both tools can therefore hand you a percentage for nothing. Its flaw is that the denominator describes your tooling's opinions about attacks, not the customer's exposure. **Denominator 2 — surface.** Numerator: entry points actually driven. Denominator: entry points the system exposes — the chat UI, the programmatic API, retrieval ingestion, file and image inputs, tool or function results fed back into context, memory carried between sessions, scheduled or agentic paths no human ever types into. No tool computes this. It comes out of an architecture walkthrough with somebody who knows the system, which is why it gets skipped and why it is the denominator that matters: most real incidents arrive through a surface nobody drove. **Denominator 3 — behaviour.** Numerator: harm classes or behaviours exercised. Denominator: a cited taxonomy, or better, the customer's own risk register. Two failure modes here: a taxonomy that does not contain the customer's actual worry, and marking a category covered on the strength of one attempt. There is a fourth number people mistake for coverage: attempts completed out of attempts budgeted. That is throughput. Report it, separately, and never as coverage. **What each one costs** | Denominator | Cost to produce | Cost to improve | |---|---|---| | Catalogue | zero — the harness prints it | run more of the same catalogue: model calls only | | Surface | 2–3 hours of an architect's time, plus the nerve to write down entry points nobody owns | often a harness adapter — a `garak` generator or a `promptfoo` provider for the untouched path — which is engineer-days, not calls | | Behaviour | half a day mapping register risks to probes or plugins | more attempts per risk: linear in `promptfoo`'s `redteam.numTests` or `garak`'s `--generations`, so linear in spend | That table is the whole reason reports drift towards the catalogue number: it is the only one you can raise by spending money instead of engineering time. **Where the number misleads** Three specific misreadings. First, the denominator moves when your toolbox does: add a second scanner with a bigger probe list, run the identical attacks, and your catalogue percentage falls — the metric tracks inventory, not exposure, and any metric that can be improved by uninstalling a tool is not measuring the system. Second, volume masquerades as breadth. `promptfoo` has two catalogue axes: `redteam.plugins` decides which harm is attempted, `redteam.strategies` decides how the attempt is encoded. Doubling `strategies` doubles generated cases and the deck reads “twice the coverage”, but the set of behaviours attempted is unchanged — that is depth per behaviour, which is worth having and is not breadth. Third, a run that fires every shipped test family at one chat endpoint is honestly 100% on catalogue and can be under 30% on surface. Both sentences are true of the same run, and only one of them is what the reader thinks they are being told. **What to write instead of “78%”** ```text Surface 5 of 9 entry points driven (4 not driven, each with a reason) Behaviour 11 of 14 register risks exercised (min 30 attempts each; 3 single-turn only) Catalogue 22 of 28 available test families run Throughput 4,180 of 6,000 budgeted attempts completed (612 unscored, see 4.3) ``` Put the surface line first: it is the one the reader's intuition is reaching for. **What I check** Can I reconstruct every fraction from the appendix, and does it match? Is each denominator named in words rather than implied? Does any breadth figure appear without a depth figure beside it — because a category counted once on one attempt and one counted after a long multi-turn campaign contribute identically to the numerator and are not the same evidence? And would this percentage survive the firm swapping tools next quarter, or would it move without the system moving?
- Why does adding a second, larger tool to your workflow lower catalogue coverage while raising real coverage?The catalogue denominator grew faster than the number of families you ran, so the percentage falls even though you exercised more distinct attacks. That inversion is the clearest argument against reporting the catalogue number.
- How do you enumerate the surface denominator for an agentic application?Walk the architecture and count every path where untrusted text can enter context: user turns, retrieved documents, tool and API responses, file or image inputs, memory carried between sessions, and messages from other agents. Each is an entry point to be driven or declared undriven.
A bare "78% coverage" is a restaurant critic saying he has visited 78% of the restaurants in the guidebook — a guidebook he wrote himself. The number is real, but it measures his own list, not the city.
saying these in an interview costs you the question
- Quoting a coverage percentage produced by a tool without knowing what it divides by
- Using the harness catalogue as the denominator in a customer-facing summary
- Counting a behaviour as covered on a single attempt
- Confusing attempts completed out of attempts budgeted with coverage