Nothing on a typical AI red-team dashboard says what was never tested. How would you put the untested part on the sheet, and why must any "coverage" percentage you publish state its denominator?
answer
- absence of testing looks like a clean result
- not-tested list on the same page
- reason codes: unauthorised, no harness, budget, by design
- three denominators: catalogue, surface, behaviour
- date the denominator; surfaces keep being added
basics
~20 sPublish an explicit not-tested list beside every result: surfaces, languages, modalities and attack classes nobody ran, with a reason for each. Coverage needs a denominator because the word means at least three things — probes run out of those available, attack surface reached, and behaviours exercised from a benchmark. A number without one is unreadable.
solid answer
~60 sEvery result page a red team publishes is silently a statement about a scope, and readers substitute their own. If you tested one chat surface in one language and the sheet says 92% coverage, the reader hears "the product is 92% tested". So make the untested part a first-class artefact. Beside the results, list what was out of scope and why: surfaces (the agent's tool-calling path, the batch pipeline, the file-upload route), languages, modalities, attack classes, and anything skipped for budget or authorisation reasons. Reason codes matter — "no authorisation from the provider", "budget exhausted", "no harness exists" are three different asks. On the denominator: "coverage" in this work names at least three different ratios — probe classes executed over probe classes available in a catalogue, attack surface reached over surface enumerated in the threat model, and behaviours exercised over behaviours defined in a benchmark. They can disagree wildly on the same engagement. Always write the ratio as a fraction with both terms named, and note who enumerated the denominator, because a surface nobody wrote down cannot appear as untested either.
go deeper
Says the report should state what was in and out of scope so a clean result is not read as a clean product.
Names the distinct denominators behind the word coverage and writes figures as an explicit fraction.
Builds the not-tested list with reason codes, drives it from the threat model's enumeration, and refuses a bare percentage.
Owns the enumeration and its refresh cadence, and treats the budget-exhausted subset as the programme's funding case.
**The asymmetry that makes this the worst failure in red-team reporting.** A programme can only report on what it ran. On a dashboard, a surface that was never tested and a surface that was tested and came back clean look identical: both contribute no findings. Every result page is therefore a silent statement about a scope, and readers substitute a scope of their own — usually "the product". The fix is not statistical, it is layout: the untested part goes on the same page as the results, never in an appendix. **What the not-tested list holds.** One line per omitted thing — surface, language, modality, attack class — with a reason code, because the reason routes the item to a different owner and a different ask. - *Not authorised* — a hosted third-party dependency whose terms or owner forbade testing. This is a risk-acceptance decision for someone else; naming it is what moves it. - *No harness* — no way to drive the surface programmatically yet. An engineering ask with a price: a tool-calling agent path typically needs a stateful driver plus per-turn scoring, which is engineer-weeks, not an afternoon. - *Budget exhausted* — the query budget or the engagement hours ran out. A funding ask, and the strongest evidence a funding review will ever see, because the cost per attempt is already measured. - *Out of scope by design* — deliberately excluded; record who decided and when. **Why the denominator is not pedantry — one engagement, three answers.** At least three different ratios travel under the word coverage, and they can disagree by an order of magnitude on the same work. | ratio | denominator, and who owns it | same engagement | |---|---|---| | catalogue coverage | probe classes in the scanner's catalogue; owned by the catalogue's maintainers | 120 of 140 = 86% | | surface coverage | attack surfaces enumerated in the threat model; owned by architecture | 1 of 9 = 11% | | behaviour coverage | behaviours a benchmark defines; owned by the benchmark's authors | 60 of 200 = 30% | Same run, and 86% is the one that is cheap to compute and therefore the one that gets published. That is exactly the misreading: a full catalogue sweep against a single chat endpoint reports near-total coverage while the agent's tool-calling path, the batch summarisation job, the retrieval ingestion path and the file-upload route were never touched. **The other ways a coverage figure misleads.** Its denominator is a moving object owned by someone else. When catalogue maintainers add sixty probe classes, your catalogue coverage falls although you ran identically — and conversely, pinning an old catalogue version is a way to hold the number up while doing nothing. A surface enumeration goes stale the moment a product team ships anything, so a percentage computed against last quarter's list overstates itself in the safe direction, which is the dangerous direction. And coverage says nothing about depth: a probe class run once at one sample counts as covered, even though a behaviour that reproduces one time in fifty was never going to appear at that sample count. Finally, a surface nobody enumerated cannot appear as untested either — the not-tested list is only as honest as the enumeration behind it, which is why the enumeration needs an owner and a refresh date of its own. **Two habits that keep it honest.** Write every coverage figure as "x of y <named things>" so the denominator travels with the number and cannot be quoted alone. And date and version the denominator — catalogue version, threat-model revision, benchmark release — so a reader can tell a real change from a denominator change. **What to check.** Ask whoever published the percentage to name the members of the denominator; if nobody can list them, the ratio is decoration. Confirm the surface enumeration has been refreshed since the last product release, and diff it against what actually shipped. For every surface marked covered, check the sample count behind it was enough to see a rare behaviour at all. And confirm every not-tested line has both a reason code and a named owner, because a gap with no reason cannot be told apart from a funding gap, and a gap with no owner will still be there next quarter.
- Which coverage denominator would you show a product owner deciding whether to launch?Surface coverage against the threat model's enumeration, with the untested surfaces named. Catalogue coverage answers a question they did not ask, and behaviour coverage is bounded by someone else's benchmark scope.
- How can an untested surface be missing from the not-tested list at all?If nobody enumerated it. The list is built from the threat model, so an undocumented surface is invisible in both directions — which is why the enumeration itself needs an owner and a refresh date.
- How does the not-tested list help in a funding review?Items with a budget-exhausted reason code convert directly into an ask, with the surface named and the cost per attempt already measured. It is the most concrete argument the programme has.
saying these in an interview costs you the question
- Publishing a bare coverage percentage with no denominator named.
- Putting scope limitations only in an appendix nobody reads.
- Treating a full sweep of a probe catalogue as evidence the architecture was covered.
- Computing coverage against a surface enumeration that has not been refreshed since the last engagement.
- No reason attached to an untested surface, so nobody can tell a funding gap from an authorisation gap.