skip to content

Measuring the Programme

The numbers a red-team programme reports upward decide whether it keeps its budget, and the easy ones reward noise over risk. Interviewers ask which metric you would defend in a funding review.

on this pageshow

explore

questions

5

An AI red-team programme reports a quarterly "findings" number obtained by summing every hit its automated LLM scanners emitted. Why is that number not the number of findings in the report, and what unit should be counted instead?

level: juniorimportance: must knowfreq 62%

answer

  1. hit = one attempt, not one issue
  2. probe x seed x sample x rerun
  3. dedup to behaviour + cause + owner
  4. count moves with the catalogue
  5. publish attempts and confirmed separately

basics

~20 s

A scanner hit is one prompt attempt its detector marked as failed. Many hits are the same behaviour repeated across prompt variants, probe classes and reruns, and some are detector false positives. The report's unit is a triaged, deduplicated issue with a cause. Count those, and report attempts separately as volume.

solid answer

~50 s

The tool's unit and the report's unit are different objects, and conflating them is the most common way a programme metric goes wrong in its first quarter. An automated LLM scanner emits one record per attempt that its detector judged a failure. A single underlying weakness — say, one weak refusal boundary — produces a hit for every paraphrase in the probe's seed set, for every temperature sample, and again on every scheduled rerun. Multiply that by however many probe classes touch the same behaviour and the count moves with the catalogue, not with the system. The reportable unit is a triaged item: deduplicated to one behaviour, attached to a cause and an owner, and confirmed by a human against the raw transcript because detectors misjudge in both directions. So publish two numbers with distinct names: attempts (volume, a cost and coverage signal) and confirmed issues (the risk signal). Never let one label carry both.

go deeper

for a junior

Says a hit is one attempt, that the same problem repeats across variants and reruns, and that the report should count deduplicated confirmed issues.

for a middle

Adds the cross-product that generates the hits, that detectors have false positives in both directions, and publishes attempts and confirmed issues as separately named numbers.

for a senior

Adds precision of the automated stage as its own metric, explains why the count moves with catalogue configuration, and refuses quarter-on-quarter comparisons across a config change.

for a principal

Frames it as a measurement-validity problem: any headline number whose error rate nobody has sampled cannot be defended when it is challenged in a review.

**What a "hit" is, mechanically.** An automated LLM scanner has three moving parts, and the headline number falls out of the third. A *probe* is the scanner's unit of test generation: a class that emits seed prompts for one behaviour it wants to elicit. A *target wrapper* sends each prompt to the system under test and captures the reply. A *detector* — the scanner's automated judge, which may be a string or regex rule, a small trained classifier, or another model asked to grade the reply — reads the response and returns pass or fail. A hit is one (prompt, sample, response) triple that the detector marked fail. The scanner writes one line per attempt to its run log and one line per failure to its hit log, and the number that gets quoted upward is, almost always, the hit log's line count. **The arithmetic that manufactures the number.** Hits are the product of a cross-product and a rate: `hits ~= probe classes enabled x seed prompts per probe x samples per prompt x reruns x observed fail rate` Concretely: 20 probe classes at 25 seed prompts each, sent 5 times apiece, is 2,500 attempts in one run. Swap the curated subset for the scanner's whole catalogue and the first factor triples. Raise the scanner's samples-per-prompt setting from 5 to 20 and everything downstream is multiplied by four. Add a second scheduled run in the quarter and it doubles again. None of those edits touched the system under test. Nothing in the pipeline deduplicates by *behaviour* either, because the scanner compares strings and detector verdicts: it has no way to know that forty failures spread across four probe classes all trace to one permissive clause in one system prompt. **What the run costs.** A mid-sized sweep is 2,500-10,000 calls to the target; if the detector is itself a model, roughly double that in grader calls. On a small hosted model that is single-digit to low-tens of dollars; with a long system prompt, retrieval context or an agent loop it is an order of magnitude more, and wall-clock runs 20-90 minutes at the concurrency a rate limit allows. The dominant cost is human and it is set by the hit count: 400 hits read at about two minutes of transcript each is roughly thirteen engineer-hours of triage, every cycle. That is the line item the metric quietly commits the team to. **Where the number misleads.** Three separate failures, and they do not cancel. - *It is configuration-elastic.* It rises with catalogue breadth, samples per prompt and rerun frequency. A quarter-on-quarter comparison across any of those edits is measuring the instrument, and it will read as the product getting worse. - *The target is stochastic.* The same prompt fails on one sample and passes on the next, so a decoding-temperature change or a silent provider model update moves the count with no weakness created or removed. - *The detector errs in both directions.* Refusal-phrase detectors mark a hedged, harmless non-answer as a failure, and mark a fluent, on-topic harmful completion that never used a refusal phrase as a pass. So the hit count is neither an upper bound (false positives inflate it) nor a lower bound (false negatives deflate it) on the number of real issues. A scanner that also reports a relative score — your model against a bag of calibration models — is describing a peer distribution, not your risk appetite; being average among peers is not a control. **What to publish instead.** Three separately named numbers, never merged into one word. *Attempts run* — volume and spend, which explains the query-budget line. *Confirmed issues* — human-triaged, deduplicated to one behaviour with one cause and one owner. *Precision of the automated stage* — confirmed issues over hits triaged, which is what tells you whether the detector is worth its triage hours. Stamp all three with a configuration fingerprint: probe list, samples per prompt, target model version, detector version. **What to check before believing any of it.** Draw a stratified sample of about thirty hits and thirty passes and read the raw transcripts. The hits give you an observed precision; the passes are the only place a false negative can be caught, and skipping them is how programmes discover a year later that their detector never fired on the failure mode that mattered. Until that sample exists, the count has an unknown error rate, and a number with an unknown error rate cannot be defended the first time someone in a funding review pushes back on it.

  • Two hits came from different probe classes but the same weak refusal boundary. One issue or two?
    One, if the fix is the same fix. Deduplicate by cause and owner, and note that two probe classes reached it — that is useful evidence the behaviour is easy to reach, and it belongs in the item, not as a second item.
  • How would you report a behaviour that reproduces on only three of ten samples of the same prompt?
    As one issue with a reproduction rate attached. Probabilistic reachability is part of the finding, not a reason to split it or to drop it; the rate is what the fix has to move.

saying these in an interview costs you the question

  • Treating the scanner's summary count as the finding count in the executive report.
  • Comparing this quarter's count to last quarter's after changing which probe classes were enabled.
  • Never reading raw transcripts, so the detector's false-positive and false-negative rates are unknown.
  • Deduplicating by prompt string rather than by underlying behaviour.

context

open as a page

If an AI red-team programme's headline number is confirmed findings per quarter, which part of its automated testing gets more investment over time, and why is that the wrong part?

level: middleimportance: must knowfreq 55%

basics

~20 s

Whatever produces findings cheapest wins the budget. That is usually a high-yield, low-severity probe family against an easy surface. The hard work — a slow multi-turn attack on a tool-using agent — yields few items and looks unproductive. The programme drifts toward whichever instrument is noisiest, not toward where the risk is.

open as a page

Nothing on a typical AI red-team dashboard says what was never tested. How would you put the untested part on the sheet, and why must any "coverage" percentage you publish state its denominator?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Publish an explicit not-tested list beside every result: surfaces, languages, modalities and attack classes nobody ran, with a reason for each. Coverage needs a denominator because the word means at least three things — probes run out of those available, attack surface reached, and behaviours exercised from a benchmark. A number without one is unreadable.

open as a page

An AI red-team programme's dashboard leads with mean time to remediate, but the fixes are made by application teams and by a model provider the programme does not control. What breaks about that metric, and what would you put on the dashboard instead?

level: seniorimportance: should knowfreq 48%

basics

~20 s

It scores teams you do not manage, so it measures their backlog, not your work. Worse, many AI findings have no fix you own: the behaviour lives in a hosted model or a third-party guardrail. Report what you control — time to triage, time to notify, and the age and acceptance status of open items.

open as a page

You get three numbers on the slide that decides whether an AI red-team programme is funded next year. Which three do you choose, and how do you defend each against the objection that the team producing it can move it?

level: principalimportance: should knowfreq 40%

basics

~20 s

Pick numbers the team cannot inflate alone. One: issues whose cause was new to the risk register, severity-banded. Two: named untested surfaces with reason codes, especially budget-blocked ones. Three: count and maximum age of open items with no owner and no acceptance. Defend each by naming who verifies it outside the team.

open as a page