skip to content

An AI red-team programme reports a quarterly "findings" number obtained by summing every hit its automated LLM scanners emitted. Why is that number not the number of findings in the report, and what unit should be counted instead?

level: juniorimportance: must knowfreq 62%

answer

  1. hit = one attempt, not one issue
  2. probe x seed x sample x rerun
  3. dedup to behaviour + cause + owner
  4. count moves with the catalogue
  5. publish attempts and confirmed separately

basics

~20 s

A scanner hit is one prompt attempt its detector marked as failed. Many hits are the same behaviour repeated across prompt variants, probe classes and reruns, and some are detector false positives. The report's unit is a triaged, deduplicated issue with a cause. Count those, and report attempts separately as volume.

solid answer

~50 s

The tool's unit and the report's unit are different objects, and conflating them is the most common way a programme metric goes wrong in its first quarter. An automated LLM scanner emits one record per attempt that its detector judged a failure. A single underlying weakness — say, one weak refusal boundary — produces a hit for every paraphrase in the probe's seed set, for every temperature sample, and again on every scheduled rerun. Multiply that by however many probe classes touch the same behaviour and the count moves with the catalogue, not with the system. The reportable unit is a triaged item: deduplicated to one behaviour, attached to a cause and an owner, and confirmed by a human against the raw transcript because detectors misjudge in both directions. So publish two numbers with distinct names: attempts (volume, a cost and coverage signal) and confirmed issues (the risk signal). Never let one label carry both.

go deeper

for a junior

Says a hit is one attempt, that the same problem repeats across variants and reruns, and that the report should count deduplicated confirmed issues.

for a middle

Adds the cross-product that generates the hits, that detectors have false positives in both directions, and publishes attempts and confirmed issues as separately named numbers.

for a senior

Adds precision of the automated stage as its own metric, explains why the count moves with catalogue configuration, and refuses quarter-on-quarter comparisons across a config change.

for a principal

Frames it as a measurement-validity problem: any headline number whose error rate nobody has sampled cannot be defended when it is challenged in a review.

**What a "hit" is, mechanically.** An automated LLM scanner has three moving parts, and the headline number falls out of the third. A *probe* is the scanner's unit of test generation: a class that emits seed prompts for one behaviour it wants to elicit. A *target wrapper* sends each prompt to the system under test and captures the reply. A *detector* — the scanner's automated judge, which may be a string or regex rule, a small trained classifier, or another model asked to grade the reply — reads the response and returns pass or fail. A hit is one (prompt, sample, response) triple that the detector marked fail. The scanner writes one line per attempt to its run log and one line per failure to its hit log, and the number that gets quoted upward is, almost always, the hit log's line count. **The arithmetic that manufactures the number.** Hits are the product of a cross-product and a rate: `hits ~= probe classes enabled x seed prompts per probe x samples per prompt x reruns x observed fail rate` Concretely: 20 probe classes at 25 seed prompts each, sent 5 times apiece, is 2,500 attempts in one run. Swap the curated subset for the scanner's whole catalogue and the first factor triples. Raise the scanner's samples-per-prompt setting from 5 to 20 and everything downstream is multiplied by four. Add a second scheduled run in the quarter and it doubles again. None of those edits touched the system under test. Nothing in the pipeline deduplicates by *behaviour* either, because the scanner compares strings and detector verdicts: it has no way to know that forty failures spread across four probe classes all trace to one permissive clause in one system prompt. **What the run costs.** A mid-sized sweep is 2,500-10,000 calls to the target; if the detector is itself a model, roughly double that in grader calls. On a small hosted model that is single-digit to low-tens of dollars; with a long system prompt, retrieval context or an agent loop it is an order of magnitude more, and wall-clock runs 20-90 minutes at the concurrency a rate limit allows. The dominant cost is human and it is set by the hit count: 400 hits read at about two minutes of transcript each is roughly thirteen engineer-hours of triage, every cycle. That is the line item the metric quietly commits the team to. **Where the number misleads.** Three separate failures, and they do not cancel. - *It is configuration-elastic.* It rises with catalogue breadth, samples per prompt and rerun frequency. A quarter-on-quarter comparison across any of those edits is measuring the instrument, and it will read as the product getting worse. - *The target is stochastic.* The same prompt fails on one sample and passes on the next, so a decoding-temperature change or a silent provider model update moves the count with no weakness created or removed. - *The detector errs in both directions.* Refusal-phrase detectors mark a hedged, harmless non-answer as a failure, and mark a fluent, on-topic harmful completion that never used a refusal phrase as a pass. So the hit count is neither an upper bound (false positives inflate it) nor a lower bound (false negatives deflate it) on the number of real issues. A scanner that also reports a relative score — your model against a bag of calibration models — is describing a peer distribution, not your risk appetite; being average among peers is not a control. **What to publish instead.** Three separately named numbers, never merged into one word. *Attempts run* — volume and spend, which explains the query-budget line. *Confirmed issues* — human-triaged, deduplicated to one behaviour with one cause and one owner. *Precision of the automated stage* — confirmed issues over hits triaged, which is what tells you whether the detector is worth its triage hours. Stamp all three with a configuration fingerprint: probe list, samples per prompt, target model version, detector version. **What to check before believing any of it.** Draw a stratified sample of about thirty hits and thirty passes and read the raw transcripts. The hits give you an observed precision; the passes are the only place a false negative can be caught, and skipping them is how programmes discover a year later that their detector never fired on the failure mode that mattered. Until that sample exists, the count has an unknown error rate, and a number with an unknown error rate cannot be defended the first time someone in a funding review pushes back on it.

  • Two hits came from different probe classes but the same weak refusal boundary. One issue or two?
    One, if the fix is the same fix. Deduplicate by cause and owner, and note that two probe classes reached it — that is useful evidence the behaviour is easy to reach, and it belongs in the item, not as a second item.
  • How would you report a behaviour that reproduces on only three of ten samples of the same prompt?
    As one issue with a reproduction rate attached. Probabilistic reachability is part of the finding, not a reason to split it or to drop it; the rate is what the fix has to move.

saying these in an interview costs you the question

  • Treating the scanner's summary count as the finding count in the executive report.
  • Comparing this quarter's count to last quarter's after changing which probe classes were enabled.
  • Never reading raw transcripts, so the detector's false-positive and false-negative rates are unknown.
  • Deduplicating by prompt string rather than by underlying behaviour.

context