An automated LLM red-team scan finishes with 480 rows flagged as hits against one chat endpoint. Why is that not 480 report findings, and what does a grouping key do?
answer
- tool unit = attempt, report unit = defect
- collapse on a key, not on prompt text
- variants plus retries inflate rows
- one item = one fix
- publish the key with the count
basics
~20 sBecause most rows are one weakness repeated: a single attack template retried with small wording changes, and each case re-sent several times. A grouping key is the field combination you collapse rows on, such as technique plus the behaviour elicited, so each report item is one distinct failure a developer fixes once.
solid answer
~50 sA scanner's unit of output is an **attempt**: one prompt sent to one endpoint, labelled hit or miss by whatever object decides that (a detector, a judge model, a regex). A report's unit is a **defect a developer can act on**. Those are different granularities, and the mapping between them is a choice you make, not a number the tool hands you. Deduplication is that mapping. You pick a grouping key — commonly the attack technique family, the control that failed, and the behaviour the model actually produced — and collapse every attempt sharing that key into one item, keeping counts and exemplars inside it. The tradeoff runs in both directions. Too coarse ("jailbreak") and a genuinely separate failure disappears inside a bucket nobody reads to the bottom of. Too fine (one item per prompt string) and the count inflates past reading, so severity gets judged by volume instead of impact. State the key you used in the report; an unexplained count is the thing reviewers distrust first.
code
text · 8 lineskey = (failed_control, observed_behaviour, entry_surface)
item = {
key: key,
attempts: 480,
hits: 61,
variants_that_worked: 9,
exemplar: strongest_attempt,
}go deeper
Should say the rows include repeated retries and template variants, so hits and findings are not the same unit, and that you group before reporting.
Should propose a concrete key, explain that prompt text is the wrong axis, and name both failure directions of a bad key.
Should describe verifying the key by reading bucket members, keeping counts and exemplars inside the item, and publishing the key so the count is auditable.
Frames the count as an editorial decision the team owns, sets a house key so reports across engagements are comparable, and expects the key stated in every deliverable.
### The unit mismatch A scanner's row is an **attempt**: one prompt sent once to one endpoint, plus a verdict from whatever object was asked to judge the response — a string or regular-expression detector, a trained classifier, or a judge model prompted to score the text. A row marked `hit` asserts exactly one thing: that this judging object said yes about this response. It asserts nothing about whether the weakness behind this row differs from the weakness behind the row above it. A report's row is a **defect**: something a named owner can change once, after which the failure stops. Converting 480 of the first kind into some number of the second kind is an editorial act performed by an analyst. No scanner performs it, and no scanner holds the information needed to — it does not know which of your controls was supposed to stop the prompt, or which change would close the gap. ### Where a number like 480 comes from Three multipliers are built into how these runs are configured, and every one of them is deliberate: | multiplier | why the tool does it | typical effect | |---|---|---| | template expansion | one technique family is authored once and machine-expanded into many surface variants, because a defence that only pattern-matches the original wording would look robust when it is not | x10 to x100 | | repeat trials | the endpoint samples its output, so the same case is sent several times (the run's repeat / generations setting) before you can say anything about whether it fails | x3 to x20 | | overlapping checks | several detectors score the same response; two of them firing on one response emits two rows | x1 to x3 | Eight technique families, twenty variants each, three sends per variant is 480 attempts before a single distinct weakness has been established. So 480 rows is entirely consistent with eight defects — and also with eighty. The rows do not distinguish those two worlds; only reading the outputs does. ### What a grouping key is A grouping key is the tuple of fields you declare to mean "same defect". Every attempt whose tuple matches collapses into one report item, and the rest of the attempt survives inside that item as a member. A key that holds up in review is built from things a remediation would change: ```text key = (failed_control, observed_behaviour, entry_surface) ``` - **failed_control** — which of your defences was supposed to stop this and did not: an input classifier that never fired, a system-prompt boundary that fired and was talked past, an output filter, a tool-call permission check. - **observed_behaviour** — what the model actually produced, read from the response, not the label the detector printed. - **entry_surface** — where the input came in: a direct chat turn, retrieved document content, a tool result, a file upload. Prompt text is deliberately excluded, because varying prompt text is precisely what the scanner was built to do. Keying on it guarantees one item per attempt and produces the raw log with headings. ### What it costs The scan is the cheap half. Four hundred and eighty attempts against a metered chat endpoint is single-digit dollars and tens of minutes of wall clock, and it runs unattended. Triage is the expensive half and it is paid in the scarcest budget on an engagement: analyst hours. At thirty seconds per row, reading all 480 is four hours, which is why the discipline is to sample two or three members per bucket rather than read exhaustively. Skipping the triage does not remove the cost — it transfers it to the developer who receives the log, and multiplies it by everyone who reads the report. ### Where the number misleads The count is a function of your key, not a property of the system. Two competent analysts working the same 480 rows can honestly publish eight items and forty, and neither is lying. This has two consequences that reviewers see constantly: - **An inflated count reads as severity.** A stakeholder skimming "480 issues" prices the risk by volume. If most of those rows are one template retried, the report has quietly claimed a scale the run never demonstrated. - **A collapsed count reads as safety, and ships an incomplete fix.** Eight tidy items look tractable; if one of them merged two defects with different owners, the developer fixes the exemplar attached and the re-run still fails. Because of that, a count published without its key is not a measurement. State the key in the report so a reader who disagrees can regroup rather than re-run. ### What you check before you believe your own count Open two or three of the largest buckets and read the actual model outputs of a sample of members — not their detector labels, which are often identical for materially different responses. Ask one question of each pair you read: would a single change close both? A yes anywhere in a bucket is fine; a no is proof the key is too coarse. Then check the opposite direction: if two buckets would be closed by the same change, they are one item. Finally, confirm the members survived the collapse with their counts intact, because the moment they are discarded nobody can recompute anything and the count becomes unfalsifiable.
- Where do the collapsed rows go — do you discard them?No. They stay as members of the item: attempt and trial counts, the strongest exemplar, and the span of variants that worked. Those are what let a reader regroup differently or recompute a rate.
- Why is the prompt string a bad grouping key?Because varying the prompt string is precisely what the scanner does. Keying on it guarantees one item per attempt and turns the report into the raw log with extra formatting.
- Two different checks fired on the same response. One finding or two?Normally one, if a single change fixes both. If the response breached two genuinely separate controls that different owners would fix, it is two items that cite the same exemplar.
saying these in an interview costs you the question
- Reporting the scanner's raw row count as the finding count.
- Deduplicating on the exact prompt string and calling it deduplication.
- Throwing away the collapsed members so no one can recompute anything.
- Choosing the key to make the number look impressive in either direction.
- Treating the grouping key as something the tool decided rather than something the analyst chose.