When collapsing a red-team scan's flagged attempts into report items, what is the difference between grouping by the attack template that produced a hit and grouping by the behaviour the model produced, and what does each key hide?
answer
- technique axis vs behaviour axis
- many-to-many grid, projected to a list
- template key hides breadth of harm
- behaviour key hides breadth of route
- carry the collapsed axis as a count
basics
~20 sGrouping by template says how the model was pushed; grouping by behaviour says what it did. The template key hides that one technique unlocked several unrelated harms. The behaviour key hides that one harm is reachable by many independent routes, so a fix aimed at one route leaves the rest open.
solid answer
~60 sThe two keys answer different questions and hide different things. **Template key** — one item per attack technique family. It reads well as an attacker narrative and maps cleanly onto "which technique should we build a defence against". What it hides is breadth of consequence: a single item can contain a policy-boundary breach, a data leak and an unsafe tool call, and whoever reads the item's title fixes only the one in the exemplar. **Behaviour key** — one item per thing the model actually did. It reads well as a risk statement and maps onto the control that failed. What it hides is breadth of reachability: an item can be reachable through six unrelated routes, and a fix built against the attached exemplar closes one of them. In practice mature reports key on the pair and let one side dominate the title. If the audience is the safety-policy owner, behaviour leads; if it is the platform team building an input defence, technique leads. Whichever you pick, the item must carry the other axis inside it as a count, otherwise the report is quietly claiming a one-to-one relationship that the run did not show.
go deeper
Should at least recognise that 'how it was attacked' and 'what the model did' are different groupings and give an example of each.
Should name what each key hides — breadth of harm versus breadth of route — and pick a default with a reason.
Should describe the composite key, choose the dominant axis by audience, and insist that the collapsed axis survives as a count with per-route exemplars.
Sets the house convention so items are comparable across engagements and re-runs, and treats a single-exemplar item on a multi-route finding as a review defect.
### The grid, and why a report is a projection of it A scan's flagged attempts naturally form a two-axis grid. On one axis are **attack techniques** — the families of manipulation the tool was configured to try, each authored once and expanded into surface variants. On the other axis are **elicited behaviours** — what the model actually produced when a technique worked: a policy-boundary breach, an internal identifier disclosed, an unsafe tool call issued, instructions from retrieved content obeyed. Any given attempt is one cell: this technique produced this behaviour on this surface. The relationship between the axes is many-to-many. One technique often unlocks several unrelated behaviours; one behaviour is usually reachable by several unrelated techniques. A report, however, is a list. Deduplication is therefore a projection of a grid onto one axis, and a projection always loses the axis you did not pick. The whole question is which loss you can afford in front of this reader. ### Projecting onto the technique axis Items read as "technique family X defeats the refusal boundary". This is the attacker's narrative and it is the shape an engineer can staff, because input-side defences — classifiers, pattern rules, prompt hardening — are usually built per technique. What it hides is **breadth of consequence**. One item can contain a policy breach, a data disclosure and an unsafe tool call. The item carries one title, one severity and, in most reports, one exemplar. The severity is set by whichever member the analyst attached, and the other behaviours inside the bucket silently inherit a rating nobody assigned to them. A reader who fixes what the title says has fixed a third of the item. ### Projecting onto the behaviour axis Items read as "the assistant will emit an internal identifier when asked in the right shape". This is the risk owner's language: it is a statement about outcome, it maps onto a named control, and it is the form in which a product or safety owner can accept or reject the risk. What it hides is **breadth of reachability**. An item may have been reached through six unrelated routes and different entry surfaces, but if it ships with one attached exemplar the reader reasonably assumes that route is the route. A patch is built against it, it passes, and the re-run fails identically through route two. This is the single most common way a red-team report produces a fix that does not fix anything, and the report — not the developer — is at fault, because the information that five other routes existed was in the analyst's hands and did not make it into the item. ### The composite key that actually holds up Mature reports key on the pair and let one axis dominate the title according to audience: ```text item_key = (behaviour, failed_control) # risk-owner audience item_body = { routes_observed: 6, technique_families: [...], per_route_exemplar: {...}, attempts: 480, flagged: 61 } ``` and the same structure inverted — technique in the key, behaviours enumerated in the body — when the reader is the platform team building an input defence. Two disciplines make it survive review: every item states the count on the axis it collapsed, and no item ships with a single exemplar when its own body says several routes reached it. If you cannot state both counts, you have not deduplicated; you have sampled. ### What it costs The composite key is not free. Assigning `failed_control` per attempt is a human judgment — the scanner does not know which of your defences was supposed to catch a prompt — and it runs at roughly a minute per member sampled, so a run with hundreds of flagged attempts is a half-day of analyst time even when you read only two or three members per bucket. Per-route exemplars multiply the evidence package a developer must be handed: six routes means six transcripts to capture, store and keep readable. The cost you are buying down is a wasted remediation cycle, which is days of engineering plus a re-test plus the credibility hit of a fix that did not hold. ### Where the number misleads Neither projection is reliably larger than the other, so the item count carries no information about which key was used — a reader cannot infer the shape of the report from its size. Worse, the two keys make the *same run* look like different engagements: a technique-keyed report of nine items and a behaviour-keyed report of nine items describe the same 480 attempts while telling incompatible stories about what to fix first. And when a report mixes the two keys across sections without saying so, the counts stop being addable at all; a "23 findings" summary line built from two different projections is not a quantity of anything. ### What you check For any item you are about to publish, ask what the collapsed axis count is. If the answer is "I do not know", open the members and count. Then check the exemplar rule: one exemplar per route on the axis you collapsed, or an explicit statement in the item that only one route was observed. Finally, confirm every section of the report used the same key, or label the sections so nobody adds their counts together.
- Which key would you default to for a report going to a product safety owner?Behaviour, with technique families carried inside the item. The safety owner is deciding whether the outcome is acceptable, not which defence to build.
- What single field most often prevents the 'fixed it, still fails' outcome?The count of distinct routes observed for the item, with one exemplar per route rather than one exemplar per item.
Grouping the same run by technique or by behaviour is like photographing one object from two sides: each picture is accurate, and each flattens away the dimension the other one shows. The item has to carry a count on the flattened dimension or the reader thinks the picture is the object.
saying these in an interview costs you the question
- Insisting one of the two keys is universally correct.
- Keying on technique and then rating the whole item by its worst exemplar without saying so.
- Attaching one exemplar to an item the analyst knows was reached six ways.
- Believing the scanner's own category labels are the grouping key.
- Being unable to state what the chosen key threw away.