skip to content

Writing the Finding

A transcript is not a finding until someone else can re-run it, rate it and act on it without you in the room. Interviewers probe here because most candidates can attack and cannot write it up.

on this pageshow

explore

questions

21

An automated LLM red-team scan finishes with 480 rows flagged as hits against one chat endpoint. Why is that not 480 report findings, and what does a grouping key do?

level: juniorimportance: must knowfreq 62%

answer

  1. tool unit = attempt, report unit = defect
  2. collapse on a key, not on prompt text
  3. variants plus retries inflate rows
  4. one item = one fix
  5. publish the key with the count

basics

~20 s

Because most rows are one weakness repeated: a single attack template retried with small wording changes, and each case re-sent several times. A grouping key is the field combination you collapse rows on, such as technique plus the behaviour elicited, so each report item is one distinct failure a developer fixes once.

solid answer

~50 s

A scanner's unit of output is an **attempt**: one prompt sent to one endpoint, labelled hit or miss by whatever object decides that (a detector, a judge model, a regex). A report's unit is a **defect a developer can act on**. Those are different granularities, and the mapping between them is a choice you make, not a number the tool hands you. Deduplication is that mapping. You pick a grouping key — commonly the attack technique family, the control that failed, and the behaviour the model actually produced — and collapse every attempt sharing that key into one item, keeping counts and exemplars inside it. The tradeoff runs in both directions. Too coarse ("jailbreak") and a genuinely separate failure disappears inside a bucket nobody reads to the bottom of. Too fine (one item per prompt string) and the count inflates past reading, so severity gets judged by volume instead of impact. State the key you used in the report; an unexplained count is the thing reviewers distrust first.

code

text · 8 lines
text
key = (failed_control, observed_behaviour, entry_surface)
item = {
  key: key,
  attempts: 480,
  hits: 61,
  variants_that_worked: 9,
  exemplar: strongest_attempt,
}

go deeper

for a junior

Should say the rows include repeated retries and template variants, so hits and findings are not the same unit, and that you group before reporting.

for a middle

Should propose a concrete key, explain that prompt text is the wrong axis, and name both failure directions of a bad key.

for a senior

Should describe verifying the key by reading bucket members, keeping counts and exemplars inside the item, and publishing the key so the count is auditable.

for a principal

Frames the count as an editorial decision the team owns, sets a house key so reports across engagements are comparable, and expects the key stated in every deliverable.

### The unit mismatch A scanner's row is an **attempt**: one prompt sent once to one endpoint, plus a verdict from whatever object was asked to judge the response — a string or regular-expression detector, a trained classifier, or a judge model prompted to score the text. A row marked `hit` asserts exactly one thing: that this judging object said yes about this response. It asserts nothing about whether the weakness behind this row differs from the weakness behind the row above it. A report's row is a **defect**: something a named owner can change once, after which the failure stops. Converting 480 of the first kind into some number of the second kind is an editorial act performed by an analyst. No scanner performs it, and no scanner holds the information needed to — it does not know which of your controls was supposed to stop the prompt, or which change would close the gap. ### Where a number like 480 comes from Three multipliers are built into how these runs are configured, and every one of them is deliberate: | multiplier | why the tool does it | typical effect | |---|---|---| | template expansion | one technique family is authored once and machine-expanded into many surface variants, because a defence that only pattern-matches the original wording would look robust when it is not | x10 to x100 | | repeat trials | the endpoint samples its output, so the same case is sent several times (the run's repeat / generations setting) before you can say anything about whether it fails | x3 to x20 | | overlapping checks | several detectors score the same response; two of them firing on one response emits two rows | x1 to x3 | Eight technique families, twenty variants each, three sends per variant is 480 attempts before a single distinct weakness has been established. So 480 rows is entirely consistent with eight defects — and also with eighty. The rows do not distinguish those two worlds; only reading the outputs does. ### What a grouping key is A grouping key is the tuple of fields you declare to mean "same defect". Every attempt whose tuple matches collapses into one report item, and the rest of the attempt survives inside that item as a member. A key that holds up in review is built from things a remediation would change: ```text key = (failed_control, observed_behaviour, entry_surface) ``` - **failed_control** — which of your defences was supposed to stop this and did not: an input classifier that never fired, a system-prompt boundary that fired and was talked past, an output filter, a tool-call permission check. - **observed_behaviour** — what the model actually produced, read from the response, not the label the detector printed. - **entry_surface** — where the input came in: a direct chat turn, retrieved document content, a tool result, a file upload. Prompt text is deliberately excluded, because varying prompt text is precisely what the scanner was built to do. Keying on it guarantees one item per attempt and produces the raw log with headings. ### What it costs The scan is the cheap half. Four hundred and eighty attempts against a metered chat endpoint is single-digit dollars and tens of minutes of wall clock, and it runs unattended. Triage is the expensive half and it is paid in the scarcest budget on an engagement: analyst hours. At thirty seconds per row, reading all 480 is four hours, which is why the discipline is to sample two or three members per bucket rather than read exhaustively. Skipping the triage does not remove the cost — it transfers it to the developer who receives the log, and multiplies it by everyone who reads the report. ### Where the number misleads The count is a function of your key, not a property of the system. Two competent analysts working the same 480 rows can honestly publish eight items and forty, and neither is lying. This has two consequences that reviewers see constantly: - **An inflated count reads as severity.** A stakeholder skimming "480 issues" prices the risk by volume. If most of those rows are one template retried, the report has quietly claimed a scale the run never demonstrated. - **A collapsed count reads as safety, and ships an incomplete fix.** Eight tidy items look tractable; if one of them merged two defects with different owners, the developer fixes the exemplar attached and the re-run still fails. Because of that, a count published without its key is not a measurement. State the key in the report so a reader who disagrees can regroup rather than re-run. ### What you check before you believe your own count Open two or three of the largest buckets and read the actual model outputs of a sample of members — not their detector labels, which are often identical for materially different responses. Ask one question of each pair you read: would a single change close both? A yes anywhere in a bucket is fine; a no is proof the key is too coarse. Then check the opposite direction: if two buckets would be closed by the same change, they are one item. Finally, confirm the members survived the collapse with their counts intact, because the moment they are discarded nobody can recompute anything and the count becomes unfalsifiable.

  • Where do the collapsed rows go — do you discard them?
    No. They stay as members of the item: attempt and trial counts, the strongest exemplar, and the span of variants that worked. Those are what let a reader regroup differently or recompute a rate.
  • Why is the prompt string a bad grouping key?
    Because varying the prompt string is precisely what the scanner does. Keying on it guarantees one item per attempt and turns the report into the raw log with extra formatting.
  • Two different checks fired on the same response. One finding or two?
    Normally one, if a single change fixes both. If the response breached two genuinely separate controls that different owners would fix, it is two items that cite the same exemplar.

saying these in an interview costs you the question

  • Reporting the scanner's raw row count as the finding count.
  • Deduplicating on the exact prompt string and calling it deduplication.
  • Throwing away the collapsed members so no one can recompute anything.
  • Choosing the key to make the number look impressive in either direction.
  • Treating the grouping key as something the tool decided rather than something the analyst chose.

context

open as a page

You are handing an AI red-team finding to an application team that does not have your scanning harness installed. What must the evidence package contain so they can reproduce the failure themselves?

level: juniorimportance: must knowfreq 52%

basics

~20 s

Ship a self-contained reproduction: the exact request they must send, the endpoint and generation settings it was sent with, the response you observed, and a plain statement of what makes that response a failure. Add how often it happened out of how many attempts. No harness install, no red-team-only dependency.

open as a page

Your red-team harness logs only the attack prompt and the model's final reply for each attempt against a hosted chat endpoint. Why can a colleague not re-run that attempt from the log, and what should the log capture instead?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Prompt plus reply says nothing about how the reply was produced. Capture the endpoint and the model identifier the provider returned, the decoding settings you sent, the application's system prompt, every earlier turn, and a timestamp with the provider's request id. Without those, a colleague who reproduces nothing cannot tell drift from a fix.

open as a page

When collapsing a red-team scan's flagged attempts into report items, what is the difference between grouping by the attack template that produced a hit and grouping by the behaviour the model produced, and what does each key hide?

level: middleimportance: must knowfreq 50%

basics

~20 s

Grouping by template says how the model was pushed; grouping by behaviour says what it did. The template key hides that one technique unlocked several unrelated harms. The behaviour key hides that one harm is reachable by many independent routes, so a fix aimed at one route leaves the rest open.

open as a page

A developer runs the reproduction from your AI red-team finding once, gets a polite refusal, and closes the ticket as unreproducible — but the behaviour is real and intermittent. What should the evidence package have contained to prevent that outcome?

level: middleimportance: must knowfreq 47%

basics

~20 s

State up front that the target is sampled, so one attempt proves nothing. Ship the observed count out of total attempts, the sampling settings you used, a script that repeats the request and tallies how many responses met the criterion, and a written verification rule such as: any hit in twenty attempts is still a failure.

open as a page

In a finding against a hosted chat endpoint you write 'set temperature to 0' as the reproduction instruction. Why is that not the same as recording a seed, and what should the finding say about determinism instead?

level: middleimportance: must knowfreq 58%

basics

~20 s

Temperature 0 records what you sent, not a guarantee you get the same tokens back. A hosted endpoint gives no seed you control and contracts no bit-identical output. State the settings you used, how many attempts you ran, and how many succeeded, so a reader re-runs the attempt the same way you did.

open as a page

An automated red-team run got a hosted chat endpoint to return another user's private record in 3 of 20 attempts. When you write the severity, does the 3-in-20 belong in the impact score, in the likelihood/exploitability score, or in both?

level: middleimportance: must knowfreq 65%

basics

~20 s

Likelihood, not impact. Impact is what one success does, and one success already leaks the whole record — a 15% rate does not make the data 15% leaked. The rate belongs on the exploitability side, and even there it counts for little when an attacker can simply retry cheaply.

open as a page

Your success rate was measured by calling a model API directly with your harness's own system prompt. Production wraps that same model in a fixed template, a 300-character user field and an input classifier. What does that do to the severity you file?

level: seniorimportance: must knowfreq 50%

basics

~20 s

It means your rate describes your harness's entry point, not the product's. Say so explicitly in the finding, then re-measure through the production path before you finalise the rating. If you cannot re-measure, file the number with the measurement point stated and rate the reachability separately, rather than quietly discounting it.

open as a page

A red-team finding you are reviewing states only "attack success rate: 15%" for an injection attempt against a hosted chat feature. What else has to sit next to that number before anyone can rate the finding's severity?

level: juniorimportance: should knowfreq 50%

basics

~20 s

The counts behind it: how many attempts and how many succeeded, because 3 of 20 and 150 of 1000 are not the same evidence. Also what decided that a response counted as a success, which endpoint and configuration it was measured against, and when. Without those, 15% is a number nobody can re-derive.

open as a page

You reran the same 20-attempt red-team script against the same hosted chat endpoint the next day: the first run succeeded twice, the rerun six times. How should the finding's severity handle a measured success rate that moved from 10% to 30%?

level: middleimportance: should knowfreq 45%

basics

~20 s

Do not rerate on it. At 20 attempts both results sit inside the same wide confidence interval, so the swing is what sampling noise looks like, not evidence the target got worse. Report both runs with their counts, widen the band you quote, and if the rate really drives the score, collect a much larger sample.

open as a page

A red-team scan sends each attack case to the endpoint ten times because responses vary, and three of one case's ten attempts were flagged as hits. When you collapse those ten attempts into one report item, what must the collapsed record keep, and what breaks if you keep only the first hit?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Keep both numbers: three hits out of ten attempts, plus the exemplars and the non-hit responses. Keeping only the first hit destroys the denominator, so nobody downstream can say how often the failure happens, and a re-run after a fix has no baseline to compare against. One hit and ten of ten read identically.

open as a page

You collapsed a red-team scan's flagged attempts into eight report items, the developer fixed all eight, and a re-run of the same scan still fails. How do you tell whether your grouping key was too coarse, and what do you inspect inside a bucket to prove it?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Take the still-failing attempts from the re-run and find which of the eight items they were grouped under. Then open that bucket's members and compare what the model produced and which control let it through. If members differ on either, the key merged distinct defects and the fix only closed the exemplar you attached.

open as a page

In your AI red-team harness, whether an attempt counted as a failure was decided by an automated model-based grader with a threshold. The application team will re-run the reproduction without that grader. What do you put in the handoff so their verdict matches yours?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Translate the grader into something they can apply. Write the decision rule in plain words with its threshold, attach the grader's output on the cases you sent, include labelled examples just either side of the line, and give a deterministic check they can assert on. If the verdict truly needs the grader, ship it pinned as a small dependency.

open as a page

One attack family in your red-team run flagged forty-plus attempts against the same chat feature. You will not attach all forty to the developer ticket. How do you choose which attempts go into the evidence package and which stay out?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Pick the set that defines the fix boundary, not the prettiest example. One canonical reproduction, plus a few variants chosen because each defeats a different plausible cheap fix, plus benign inputs that must keep working after the change. Link the full log as an appendix and say why each attached case is there.

open as a page

Your multi-turn red-team harness generated each follow-up prompt with an attacker model rather than reading a fixed script, and it succeeded against a hosted chat endpoint. What goes in the finding so the result can be re-run, given that re-running the harness produces different prompts every time?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Ship the realised conversation, not the generator. Store every turn verbatim in order as a fixed replay script a reader can send as-is, plus the target's decoding settings and system prompt. Record the attacker-model configuration separately, as provenance for how the conversation was found, not as the reproduction steps.

open as a page

A red-team finding you filed three weeks ago against a hosted chat endpoint no longer reproduces. What captured from the original run would let you distinguish a silent provider-side model change from a shipped fix, and what do you do if you captured none of it?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Compare the model identifier the response echoed then and now, the application's system prompt then and now, and your request ids and timestamps. A changed identifier or prompt points at drift; both unchanged points at a fix or variance. With none captured, you cannot attribute it — re-run, capture properly this time, and say so.

open as a page

Two findings against the same chat feature: A succeeds in 18 of 20 attempts and returns the assistant's own refusal-policy wording; B succeeds in 1 of 50 attempts and returns another customer's record. Which do you rate higher, and what do you tell the team about B's 2%?

level: seniorimportance: should knowfreq 55%

basics

~20 s

B, clearly. Impact decides the order: one success in B is a real disclosure of another customer's data, while A leaks the product's own policy text. Tell the team that 2% is not rare for an attacker — around fifty cheap retries make success roughly even odds, and nothing in the rate makes the leak smaller.

open as a page

A red-team scan is re-run against the same application every week. How would you give each report item an identity that survives across runs, so a fixed item does not reappear as brand new and a genuinely new failure is not silently absorbed into an old one?

level: principalimportance: should knowfreq 30%

basics

~20 s

Derive the identity from things a fix would change — the failed control, the behaviour, the surface — never from the prompt text, the attempt index or the run timestamp. Keep the mapping in a register you own, review unmatched items by hand each week, and record splits and merges with lineage instead of new ids.

open as a page

The core evidence for an AI red-team finding is a model response containing genuinely harmful, actionable content, and the fix requires developers to reproduce it. Your issue tracker is readable by most of the company. How do you package the finding?

level: principalimportance: should knowfreq 29%

basics

~20 s

Split the package. The broadly readable ticket carries a characterisation of the harm, the criterion, the rate and the fix requirement. The reproduction input and the full response go to an access-controlled store, referenced by identifier and hash, granted to the engineers doing the fix. Agree the split, the retention and the deletion date in advance.

open as a page

You lead a team running several different AI red-team tools, each writing its own log format. Define the capture standard that keeps any finding re-runnable a quarter later against hosted targets: what is mandatory, and what do you deliberately not store?

level: principalimportance: should knowfreq 30%

basics

~20 s

Mandate a small common envelope every tool must emit: target identity and echoed model id, effective decoding settings, system-prompt digest, verbatim turns, attempts and successes, timestamps and request ids, harness build. Deliberately skip credentials, real customer data, and full bodies for non-hit attempts beyond a sampled retention window.

open as a page

You own the severity rubric for AI findings across an organisation where every control the findings target fails some fraction of the time. How do you let a measured success rate into the rating, and what do you refuse to let it do?

level: principalimportance: should knowfreq 35%

basics

~20 s

Let the rate modify exploitability only, alongside retry cost and who can reach the entry point. Refuse to let it touch impact, refuse to let it gate whether a finding gets a severity at all, and set a floor so any reproducible success keeps a minimum rating. Require counts, a measurement point and a date on every filed rate.

open as a page