skip to content

One attack family in your red-team run flagged forty-plus attempts against the same chat feature. You will not attach all forty to the developer ticket. How do you choose which attempts go into the evidence package and which stay out?

level: seniorimportance: should knowfreq 44%

answer

  1. fix boundary, not best example
  2. each variant defeats one cheap fix
  3. benign controls against over-blocking
  4. rationale line per attachment
  5. full log as appendix

basics

~20 s

Pick the set that defines the fix boundary, not the prettiest example. One canonical reproduction, plus a few variants chosen because each defeats a different plausible cheap fix, plus benign inputs that must keep working after the change. Link the full log as an appendix and say why each attached case is there.

solid answer

~50 s

Two opposite mistakes bracket this. Attaching all forty moves your triage cost onto the application team, and the ticket gets skimmed. Attaching one moves the fix's scope to that one string — the developer blocks the literal phrase, the ticket closes green, and the behaviour is still reachable. So select for **fix boundary**: - **One canonical case**, cleanest and fully reproducible, with the observed rate. - **Two or three variants that each break a different naive remedy** — a paraphrase (defeats literal string matching), a different delivery channel such as retrieved content rather than user input (defeats input-only filtering), a different surface phrasing of the same objective (defeats a narrow classifier rule). - **Benign controls**: legitimate inputs that must still succeed after the fix, or the team ships an over-broad block and hears about it from users. - **A one-line rationale per attached case** naming the fix it is there to defeat. The remaining thirty-odd stay in a linked, grouped appendix. Volume is a severity input, not evidence.

go deeper

for a junior

Knows to attach a clear reproduction rather than the whole log, and that one example is thin.

for a middle

Selects a few varied cases and includes benign inputs that must keep working.

for a senior

Chooses attachments explicitly to defeat each cheap remedy, annotates why each is there, and notices when a variant is really a separate finding.

for a principal

Makes the two-sided acceptance criterion — attacks blocked, controls preserved — the standing definition of done for this class of finding across teams.

The selection question is really a question about what the developer will do next. They will write the narrowest change that makes your evidence stop reproducing. Your attachment set therefore *specifies* the fix, whether you meant it to or not. ## Forty is smaller than it looks Before selecting anything, work out what produced forty. Scanners multiply: garak sends each prompt in a probe `--generations` times, so a probe holding eight prompts at five generations yields forty attempts, and the flagged subset counts attempts, not distinct inputs. promptfoo multiplies test cases by the strategies applied to them, so one seed case can appear a dozen times in different clothing. PyRIT's multi-turn orchestrators emit a conversation piece per turn, so a handful of conversations can look like a wall of rows. Collapse first to distinct inputs, then to distinct objectives. Forty flagged attempts commonly reduce to five or six genuinely different ways in — and that reduction is the judgment the developer cannot make for themselves, because it needs the run's configuration to interpret. ## Select against remedies, not against attempts Enumerate the cheap fixes a hurried team reaches for: a literal denylist entry, an input-side classifier, a system-prompt instruction, a regex on the output, a rate limit. Then choose attachments so that each cheap fix leaves at least one attached case still failing. In practice that is a canonical case plus two or three boundary cases: a paraphrase (defeats literal string matching), the same objective arriving through a different channel such as retrieved document content rather than the user turn (defeats input-only filtering), a different surface phrasing of the same objective (defeats a narrow classifier rule). Three to five attachments, each with a one-line rationale naming the remedy it exists to defeat, does more work than thirty near-duplicates and tells the developer the required shape of the fix without you dictating the fix itself. ## The controls are not optional Attach benign inputs that sit closest to the attack — same topic, legitimate intent — and that must still succeed afterwards. This makes the acceptance criterion two-sided: attacks stop, controls still pass. Without it an over-broad block passes verification and real users discover the regression instead. That pair is also exactly what becomes the regression test the team keeps. ## What it costs Every attachment is a recurring bill on the other team. Reading a case, reproducing it and deciding what change covers it runs five to fifteen minutes, and most attachments then become permanent test cases that execute on every pipeline run. Three to five costs them under an hour once and a small test file forever. Forty costs half a day of triage you transferred to them, and the predictable outcome is a skim, a fix aimed at the first example, and a green closure. On your side, honest selection costs about an hour: re-running each candidate to confirm it still reproduces today, and hand-marking it against the written criterion so you are not attaching a case you have never read. ## Where the number misleads "Forty flagged attempts" gets read three wrong ways. As forty vulnerabilities: it is one family, and the report's unit is not the tool's unit. As a severity signal: raising the repeat count raises the flagged total linearly while the target does not get any worse, so volume is largely an artefact of your configuration and at most an input to rating. And as a success rate: the denominator is the number of attempts your config chose to make, by an operator deliberately steering, which is not a population of user sessions. There is a sharper trap underneath. A family with forty near-identical hits and a family with one hit can carry identical impact, and the single-hit family is often the more interesting one, because it is precisely the case the obvious fix will miss. If, while selecting, you find a case that no cheap fix for the others would touch, treat that as evidence it was never the same family and file it as its own finding rather than burying it as attachment number four. ## What to check before you send Apply the most obvious one-line fix on paper and ask whether anything in your attachment set still fails. If nothing does, the package is under-specified and the ticket will close prematurely. Then confirm each attachment still reproduces today, that each carries its rationale line, and that at least one benign control is present so "blocked everything" cannot pass as a fix.

  • How many attached cases is right?
    As many as there are distinct cheap fixes to defeat, plus controls — typically three to five. If you need ten, you are probably looking at more than one finding.
  • The developer asks for the other thirty-five anyway. What do you send?
    The grouped appendix with counts and one line on what each group shares, not a raw dump — the grouping is the part that took judgment.

Counting flagged attempts as findings is like counting a photocopier's output as new documents: the repeat setting multiplies copies, not distinct failures.

saying these in an interview costs you the question

  • Attaching the whole run output and calling selection the developer's job.
  • One example only, with no variant that survives a literal string block.
  • No benign control cases, so the team cannot tell over-blocking from a fix.
  • Prescribing the exact remedy instead of specifying what the fix must cover.
  • Treating the number of flagged attempts as if it were the evidence.

context