skip to content

You collapsed a red-team scan's flagged attempts into eight report items, the developer fixed all eight, and a re-run of the same scan still fails. How do you tell whether your grouping key was too coarse, and what do you inspect inside a bucket to prove it?

level: seniorimportance: should knowfreq 44%

answer

  1. map re-run failures back to the item
  2. read members, not labels
  3. behaviour / failed control / surface
  4. two members, two different fixes = too coarse
  5. lowered rate is not the same as bad grouping

basics

~20 s

Take the still-failing attempts from the re-run and find which of the eight items they were grouped under. Then open that bucket's members and compare what the model produced and which control let it through. If members differ on either, the key merged distinct defects and the fix only closed the exemplar you attached.

solid answer

~60 s

Start from the re-run, not the report. Map each still-failing attempt back to the item it would have been keyed into. Three outcomes, and they need different responses. 1. **It maps to a fixed item.** Now open that item's members. Compare, member by member, the behaviour produced, the control that failed, and the surface reached. Variation on any of those inside one bucket is the proof: the key merged defects that need separate fixes, and the developer shipped against the single exemplar attached. 2. **It maps to no item.** The re-run reached something the first run did not, or the fix moved the failure. That is a new item, not a dedup defect. 3. **It maps to a fixed item and the members are homogeneous.** The key was fine; the fix is incomplete or the control is probabilistic and the first run's exemplar was simply the loudest case. The general rule: a bucket is too coarse when two of its members would be closed by two different changes. That is the only test that matters, and it is checked by reading outputs, not by counting rows.

go deeper

for a junior

Likely goes straight to re-running the scan; may not think to map failures back to the item they were grouped under.

for a middle

Should propose opening the bucket and comparing members, and name at least behaviour and failed control as the fields to compare.

for a senior

Should run the three-way mapping, use the 'two members, two fixes' test, and separate a dedup defect from a probabilistic control whose rate merely fell.

for a principal

Turns it into a standing pre-publication invariant, keeps split lineage against the original item so re-splits do not read as new findings, and owns telling the developer the first report under-counted.

### What the scenario is actually telling you Eight items, eight fixes, still failing. The instinct is to suspect the fixes. The more likely explanation is that the report was short because it was over-collapsed, and shortness felt like quality at the time. Deduplication makes a claim — *everything inside this bucket would be closed by one change* — and a re-run is the first experiment that can falsify it. Treat the re-run as that experiment rather than as a verdict on the developer. ### The three-way mapping Start from the re-run, not the report. For each still-failing attempt, recompute the grouping key you used and look up which of the eight items it would have landed in. There are exactly three outcomes and they need different responses: 1. **It maps to a fixed item.** This is the case worth investigating. Open that item's stored members and compare them. 2. **It maps to no item.** The fix moved the failure, or this run reached something the first did not. That is a new item, not a deduplication defect, and saying so protects your own credibility. 3. **It maps to a fixed item and the members are homogeneous.** Your key was fine. Either the fix is incomplete, or the control is probabilistic and the fix lowered a rate without eliminating it. This mapping only works if you kept the members. If deduplication threw the collapsed attempts away, there is nothing to inspect and the only remaining move is a fresh run with a finer key — which costs another full scan and another triage pass. ### What to inspect inside a suspect bucket, in order - **Observed behaviour.** Read a sample of the members' actual model outputs, not the detector labels. Two responses flagged by the same check can be materially different — one hedged paraphrase and one direct disclosure carry the same label and are not the same defect. - **Which control failed.** An input classifier that never fired and a system-prompt boundary that fired and was argued past are different defects with different owners, even when the technique and the visible output look alike. This is the field most often missing from a bucket and the one that most often proves it should split. - **Surface reached.** The same behaviour through a direct chat turn and through retrieved document content is two items. A fix in the chat path does not touch the retrieval path, and vice versa. - **Variant span.** If the successful variants cluster tightly around one framing, one change plausibly covers them. If they are spread across unrelated framings, the bucket is a category, not a defect. The general test underneath all four: **two members that would be closed by two different changes prove the key was too coarse.** That is the only decisive test, and it is answered by reading outputs, not by counting rows. ### What it costs Cheap, if the members were kept: recomputing keys over a re-run is scripted and runs in seconds, and reading three or four members from each suspect bucket is under an hour of analyst time. Expensive if they were not: a fresh scan plus triage is hours of wall clock, another metered-endpoint bill, and a second demand on the developer's attention. The real cost, though, has already been paid — a wasted remediation cycle is days of engineering, a re-test, and the trust you spend when a stakeholder learns the first count was wrong. That asymmetry is the argument for spending the hour on bucket inspection before publication rather than after. ### Where the number misleads Three readings of this situation are wrong in ways that look reasonable: - **"Eight items, eight fixes, still failing, so the fixes were bad."** The count of fixes matches the count of items by construction; it says nothing about the count of defects. Matching numbers give a false sense that the ledger balanced. - **"The bucket with the most attempts is the one that was too coarse."** Bucket size reflects how hard the tool pushed that area — how many variants the template expanded into, how many trials were configured — not how many distinct defects it holds. A four-member bucket can hide two defects while a two-hundred-member bucket holds one. - **"Any post-fix hit proves the grouping was wrong."** Probabilistic controls fail at a rate. A fix that took a case from three-in-ten to one-in-two-hundred still produces hits in a large enough re-run, and splitting a correctly-grouped bucket in response manufactures findings that no change will ever close. Separate the two before accusing your own key: a homogeneous bucket plus a fallen rate is a rate story, not a grouping story, and you can only see the fall if both runs published their attempt counts. ### What you change afterwards Split the bucket, attach one exemplar per split, and record the split against the original item id so a reader sees the total did not really jump — a silent re-split reads as new findings appearing from nowhere. Then make it structural: a pre-publication check that no item ships with a single exemplar when its own members disagree on behaviour, failed control or surface. That check would have caught this before the developer ever saw the report.

  • How do you tell an over-coarse bucket from a fix that merely lowered a probabilistic failure rate?
    Homogeneity of the bucket. If members still agree on behaviour, failed control and surface, the grouping was right and you are looking at a rate that dropped but did not reach zero.
  • What do you change so the same mistake is caught before publication next time?
    A check that no item ships with one exemplar when its own members disagree on behaviour, failed control or surface — plus recording the split against the original item id.

saying these in an interview costs you the question

  • Blaming the developer's fix without opening the bucket.
  • Deciding the key was too coarse from row counts rather than from reading outputs.
  • Re-issuing split items with fresh ids and no lineage, so the report looks like new failures appeared.
  • Never having kept the collapsed members, leaving nothing to inspect.
  • Assuming any post-fix hit proves bad grouping, ignoring that the control is probabilistic.

context