A red-team scan is re-run against the same application every week. How would you give each report item an identity that survives across runs, so a fixed item does not reappear as brand new and a genuinely new failure is not silently absorbed into an old one?
answer
- identity from what a fix would change
- no prompt text, no run timestamp
- register plus manual match queue
- splits and merges carry lineage
- absent this week is not fixed
basics
~20 sDerive the identity from things a fix would change — the failed control, the behaviour, the surface — never from the prompt text, the attempt index or the run timestamp. Keep the mapping in a register you own, review unmatched items by hand each week, and record splits and merges with lineage instead of new ids.
solid answer
~1 minIdentity across runs is a different problem from deduplicating within one run, and the failure modes are opposite. Within a run, a bad key merges or explodes the count. Across runs, a bad identity destroys the trend line — the thing leadership actually reads. The design has three parts. **A key made of durable attributes.** Failed control, observed behaviour, entry surface. Anything the scanner varies (prompt text, ordering, seed) and anything about the run (timestamp, attempt index) must be excluded, or every week produces a fresh population of items. **A register, not a recomputation.** Store the assigned identity with its key and its members. Each week, match new items against the register, and route anything that matches nothing to human review rather than auto-creating it. **Explicit lineage for splits and merges.** When triage discovers a bucket was two defects, the new items point at the parent. When two items turn out to be one, the survivor records both. Without this, a legitimate re-split reads as regression and a merge reads as a fix. The cost is honest to state: this is a maintained dataset with an owner, and its accuracy decays if unmatched items are auto-created instead of reviewed.
go deeper
Likely proposes hashing the prompt or reusing the tool's own case id, and does not yet see that both churn between runs.
Should build the key from durable attributes and see why run-specific fields must be excluded.
Adds the register with a manual queue for unmatched items and separates 'absent this week' from 'fixed'.
Owns the whole scheme: lineage for splits and merges, metrics that detect key drift, a named owner, and the willingness to drop the register when the trend is not worth its maintenance cost.
### Why this is a different problem from deduplicating one run Within a single run, a bad grouping key merges defects or explodes the count, and the damage is contained to that report. Across runs, a bad identity destroys the **trend line** — and the trend is the only thing most stakeholders actually read from a recurring programme. "Are we getting better?" is answerable only if "the same finding" means the same thing in week one and in week nine. Everything about a recurring scan works against that. Generated prompts are re-generated. The tool's own catalogue of techniques shifts between versions, renaming and re-numbering cases. The target application changes underneath, so surfaces appear and disappear. Any identity built on those things is re-minted every week. ### The design, in three parts **A key made of durable attributes.** The identity is derived from what a remediation would change: the control that failed, the behaviour the model produced, the entry surface it was reached through. Everything the scanner varies (prompt text, ordering, random seed) and everything about the run itself (timestamp, attempt index, the tool's generated case id) is excluded by construction. The diagnostic tell that a team skipped this step is a weekly report whose item count is stable while almost no identifiers repeat — the underlying defects are the same and the scheme cannot see it. **A register, not a recomputation.** Store each assigned identity with its key and its members, and match each week's items against that store. Do not recompute identity from scratch per run, because a key that shifts even slightly then reassigns everything at once. **Explicit lineage for splits and merges.** Triage learns things: a bucket turns out to be two defects, or two items turn out to be one. Both are normal and both must be expressed as parent and child rather than as deletion and creation. Without lineage, a legitimate re-split reads as a regression and a merge reads as a fix, and the trend line reports application changes that never happened. ### Three rules that carry most of the value **Fail toward review, never toward creation or merge.** An item matching nothing in the register lands in a human queue. Auto-creating grows the register forever and makes the count meaningless. Auto-merging on a fuzzy match is the more dangerous reflex, because it hides genuinely new defects inside old identifiers and is invisible in the numbers — nothing looks wrong, the count is stable, and a new failure is silently attributed to something already being worked. **Status is separate from presence.** An item absent from this week's run is not fixed. It may not have been probed at all — catalogue changed, surface unreachable, budget cut short — or the control is probabilistic and simply did not fail this time. Closure requires positive evidence: an accepted fix, or a targeted re-test at a stated attempt count. "Not observed this run" is its own state and belongs in the schema. **A named owner.** The register is a maintained dataset. Without someone accountable for the weekly match queue it decays within a couple of months into a list nobody trusts. ### What it costs The scan is automatable; the register is not. Realistic ongoing cost is a standing hour or two of analyst time a week to work the unmatched queue and record lineage, plus the engineering to keep the key computable as the tool changes. That is the item to be honest about with whoever asked for weekly scanning: if the trend is not worth roughly a day a month of a skilled person, the correct answer is to publish per-run reports and stop claiming a trend exists. A register maintained badly is worse than none, because it produces a graph that looks authoritative and is not. ### Where the number misleads - **Item-count stability is not stability.** A flat count with total identifier turnover means the scheme is re-minting, not that the application is steady. Read turnover, not totals. - **A drop in open items after a tooling upgrade is almost never remediation.** When the tool renames or regenerates its catalogue, keys reshape and old items stop matching. That looks like closure. Any week where identifier turnover spikes without a corresponding application change should be investigated as a tooling change first. - **"Closed" built from absence inflates the fix rate.** Track how many items were closed by re-test versus by silence; if the second number is large, the programme's headline improvement is mostly bookkeeping. - **A rising manual-match share is a leading indicator, not noise.** It means the key is drifting away from what the tool now emits, and the trend is degrading before anyone notices in the graph. ### What you check Weekly: the manual-match share, the split of closures into re-tested versus not-observed, and identifier turnover against the application's own change log. Before publishing any trend to leadership: confirm the tool version and its catalogue did not change between the two points being compared.
- Should an unmatched item auto-create or wait for review?Wait for review. Auto-creation inflates the register forever, and the opposite reflex — auto-merging on a fuzzy match — hides genuinely new defects inside old ids, which is worse because it is invisible.
- The item did not appear in this week's run. Is it fixed?No. Absence can mean it was not probed, or the control is probabilistic and did not fail this time. Closure needs an accepted fix or a targeted re-test at a stated attempt count.
- Item ids turned over heavily this week with no application change. First hypothesis?The tooling changed — a regenerated catalogue or renamed categories reshaping the key — not the application. Check the key's inputs before reporting a regression.
Keying a finding on the prompt that triggered it is like filing a customer record under the phone number they called from: it works until they call from somewhere else, and then every week's callers look like brand-new customers.
saying these in an interview costs you the question
- Hashing the prompt text or the tool's generated case id as the identity.
- Closing an item because it did not appear in the latest run.
- Auto-merging unmatched items into the nearest existing one.
- Re-issuing split items as new findings with no parent reference.
- Presenting a week-over-week trend without checking whether the tooling's catalogue changed.