skip to content

Continuous automated LLM red-team suites make successful harmful outputs pile up as an archive that doubles as the regression set proving fixes hold. How do you decide what is kept, for how long, and who may open it?

level: principalimportance: should knowfreq 28%

answer

  1. regression replays input, re-scores output
  2. keep input + verdict + hash, not the text
  3. retention by harm class, default expiry
  4. two-person logged access, keys split from CI
  5. threat-model the archive; rotate reviewers

basics

~20 s

Split the artefact. Attack inputs are what regression testing needs, so keep those under access control; harmful completions are usually re-derivable by re-running and re-scoring, so keep only a verdict and a hash. Set per-class retention with default expiry, two-person access with logging, and a named owner outside the red team.

solid answer

~60 s

Start from what regression actually consumes. A regression run replays the **input** and re-scores fresh output; it almost never needs the stored completion. That single observation shrinks the archive to attack inputs plus verdicts and hashes, and reserves stored completions for cases where the harm claim genuinely cannot be re-derived. Then price the asset honestly. A curated set of validated working attacks against your own systems is one of the highest-value targets on your network and your largest insider-risk concentration — and it is also a compliance surface and a discovery risk. Value on the other side: regression evidence, calibration of your automated grader against human judgement, and trend measurement over releases. Design: per-class retention (some classes never stored), default expiry with explicit renewal rather than indefinite keep, access by name with two-person approval and logging, keys separate from the CI identity that writes the archive, and the archive listed in the data inventory and in incident-response planning. Ownership sits with a named accountable owner outside the red team; the red team classifies, it does not self-authorise.

go deeper

for a junior

Says the archive must be access-controlled and not kept forever; may not yet separate the attack input from the harmful output.

for a middle

Argues that regression needs the input and a fresh score rather than the stored completion, and proposes expiry plus restricted access.

for a senior

Adds per-class retention, split write/read privileges from the CI identity, logged break-glass access, and a documented destruction path that includes backups.

for a principal

Assigns accountable ownership outside the red team, threat-models the archive as a top-value target in incident response, puts it in the data inventory, and treats reviewer rotation and welfare as part of the policy.

The tension is genuine and does not dissolve: destroy everything and you cannot prove a fix held; keep everything and you have built a growing, indexed, validated attack corpus against your own products, sitting on your own network with internal access. ## Decompose the artefact first A stored run has four separable parts: the **attack input**, the **model output**, the **verdict**, and the **run metadata**. A regression run consumes the input and the metadata, sends the input to the current build, and produces a *fresh* output and a *fresh* verdict. Replaying a stored output would test nothing at all — it would assert that a string you already have is still the string you have. That single observation does most of the work. The default becomes: keep input, metadata and verdict; keep a hash of the completion; and keep the completion itself only where the harm claim rests on properties of that specific text and cannot be re-derived. Those exceptions are real but narrow — grader calibration over time, and defending a contested severity months later — and they belong in the tightest access tier. ## Retention by harm class, not by uniform age Some classes never enter the archive. Others carry a short clock. Regression-critical inputs may justify a longer one, but should still expire by default with renewal as an explicit, signed act by a named person. Indefinite retention is never a decision anybody made; it is the state every archive drifts into when expiry is opt-in. ## Access, ownership, and the archive as a target Access should be by name, with two-person approval, every read logged and the log reviewed — not a group membership granted at onboarding, which turns the corpus into ambient reading for a population that only grows. Read keys should not be held by the CI identity that writes the archive; write and read are different privileges here, and a compromised pipeline should not yield the whole corpus. Ownership belongs to a named accountable person outside the red team, typically in legal or trust-and-safety: the team that benefits from keeping evidence should not be the team that authorises keeping it, and the red team's job is to classify findings into the schedule rather than to set it. Then threat-model the archive itself and put it in incident-response planning, answering plainly: if this store is breached, what does the adversary now have? The answer is a curated, validated playbook of attacks that work against your live products — a higher-value target than most of what the security team is defending. Finally, the humans who review this material need rotation, consent and support; concentrating it on one reviewer is both a welfare failure and a single point of judgement on every severity call. ## What it costs — and this is where the policy actually breaks Storage is trivially cheap, which is why nobody argues about storage. The real cost is regeneration. If the regression suite holds two thousand attack inputs and each is sampled three times to cope with sampling variance, one full run is six thousand model calls; nightly, that is on the order of a hundred and eighty thousand calls a month, plus the grader calls on top if scoring is itself a model. Against a metered frontier endpoint that is a line item somebody will notice, and the tempting saving — cache the completions and re-score the cache — is exactly the move that destroys the point of the run. The honest response is to shrink or tier the suite (a small nightly set, the full set weekly), not to stop regenerating. There is also human cost: access review, the owner's time, and expiry automation that somebody has to maintain. ## Where the number misleads A green regression run is read as "the model is fixed". It supports something much narrower: *these archived inputs no longer succeed against this build*. It says nothing about the attack space outside the archive, and — the sharp edge — a fix narrowly targeted at the archived strings, whether a guard rule from your own team or a provider-side patch, produces exactly the same green result as a genuine fix. An archive frozen at the moment findings were logged is a set the defence can overfit to, and the longer it is used as the pass/fail gate the more the system is tuned to it. Treat archive pass rate as a regression alarm only, never as a coverage measure, and keep a rotating set of fresh, unarchived variants of the same attack family as the control that tells you which kind of fix you got. ## How I would know the policy is more than paper Sample the store against the schedule and confirm items past expiry are actually gone. Read the access log for named, justified reads rather than routine browsing. And verify that a regression run genuinely regenerates output from inputs — because if it has quietly come to depend on stored completions, the entire argument for not retaining them has already collapsed and nobody noticed.

  • What is the strongest argument for keeping some harmful completions rather than only inputs and verdicts?
    Grader calibration and contested severity: to check that your automated judge agrees with human reviewers over time, and to defend a severity rating months later, you sometimes need the specific text — which is why that tier exists, narrowly and with the tightest access.
  • How would you detect that the retention policy exists on paper only?
    Sample the store against the schedule for items past expiry, review access logs for routine rather than justified reads, and confirm a regression run reproduces findings from inputs alone.

A regression archive should keep the keys that opened the door, not the video of the door opening — replaying the video proves nothing about today's lock. And if the new lock only resists the keys already in your kit, you have a lock cut to match the kit rather than a better lock.

saying these in an interview costs you the question

  • Keeping every successful transcript indefinitely because 'it might be useful for regression'
  • Granting archive access by team role rather than by name, with no access logging
  • Letting the CI identity that writes the archive also hold read keys
  • Never asking what an attacker gains by breaching the archive
  • Concentrating all review of harmful material on one person with no rotation

context