You have run a prompt-template mutation fuzzer against a vendor chat endpoint for months. Seeds that stop producing hits get dropped and the corpus is re-seeded from survivors. What goes wrong with that corpus over time, and how do you stop the tool testing one family forever?
answer
- survivorship: corpus keeps only what works
- dead seed = regression control
- frozen set vs exploration set
- per-family quotas, tagged seeds
- pin corpus before trending a number
basics
~20 sThe corpus survivorship-collapses. Dropping seeds that stop landing leaves only the family the current model version is weakest to, so the hit rate measures your corpus, not the model. Keep a frozen regression set that includes the dead seeds, quota seeds by family, and reserve part of every campaign for families the corpus has never held.
solid answer
~50 sTwo failures compound. **Survivorship**: the corpus keeps only what still works, so it narrows toward one family and the run's output becomes a description of your own selection process. **Lost regression signal**: the seeds you dropped were the ones a fix had just closed, and a dropped seed can never tell you the fix regressed in the next model update. The fix is to split the corpus by purpose. A **frozen regression set** never loses members: seeds that produce zero hits stay in and their zero is recorded as the result. A **exploration set** carries seeds tagged by family, with a per-family quota so no family can crowd out the rest, refreshed from human authoring and published taxonomies rather than from survivors. And compare only like with like: a hit-rate trend across model versions is meaningless unless corpus, operators and budget are pinned. If the corpus changed, the trend is your corpus adapting, not the model.
go deeper
Should notice that if you only keep what works, the corpus stops representing anything but the current soft spot.
Names survivorship explicitly and proposes keeping a fixed set of seeds, including ones producing zero hits, so results stay comparable.
Separates regression from exploration corpora, tags and quotas families, keeps zero results as evidence, and pins corpus plus operators before trending anything across model versions.
Owns the reporting consequence — untested families reported beside hit counts, corpus versioned with results — and sets the cadence that balances metered query cost against regression coverage.
**What survivorship collapse is, mechanically.** The policy has two steps: drop seeds that stopped producing hits, and re-seed the corpus from survivors. Each step is individually defensible and together they form a selection loop with no counterweight. Every cycle, the corpus loses the templates the current model version handles well and gains descendants of the templates it handles badly. After a few months the corpus is not a sample of jailbreak families at all; it is the fixed point of the vendor's own patching — a concentrated description of whatever the current model version happens to be softest on, restated in a few hundred rewordings. **Why nobody notices.** The failure produces no error and no empty dashboard. Hit counts stay comfortably non-zero, because a narrowing corpus is a corpus increasingly made of things that work. Cost per hit even improves, which reads as the programme getting more efficient. What changed silently is the question the suite answers: it started as *is this assistant robust across jailbreak families* and became *has my corpus found the current soft spot*. The second question is worth something, but it is not what the number on the slide is being read as. **The regression cost is the sharper one.** Think about when a seed stops landing. Almost always it is because the vendor closed something, or a fix on your own product shipped. That is the exact moment the seed becomes valuable, because it stops being an attack and starts being a control: it is now the thing that will tell you if a later model update quietly reopens the behaviour. Deleting it converts a durable regression signal into permanent silence. Zero hits from a retired seed is not nothing — it is the result a regression suite exists to produce, and it must be recorded and plotted as a zero rather than as an absence. **Structure that survives this.** - **Two corpora, separately versioned.** A *regression set* that is append-only and never pruned: every member runs every cycle, and its zeros are recorded as results. An *exploration set* that churns freely and is allowed to be wasteful. - **Tag every seed with the family it represents and enforce per-family quotas**, so one prolific family cannot eat the exploration budget simply because it is currently productive. - **Reserve a fixed slice of every campaign for seeds no mutation produced** — hand-authored against a published taxonomy, or lifted from incident reports. This is the only line item in the whole programme that can produce evidence about a family the corpus has never held; mutation structurally cannot. - **Track two counters over time**: families seeded, and families seeded that returned nothing. The second one rising is the signal that fixes are holding, and it is invisible under a prune-the-dead policy because the dead were deleted. - **Pin operators, variant budget and decoding settings** for anything you intend to trend, and version the corpus alongside the results, so a later reader can tell whether a change in the number came from the model or from you. **What it costs, and the honest answer to the cost objection.** Re-running dead seeds spends metered queries for a predictable zero, and that is the argument people use for pruning. Answer it with cadence rather than deletion: run the full regression set at release boundaries and after each shipped fix, and a stratified sample of it in between. A 300-seed regression set at three re-sends is a few thousand requests, which against a rate-limited endpoint is an hour or two of wall clock — cheap next to the alternative, which is discovering a reopened behaviour from a customer report. The seed keeps its place in the corpus; only its frequency drops. **How the resulting number misleads if you skip all this.** A hit-rate trend plotted across model versions from a corpus that was pruned between runs is uninterpretable in a specific way: one of the two inputs to the ratio moved at every point, so a dip is equally consistent with the model improving, with pruning having removed the productive seeds, with an operator change, or with ordinary sampling variance. Presented as a robustness trend it is worse than no chart, because it licenses a decision that the data cannot support. **What I would check on inheriting such a suite.** When each seed entered the corpus and when it last produced a hit; how many genuinely distinct families the current corpus represents versus how many the taxonomy lists; whether zeros are stored at all; whether any result older than the last corpus edit is still being plotted on the same axis as today's; and whether a documented rule exists for what may be removed and who approves it.
- A dead seed suddenly lands again after a model update. What does that tell you and what do you do first?It is a regression: a previously closed behaviour reopened. First confirm it reproduces across several re-sends and that the variant still carries the original request, then report it as a regression with the date it last passed — that history is what makes it actionable rather than just another hit.
- How do you get any evidence about a family your corpus has never contained?Not from this tool. Mutation only explores around what you seeded, so an unseeded family needs a seed authored some other way — a human writing to a taxonomy, an incident report, an externally published family — and until that seed exists the honest report line is 'untested', not 'no findings'.
Removing seeds that no longer land is like taking down the smoke detectors that have never gone off: the ones that stayed quiet are exactly the ones proving nothing is burning, and once they are gone a new fire is silent too.
saying these in an interview costs you the question
- Prunes seeds that stop landing and sees no downside — the dead seed was the regression control.
- Plots a hit-rate trend across model versions from a corpus that changed between runs.
- Argues query cost justifies deletion rather than lowering the cadence of the regression set.
- Refreshes the corpus only from its own survivors, so no genuinely new family can ever enter.
- Treats 'zero hits this cycle' as nothing to record instead of the result a regression suite exists to produce.