You reported that an automated jailbreak search had saturated on a target. A colleague then ran the same search class against the same target with a different seed corpus and objective and found successful templates yours never produced. What did your saturation measurement actually measure, and how should it have been worded?
answer
- flat curve describes the search, not the target
- seeds, objective, operators, scoring step
- coverage fraction with no denominator
- word it as this configuration exhausted
- run disjoint configurations, measure overlap
basics
~20 sIt measured that one search stopped finding new templates, not that the target has none left. Saturation is relative to the seeds, objective, scoring step and mutation operators that run used; change any of them and the reachable region changes. Word it as this search exhausted, with those settings named, never as an absence of weaknesses.
solid answer
~50 sA discovery curve flattening is a statement about the search, not the target. Every automated jailbreak search explores a region fixed by four things: the seed corpus it starts from, the objective it optimises, the operators it may apply, and the scoring step that decides what counted. A template outside that region is not merely undiscovered — it is unreachable, and no extra allowance would have produced it. So the honest claim is bounded: *this configuration stopped producing new distinct templates after N queries*, with those four settings and the clustering rule named alongside. That wording survives your colleague's result; "the target saturated" does not. Saturation is therefore a signal to redirect, not to conclude. Running several deliberately different configurations and comparing which templates each reaches says far more about reachability than one long run that plateaus cleanly.
go deeper
Recognises that finding nothing new does not mean nothing is there, and that another starting point can find more.
Names the settings that bound the reachable region and rewords the claim to be about the run rather than the target.
Explains why a coverage fraction has no denominator here, and proposes several disjoint configurations with overlap measured between them instead of one long run.
Owns how the finding is consumed downstream: ensures a plateau is never converted into an assurance claim, and funds configuration diversity as the programme's actual coverage strategy.
## Why a flat curve feels like completion A discovery curve that flattens has the *shape* of a task finishing, and that is the whole trap. Look at what the number would need in order to mean completion: a fraction, distinct templates found over templates there were to find. The bottom of that fraction does not exist. There is no enumerable universe of jailbreak templates to divide by, so a coverage claim built on a saturation curve is a fraction with no denominator, and every reader who converts it into an assurance statement is inventing one. What the curve does measure is real and useful — it says the marginal query in *this* configuration stopped buying new mechanisms. That is a resource-allocation fact about your run. It is not a property of the target. ## What bounds the reachable region An automated jailbreak search can only ever reach a region fixed by four settings. A template outside that region is not merely undiscovered; it is unreachable, and no extra allowance would ever have produced it. - **Seeds.** The corpus of starting prompts. A mutation or attacker-model search explores the neighbourhood of its seeds; whole families sharing no structure with any seed lie outside it entirely. - **Objective.** The behaviour the search optimises towards, as encoded in the scoring step's target. Candidates that elicit some *different* harmful behaviour score zero and are discarded even though they worked. - **Operators.** What the search is permitted to change — tokens, phrasing, encoding, conversational structure. An operator set that never restructures a dialogue cannot find a multi-turn shape, however long it runs. - **Scoring step.** Anything the scorer will not accept does not exist to the search. A narrow scorer prunes the frontier invisibly, and because pruning happens before you ever see a candidate, there is no trace of it in the output. ## The wording that survives contact Replace "the target saturated" with a sentence that carries the configuration: which seed corpus, which objective, which operator set, which scoring method, how many queries were spent, and the clustering rule used to count distinct templates — plus an explicit line that this bounds the search and not the target. That version survives a colleague finding more; the short version does not, and the difference is not pedantry. Downstream readers making a ship decision or a funding decision will otherwise convert your curve into an assurance claim it cannot support, and you will not be in the room when they do. ## Making the bound less severe, and what that costs The fix is deliberate configuration diversity: several runs with disjoint seed corpora and different objectives, each necessarily smaller, instead of one long run. Then measure something more informative than any single curve — the *overlap* between what each configuration reached. Heavy overlap is evidence the reachable region really is small. Near-disjoint results are direct evidence that a plateau in any one of them meant very little, and that finding is worth reporting on its own, above any individual run's numbers. The arithmetic is worth stating plainly. Splitting a thirty-thousand-query allowance into three ten-thousand-query runs buys you the overlap measurement and costs you depth in each: each run plateaus earlier and may miss templates a longer run in the same region would have found. The binding constraint is usually not queries at all but engineer hours — curating a seed corpus, restating an objective, and re-calibrating a scoring step is fixed setup per configuration, and that fixed cost, not the query price, is what actually limits how many configurations a team runs. Budget it explicitly or configuration diversity quietly never happens. ## What I would check when handed a saturation claim - The seed corpus and the objective, named specifically enough that someone could run something deliberately different. - The clustering rule and its version, since a plateau produced by an under-splitting rule is not a plateau at all. - Queries spent and the billed-call multiplier, so the stop can be read as exhausted-search or exhausted-budget. - Whether a second configuration was run at all. If the answer is no, the claim is a single sample and should be read, written and funded as a single sample.
- Two deliberately different configurations against one target produce almost disjoint sets of templates. What do you report?That reachability is configuration-dominated here, so a single run's plateau carries very little information, and that the programme should fund breadth across configurations rather than depth in one.
- Can a saturation curve ever support a coverage number?Only against an enumerable denominator, such as a fixed catalogue of behaviours or seeds you agreed to attempt. It can never support coverage of the space of possible jailbreaks, because that space is not enumerable.
Reporting a plateau as a property of the target is like a fisherman with one net in one bay concluding the ocean is empty. The net stopped catching, which tells you where the net was, not what is in the water.
saying these in an interview costs you the question
- Writing that the target saturated, or that no further jailbreaks exist, on the strength of one search.
- Reporting a saturation curve as a coverage percentage.
- Letting a downstream reader convert an exhausted search into a ship-safe assurance claim without correction.
- Running only one configuration and treating its plateau as the programme's answer.
- Not naming the seed corpus, objective and clustering rule alongside the number.