An automated jailbreak search has produced three thousand prompts its judge marked successful, but they are heavy rewrites of two underlying shapes. Why is the raw count of successful prompts a poor signal for whether to keep spending the remaining queries?
answer
- count clusters, not prompts
- optimiser stays in its basin
- distinct templates per thousand queries
- watch the derivative, not the level
- volume implies breadth it does not have
basics
~20 sThe raw count measures how often the search repeated itself, not how much it found. Thousands of near-identical rewrites are one weakness discovered many times. The signal that matters is distinct successful templates: when that stops rising while queries keep being spent, the remaining allowance is buying duplicates.
solid answer
~50 sA success count and a discovery count are different quantities. An automated search that mutates a working prompt will keep producing children of that prompt, and every child that still works increments the success count without adding anything a defender must fix separately. Two shapes discovered three thousand times is two results. So the operating metric is **distinct successful templates per thousand queries spent** — count clusters, not prompts, and plot that against queries. A rising curve means the search is still exploring; a flat curve with queries still draining means it is exploiting one basin. The tradeoff is that the number depends entirely on how you define "distinct", so the definition has to be fixed before the run, not chosen afterwards when it flatters the result. The simple story breaks when the two shapes are genuinely different mechanisms with different fixes: then two is a real, useful answer and the volume is noise on top of it.
go deeper
Says that the same jailbreak found many times is still one finding, and that the run should be judged on how many different working prompts it found, not how many total.
Names the metric as distinct successful templates per thousand queries, explains that the optimiser is designed to stay near what worked, and notes the definition of distinct must be fixed up front.
Adds that the decision point is the flattening of new-cluster arrivals, that success counts cannot rank two runs, and that a loose judge inflates and self-similarises the successes at once.
Frames the success count as an incentive problem across a programme: it is the number teams will optimise if it is the number reported, so the programme reports discovery, not volume.
## What the search is doing when it "succeeds" An automated jailbreak search is a loop with three parts. A **generator** produces a candidate prompt; a **target** — the model under test, reached through a metered API or loaded locally — answers it; a **scoring step**, called a judge in some tools and a scorer or detector in others, decides whether that answer counts as a success. Whatever the family, the loop keeps the candidates that scored well and builds the next batch out of them. A token-level optimiser edits a suffix under a numeric signal. An attacker-model loop hands a second language model the last refusal and asks it to rewrite the request. A mutation loop paraphrases and recombines entries in a pool of hand-written framings. All three are hill-climbers, and their entire design is to stay near whatever last worked. That design fixes the shape of the output. Once the loop finds one working shape, the cheapest way to score again is to perturb that shape, so children of the hit arrive in volume and keep arriving. Three thousand accepted prompts descended from two parents is not three thousand results. It is two results, plus a measurement of how reproducible each one is — which is worth knowing, and is a completely different quantity from how much the run discovered. ## What the run costs, in the units you are billed in Queries, not prompts, are the meter, and one *attempt* is usually more than one billed model call. A judged attacker-model loop calls three models per attempt: the attacker that writes the candidate, the target that answers it, and the judge that scores the answer. A campaign described as "ten thousand attempts" is therefore on the order of thirty thousand completions, and a multi-turn loop multiplies the target and judge legs again by turn depth while the target's prompt leg grows with the conversation it is carrying. Against a commercial endpoint every one of those legs is priced per token. Wall-clock is normally set by a tokens-per-minute or requests-per-minute ceiling rather than by compute, so "run it longer" often means "wait longer", not "spend more machine". And the human cost lands after the run: reading three thousand accepted transcripts to discover there were two mechanisms is a day of engineer time that a live cluster curve would have saved. The reason to state all this is that the marginal query has a price, so "it is still finding things" has to be weighed against something rather than treated as self-evidently good. ## Where the number misleads **Budget.** An operator watching successes per hour sees a healthy rate and lets the allowance drain. The rate really is healthy — it is measuring reproduction of a hit already recorded, and each further query buys a duplicate. **Breadth.** "Three thousand successful jailbreaks" reads to any downstream audience as wide coverage of the target's failure modes. The run has deep coverage of two. That gap is exactly where a report turns into a wrong decision in either direction: it can alarm a team about volume that is one weakness, or it can imply a thoroughness the search never had. **Comparison.** Two runs cannot be ranked by success count at all. A run that finds one highly reproducible shape will out-count a run that finds six fragile ones, and the second run is the better result. Any leaderboard built on raw hits pays teams to exploit rather than explore. **The judge underneath.** If the scoring step accepts on a surface cue — a compliance-shaped opening, the mere absence of a refusal phrase — rather than on the behaviour the objective named, false accepts enter the count. Those false accepts are strongly self-similar, so they inflate the total *and* make the run look like clean exploitation of one basin at the same time. A count that is wrong while the curve looks explicable is the worst of the four cases, because nothing in the dashboard flags it. ## What to check - Cluster the accepted prompts and plot **distinct clusters against queries spent**, updated live rather than at the end. The level of that curve is the result; its derivative is the stop signal. - Read the cluster-size distribution. One enormous cluster with a tail of singletons usually indicts the grouping rule rather than describing the target. - Hand-audit a sample of accepted successes: does the response actually contain the behaviour the objective named? Do this before believing anything computed downstream of the judge. - State your denominator out loud. "Hit rate" is not a stable field name across red-team tooling and different tools headline different quantities; say whether you mean accepted attempts over attempts made, or distinct templates over queries spent.
- If the two shapes really are two different mechanisms, is the run a failure?No. Two distinct mechanisms is a real result, and the volume around them is just the optimiser's exploitation. The failure would be reporting three thousand results, or continuing to spend queries that only deepen the same two.
- What would make the success count itself untrustworthy, before any clustering?A judge that fires on surface cues rather than the behaviour under test. Those false accepts are highly self-similar, so they inflate the count and simultaneously make the run look like it is exploiting one basin.
- You have only the success count from someone else's run. What single extra number do you ask for?Distinct successful templates and the definition used to cluster them, together with queries spent. Without a denominator and a clustering rule, the count cannot be interpreted or compared with any other run.
Counting successful prompts is like counting the photographs you took of one landmark and reporting the total as the number of landmarks you visited. The camera was busy; the itinerary never grew.
saying these in an interview costs you the question
- Reporting the number of successful prompts as the run's headline result with no notion of how many are the same thing.
- Treating a high success rate as evidence the search is still productive.
- Comparing two searches by success count without fixing a shared definition of a distinct template.
- Assuming every prompt the judge accepted actually produced the harmful behaviour.