You want a live saturation signal for an automated jailbreak search: distinct successful templates per thousand queries. How do you count 'distinct', and why does clustering the successful prompts by surface text similarity mislead you in both directions?
answer
- groups over queries spent
- over-split: optimiser noise, one mechanism
- under-split: shared scaffolding, two mechanisms
- lineage and response fingerprint beat edit distance
- freeze the rule before the run
basics
~20 sGroup the successful prompts and count groups per thousand queries spent. Surface-text similarity misleads both ways: a search mutates wording heavily while keeping one mechanism, splitting one template into many groups, and two unrelated mechanisms that share boilerplate collapse into one. Cluster on the mechanism and on what the target actually did.
solid answer
~50 sThe numerator is groups, the denominator is queries spent, and the whole value of the metric lives in the grouping rule. Surface similarity is the tempting rule and it fails in two opposite directions. **Over-splitting:** a token-level optimiser produces successes whose visible text is near-random and mutually dissimilar, yet every one is the same mechanism attached to the same request, so edit distance reports dozens of discoveries where there is one. **Under-splitting:** an attacker-model loop wraps everything in the same scaffolding, so two genuinely different mechanisms sit inside identical boilerplate and collapse into one group. A workable rule keys on the invariant rather than the wording: what stays constant across the successful children — the mechanism and the target behaviour elicited — with text similarity used only as a cheap pre-filter. Freeze the rule before the run, apply it identically across campaigns you intend to compare, and store the raw prompts so grouping can be recomputed when the rule changes.
go deeper
Knows the metric is groups divided by queries spent and that identical-looking prompts should be one group.
Explains both failure directions of surface similarity, proposes lineage or mechanism as the key, and insists the rule is fixed before the run.
Adds the online-versus-end-of-run distinction, the per-campaign threshold problem, judge independence, and the cluster-size distribution as a diagnostic on the rule itself.
Treats the grouping rule as a governed artefact: versioned, shared across campaigns, and immutable within a run so that stop decisions across a programme remain comparable.
## Two jobs that look like one Grouping near-identical successes so a human reading a finding list is not drowned is an *editorial* job: done once, at the end, with a person in the loop, optimised for readability. A saturation signal is an *operational* job: computed continuously during the run, cheaply, on every new accepted prompt, so the operator can watch the curve flatten while there is still allowance left to redirect. They can share a rule, but they have different requirements, and building only the editorial one is the common mistake — by the time it runs, the money is gone. The operational version needs a grouping rule with three properties: cheap enough to run online, stable when the search changes its wording, and independent of the judge's confidence value so that a drifting judge does not silently move the cluster count. ## Candidate grouping keys | grouping key | what it costs at runtime | what it gets right | how it fails | |---|---|---|---| | Seed or strategy lineage — which parent the candidate was mutated from | free, if the search records lineage at generation time | one mechanism stays one group no matter how far the wording drifts | two different seeds converging on the same mechanism stay apart, so it over-counts | | Target-response fingerprint — the shape of what the model actually produced | one extra pass over each accepted transcript | catches exactly the convergent discovery lineage misses | when the elicited behaviour has one canonical output shape, genuinely different routes merge | | Prompt embedding proximity | one embedding call per accepted prompt, plus incremental centroid comparison | language-aware, tolerant of paraphrase | shared scaffolding raises baseline similarity; high-entropy optimiser output destroys the geometry | | Edit distance or n-gram overlap | negligible | catches exact and near-exact repeats | everything else | The practical answer is a composite: lineage as the primary key because it is free and near-perfect for the exploitation case, merged by response fingerprint to catch convergence, with text similarity used only as a pre-filter. Note the cost asymmetry — lineage costs *engineer time before the run* (instrumenting the generator to stamp a parent id on every candidate) and nothing after; embedding clustering costs *money during* every run, forever. Instrument first. ## Where the number misleads **Over-splitting.** A token-level optimiser produces accepted prompts whose visible text is near-random and mutually dissimilar, yet every one carries the same request and the same mechanism. Edit distance reports dozens of discoveries where there is one, and the curve keeps rising, so the run never trips its stopping rule and drains the whole allowance while discovering nothing. **Under-splitting.** An attacker-model loop that wraps every candidate in the same role-play preamble and formatting instruction raises pairwise similarity across the *entire* accepted set. Genuinely different mechanisms fall under one threshold, the count flatlines early, and the run stops with allowance unspent. Worse, a threshold tuned on one campaign silently reports saturation on the next, because the constant scaffolding changed and nobody re-tuned. Strip known-constant spans before embedding, or recompute the threshold per campaign. **Judge coupling.** Deriving clusters from the judge's confidence value ties the metric to a component that is itself under test and gets retuned. Two runs then stop being comparable for reasons that have nothing to do with the targets. **Post-hoc tuning.** The tempting move when a run looks unproductive is to loosen the rule until the cluster count rises. A metric chosen after seeing the result is not a metric. The rule, the threshold and the clustering version belong in the run plan next to the query allowance, before anything is spent, and the raw accepted prompts must be stored so grouping can be honestly recomputed when the rule is deliberately revised. ## Cost of computing it Small in absolute terms, and worth sizing anyway. Embedding a few thousand short accepted prompts is fractions of a cent per thousand on a commercial embedding endpoint and milliseconds of compute; naive pairwise comparison at three thousand successes is about four and a half million distance computations, trivial in process. It stops being trivial around six figures of successes, where you must keep incremental centroids rather than recompute a full similarity matrix each time a success arrives. Against that, the cost of *not* computing it live is measured in the fraction of a metered allowance spent after the curve had already flattened — usually far larger than the clustering itself. ## What I would check on a live run - Cluster count and queries spent plotted together, refreshed as successes arrive. - The cluster-size distribution. One cluster holding ninety percent of the successes with a long tail of singletons almost always means the grouping rule, not the target, is doing the talking. - A spot-read of a handful of prompts from the largest cluster, confirming by eye that they really are one mechanism. - Whether the clustering rule version and threshold are recorded on the run alongside the number, so a later reader can tell a genuine plateau from a rule change.
- Your cluster-size distribution is one cluster with ninety percent of the successes and a long tail of singletons. What is the first thing you suspect?The grouping rule, not the target. Either shared scaffolding is collapsing distinct mechanisms into the giant cluster, or the threshold is too loose. Spot-read prompts from that cluster before trusting any saturation reading built on it.
- Why not just key the clusters on the judge's confidence value?Because the judge drifts and is itself under test. A grouping rule that depends on judge output moves whenever the judge is retuned, so two runs stop being comparable for reasons that have nothing to do with the target.
- The search knows which seed each candidate descends from. Why is that not the whole answer?Lineage misses convergent discovery — two different seeds arriving at the same mechanism stay in separate groups and inflate the count. Pair lineage with a response or mechanism fingerprint to merge them.
Grouping successful prompts by how similar their text looks is like sorting keys by the colour of their plastic tags instead of by the lock each one opens. Two keys to the same door land in different piles, and keys to different doors land in the same one.
saying these in an interview costs you the question
- Choosing or loosening the clustering rule after seeing the result.
- Using raw edit distance on the output of a token-level optimiser and reporting each success as a discovery.
- Reusing a similarity threshold across campaigns whose prompt scaffolding differs.
- Deriving the cluster count from the judge's confidence value, so judge drift silently moves the metric.
- Confusing this with end-of-run report grouping and computing it only once, when the allowance is already spent.