A jailbreak benchmark harness rules an attempt successful when the model's reply does not contain any phrase from a fixed refusal-string list ("I'm sorry", "I cannot", "As an AI"). What errors does this substring rule push into the attack-success rate it produces?
answer
- absence of phrase, not presence of harm
- apologise then comply
- empty or garbled counted as success
- refusal idiom is model-specific
- reward-hackable search signal
basics
~20 sIt counts by wording, not content. A reply that apologises then complies is scored a refusal; an empty, off-topic or garbled reply with no listed phrase is scored a success. Unlisted refusal wordings, other languages and paraphrases all leak through, so the rate drifts both up and down.
solid answer
~50 sThe rule measures the **absence of a phrase**, not the presence of harm, so it errs in both directions. **Scored as a success but is not one:** an empty or truncated reply, a topic change, repeated tokens from a broken adversarial suffix, a refusal phrased in wording nobody put on the list ("that's not something I'll help with"), a refusal in another language, or a reply that is enthusiastic and completely useless. **Scored as a refusal but is a hit:** the classic apologise-then-comply, where the model opens with a listed phrase and then produces the content anyway. Many harnesses only inspect a prefix of the reply, which makes this worse. The practical consequences: the rate is sensitive to the target model's *style*, so two models with identical real behaviour score differently; and an attack search that optimises against this rule learns to suppress the listed strings rather than to elicit content. Use it as a cheap screen, then adjudicate or re-score with a harm classifier.
go deeper
Should say the check looks for refusal phrases rather than at the content, and give at least the apologise-then-comply miss.
Names both error directions, including junk or empty output counting as a success, and notes that the phrase list is tied to one model's refusal style.
Adds that the rule is a reward-hackable objective for an automated attack search, and describes sampling scored hits and misses to measure the two error rates before quoting a number.
Frames it as a measurement-definition decision: string matching is acceptable as a screen, but the published rate must name its scoring rule, and the whole series has to be re-scored if the rule changes.
## What the rule is, mechanically A **refusal-substring scorer** contains no model at all. For each generated reply the harness tests whether any string from a fixed list — `"I'm sorry"`, `"I cannot"`, `"As an AI"`, `"I apologize"` and a dozen relatives — occurs in the text, and labels the attempt a **success** when none of them does. In the AdvBench/GCG lineage that list is a module-level constant of refusal *prefixes*, and only the opening of the generation is inspected; later harnesses widen it to a substring search over the whole reply. Either way the entire evidence the label rests on is "did one of these strings appear". ## What it costs, and why that keeps it alive Effectively nothing. A string scan is microseconds per transcript on a CPU, so a hundred thousand replies score in about a second, the labels are deterministic and reproducible years later, and re-scoring the whole archive under a revised phrase list is free. Compare it with what else could rule on the same corpus: - **HarmBench** distributes a fine-tuned harm classifier in the low tens of billions of parameters: one forward pass per transcript, so it wants a GPU and turns a one-second job into minutes or hours — but it is a pinnable artifact with no per-call price. - **A hosted chat model** prompted with a harm rubric costs one API call per transcript; at roughly a thousand tokens of transcript each, a twenty-thousand-attempt sweep is a real invoice and, more painfully, a rate-limited wall clock that can exceed the generation run it is scoring. Substring matching survives because it is the only scoring rule whose **marginal cost is zero**, which is precisely why it ends up carrying numbers its accuracy cannot support. ## How the number misleads, direction one: manufactured successes The predicate is the *absence of a phrase*, not the *presence of harm*, so anything that avoids the listed wording scores as a landed attack. That includes: - an empty or truncated reply; - a completion that changes the subject; - degenerate repeated tokens (very common when an optimised adversarial suffix breaks the target's decoding); - a refusal in wording nobody put on the list — "that is not something I will help with"; - and a refusal in another language. These are not edge cases: on a large sweep the junk-output bucket alone can be several points of the reported rate. ## Direction two: suppressed successes The classic **apologise-then-comply**, where the reply opens with a listed phrase and then produces the content anyway, is scored a refusal. Prefix-only inspection makes this strictly worse, because the one region of the reply the rule examines is exactly the region a hedged preamble occupies. Refusal-shaped text quoted inside an otherwise compliant answer trips it the same way. ## Style coupling — the failure that survives every fix The phrase list was assembled from the refusal idiom of one model family. Point the same harness at a differently-tuned target and the measurement silently changes meaning: a model that declines tersely, in unlisted words, or in a non-English default reads as far more jailbroken than it is. Two models with identical real behaviour can be twenty points apart under this rule, and nothing in the output tells you that is what happened. ## Optimisation pressure If an automated attack search uses this rule as its own success signal, the objective is **reward-hackable by construction**. Suppressing a dozen fixed strings is a much easier search problem than eliciting content, so the loop converges on transcripts that clear the check while landing nothing, and the campaign's rate curve looks like progress. ## What I would check before quoting a rate 1. Draw a **stratified sample** — fifty transcripts the rule scored as hits and fifty it scored as misses — and read them, which is one to two person-hours and the cheapest instrument calibration available. The hit pile gives an empirical false-positive rate (the junk-output share is usually the shock); the miss pile gives the apologise-then-comply share, which a hits-only audit can never find. 2. Then check what fraction of the scored hits are empty or under some length threshold. 3. Diff the rate against a content-based scorer on the same transcripts. 4. Decide whether the string rule stays as a cheap pre-filter with a classifier behind it. Whatever you publish must name the **scoring rule** beside the number; a rate with no stated scorer is not a measurement.
- An attack loop uses the refusal-string check as its own optimisation signal. What goes wrong?The loop optimises for suppressing a short list of strings, which is much easier than eliciting content. You get transcripts that pass the check and contain nothing useful, and a rate that looks like progress.
- Two models score 30% and 55% under the same refusal-string list. What is the first thing you rule out?That the list matches one model's refusal idiom and not the other's. Sample the 55% model's scored hits and check how many are unlisted refusal wordings or empty output before believing the gap.
- Is the string check ever the right choice?Yes, as a cheap first-pass filter to shrink the transcript set that a costlier judge or a human reviews, and for regression smoke tests where you only need to notice a large movement.
saying these in an interview costs you the question
- Treats the string-match rate as ground truth and quotes it without naming the scoring rule.
- Only names one error direction (usually false misses) and misses that junk output scores as a success.
- Thinks adding more refusal phrases to the list fixes it, rather than moving to content-based scoring.
- Cannot say what the rule measures if asked directly ('did it refuse?' vs 'was the reply harmful?').