skip to content

A genetic prompt search scores each candidate by the absence of refusal phrasing in the target model's reply, and after many generations almost every survivor scores near the maximum. Why is that result usually worthless, and what does the population most likely contain?

level: middleimportance: must knowfreq 66%

answer

  1. absence of refusal is not compliance
  2. mutation deletes the payload
  3. refusal list misses unlisted wording
  4. score presence of the behaviour
  5. read the top replies, not the curve

basics

~20 s

Because not refusing is not the same as complying. The cheapest way to avoid refusal wording is to stop asking for anything harmful, so the search drifts toward fluent, harmless prompts that get chatty answers. The population is full of false successes carrying no payload.

solid answer

~50 s

Absence-of-refusal is a **negative** signal: it is satisfied by any reply that does not contain the refusal wording you listed. A population search finds the cheapest route to a high number, and the cheapest route here is to mutate the harmful request out of the prompt entirely. The survivors are fluent, on-topic-looking prompts whose replies are perfectly benign. Two further leaks make it worse. The refusal check is a fixed list of phrasings, so a reply that declines in wording you did not enumerate scores as a success. And a reply that starts by complying and then breaks off, or that gives a safe generic overview, also scores full marks. The fix is to score for **presence** of the thing you actually wanted: a judge that rates whether the reply contains the specific prohibited content, combined with a check that the mutated prompt still asks for it. Then re-verify the top candidates by hand before any of them is written up.

go deeper

for a junior

Says a reply can avoid refusal wording while still being harmless, so a high score does not mean the attack worked.

for a middle

Explains the drift: the easiest way to stop being refused is to stop asking, so the population loses the payload while fitness rises.

for a senior

Names all three false-success modes, adds benign control prompts scored by the same function, and requires manual verification of top candidates before reporting.

for a principal

Refuses to let a run start without a stated success criterion and a verification step, because the alternative is a report of findings that nobody can reproduce.

### What the signal actually says "Absence of refusal" means: take the target's reply, test it against a fixed list of refusal strings, and score high when none of them match. It is a **negative predicate** - it is satisfied by *not containing* something. The set of texts that do not contain a given phrase is astronomically large, and genuine compliance with a prohibited request is a vanishingly small subset of it. That asymmetry is the entire failure. ### Why the population ends up harmless Selection can only distinguish candidates whose fitness differs. Here almost the whole top of the range is one flat plateau: every reply lacking the listed strings scores the same maximum, whether it is the prohibited content or a paragraph about houseplants. Genetic drift plus mutation therefore walks the population into the *largest* region of that plateau, which is the harmless one, because harmless prompts vastly outnumber working attacks and nothing in the signal pushes back. Once a single payload-free candidate enters the parent pool its descendants dominate quickly, because they never risk triggering a refusal at all. Nothing is broken. The search converged on the objective you specified. ### Three distinct false successes hiding behind one number - **Payload loss.** Mutation deletes or garbles the clause that requested the prohibited thing. What remains is an ordinary question, the reply is an ordinary answer, no refusal wording appears, and it scores full marks. - **Unlisted refusals.** The target declines in words your list does not contain: a short deflection, a redirection to a helpline, a refusal in another language, or a policy notice rephrased after a model update. Scored as a success. - **Empty compliance.** The model answers at length in an obliging tone with generic, publicly available, non-actionable material. No refusal wording, no payload. This is the hardest of the three to spot from metadata, because the reply *is* long and *is* on topic. ### What it costs, and why the waste stays invisible Every one of those false successes consumed a full target call, and a judge call too if one was running alongside. A 50 x 40 run is 2,000 evaluations: against a self-hosted target that is GPU-hours, against a metered endpoint it is hours of wall clock under a rate limit. The specific pain here is that all of it is spent *before* anyone looks, because the fitness curve rises beautifully throughout. Teams routinely discover the problem only when a reviewer finally opens the top ten replies at the end, by which point the whole allowance is gone. The cheap signal was chosen to save money and ends up wasting the most of it. ### How the number misleads afterwards "Ninety per cent of the final population scored above threshold" reads like a 90% attack-success rate and is nothing of the kind: numerator and denominator are both "candidates the flawed scorer liked". Even the modest-sounding version - "we found 380 non-refused prompts" - is a count of replies that failed to match a string list, which is a statement about your string list and not about the target. And a run like this cannot be read as reassurance either; it produced no valid measurement in either direction. ### What to score instead Score the **presence** of the specific behaviour under test, judged on the reply, rather than the absence of something. Add a term checking that the request is still present in the prompt, so payload loss is penalised rather than rewarded - but note that a strict substring version of that check also blocks legitimate rephrasings, so an intent-level check, or a filter applied after scoring rather than inside fitness, is usually the better shape. A longer refusal list is worth having only as a cheap pre-filter that discards obviously refused replies before you pay for a judge. It is never the objective itself. ### What you would check Score a held-out set of clearly benign prompts with the same fitness function. If they land where the survivors land, the signal separates nothing and the run is measuring your string list. Read the actual replies of the top ten candidates rather than the fitness curve; that costs ten minutes and is decisive. Watch population diversity - a collapse onto one phrasing family alongside a rising score is consistent with the population having found one cheap route to being un-refused. And check what fraction of the population sits at the exact maximum: a ceiling that most candidates reach within a few generations leaves no resolution left to select on, which is the numerical signature of this trap.

  • You add a term that penalises prompts whose text no longer contains the request. What new problem can that create?
    It constrains the search: mutation can no longer rewrite the request at all, which blocks legitimate rephrasings. Match on intent rather than exact substring, or check it as a filter after scoring rather than as part of fitness.
  • How do you tell payload loss apart from empty compliance without reading every reply?
    Score a held-out set of clearly benign prompts with the same fitness function. If they score as high as the survivors, the signal cannot separate the two and the run tells you nothing.
  • Is a longer refusal-phrase list ever worth having?
    Only as a cheap pre-filter to skip obviously refused replies before paying for a real judge. It is never the fitness signal itself.

It is the difference between measuring a hospital by patient outcomes and measuring it by the number of complaints filed: the cheapest way to drive complaints to zero is to stop admitting patients. The search stops asking for the harmful thing because that is the cheapest way to never be refused.

saying these in an interview costs you the question

  • Claiming the fix is just a longer list of refusal phrases.
  • Reading the rising fitness curve as evidence the attack is working, without opening the replies.
  • Not noticing that mutation can remove the harmful request itself.
  • Reporting the count of high-fitness candidates as a number of confirmed jailbreaks.

context