skip to content

A genetic prompt search uses an automated judge model to rate whether the target model's reply is harmful, and uses that rating as fitness. Over a long run the rating climbs steadily, but manual spot-checks of the top candidates find nothing actually harmful. What is happening, and how do you change the setup?

level: seniorimportance: must knowfreq 52%

answer

  1. you attack whatever you optimise
  2. judge blind spots, not model weakness
  3. rating climbs, review finds nothing
  4. second rater that never selected
  5. no verified hits is not robustness

basics

~20 s

The search is optimising the judge, not the target. Any judge has blind spots, and selection pressure finds them: replies that look harmful to the rater but are not. Fix it by verifying top candidates with a different checker that never drove selection, and by holding out judged examples to measure the rater.

solid answer

~60 s

Whatever you put in the fitness slot becomes the thing under attack. Running thousands of generations against a fixed automated rater is an adversarial search *against that rater*, so the population converges on its failure modes: surface features that correlate with harm (alarming vocabulary, confident instructional tone, a numbered list) without the prohibited content being present. The telltale signs are exactly what you describe: a smooth fitness climb, low diversity at the top, and manual review finding nothing. A useful confirmation is to score the top candidates with a **different** rater; if the second one disagrees sharply, you were climbing the first one's artefacts. What to change: keep a labelled held-out set so you know the rater's own error rate before trusting it; use a strict independent verification step for anything that will be reported; rotate or ensemble the rater so no single one is optimised end-to-end; and cap how long you let a run push against one fixed rater. The final report counts confirmed items, never the rater's peak score.

go deeper

for a junior

Recognises that the automated rating may be wrong and that someone should read the actual replies.

for a middle

Names it as optimising the proxy: the search finds inputs the rater misjudges, and the fix is independent verification.

for a senior

Separates the selecting component from the verifying one, measures the rater on held-out labels, rotates or ensembles it, and time-boxes runs against a fixed rater.

for a principal

Insists the write-up reports confirmed items only and states plainly that a run which measured its own rater says nothing about the target's robustness.

### What the run is really optimising A *judge* - equivalently a scorer, grader or detector - is itself a model with an error rate: some replies it rates harmful are not, and some it rates safe are. Put its rating in the fitness slot and the genetic search becomes a black-box adversarial search **against the judge**, funded for as many generations as you pay for. It sees only the judge's output, it varies inputs, it keeps whatever scores higher. Its optimum is therefore not "maximally harmful reply" but "maximally *rated* reply", and those two coincide only in the region where the judge is accurate. A search always takes the cheapest route upward, and the cheapest route runs straight through the judge's blind spots. Producing text carrying the surface features a rater keys on - alarming vocabulary, a confident instructional register, a numbered list, a technical-sounding preamble - is far easier than getting a hardened target to actually emit prohibited content. So that is where the population goes. ### Why long runs specifically Early generations sit in the region where the judge is roughly right, so early climbing is honest and reassuring. Divergence grows with generations, because each round of selection acts on the residual error. A smooth, monotone climb over thousands of generations against a *fixed* judge is itself the warning sign: real target behaviour is lumpy - a defence either holds or it does not - while a proxy's error surface is smooth and can be climbed continuously. ### Diagnostics, cheapest first - **Read the top-ranked replies.** Free, immediately decisive, and routinely skipped in favour of watching the fitness curve. - **Re-score the top candidates with a judge that took no part in selection** - a different model, a different rubric, or a stricter prompt. Sharp disagreement is the signature. Agreement is weak evidence but not nothing. - **Score a labelled held-out set with the selecting judge** to obtain its false-positive rate on this behaviour. Do this before the run, so you know what error rate you handed the search. - **Track population diversity.** Collapse onto a single phrasing family alongside a rising score is consistent with one artefact being exploited. ### What it cost Everything the run spent - target calls, judge calls (double the inference bill of a string-scored run), GPU or wall clock under rate limits, and the reviewer hours now being spent finding nothing - bought a characterisation of your judge. That is genuinely useful feedback about the judge, and it is not the deliverable anyone asked for, and it was not budgeted. A multi-day run that ends here has usually consumed the engagement's whole automated-testing allowance. ### The two misreadings, and the second is the dangerous one **Peak rating as a result.** "Our search reached a 0.94 harm score" is a statement about a rater that was under attack for the entire run. Its top-scored items are exactly where its false positives concentrate, so the ratings that look best are the ones least worth trusting. The ranking is inverted from what intuition suggests. **No confirmed hits read as robustness.** This is the one that causes real harm. The run measured the judge; it produced no valid measurement of the target in either direction. Writing "the assistant withstood 40,000 automated attack attempts" from this run manufactures an assurance that will be quoted in places you cannot later correct. The honest write-up says the signal failed, no findings were produced, and here is what will be changed before the next run. ### Changing the setup Split the roles: one component drives selection, a different and stricter one decides what may be written up, and never the same component in both slots. Ensemble or rotate the selecting judge across generations so no single decision boundary can be tracked end to end - this raises cost and injects noise into selection, so it is a trade rather than a free fix. Time-box runs against any fixed judge and re-baseline that judge on its labelled set between runs. Sample a small share of high-scoring candidates for human review continuously, every few generations, so a widening gap between rater and reviewer surfaces within hours instead of after the budget is gone. ### What not to do Do not simply raise the score threshold: the search re-optimises against the new boundary, so you have moved the exploited region rather than removed it, and made the signal sparser at the same time. Do not add the failing candidates to the judge's training data and continue the same run - you have now co-evolved the judge with the population, and fitness values before and after are not comparable. And do not quietly report the run as a negative result about the target.

  • Why does raising the score threshold not solve it?
    The search re-optimises against the new boundary. You have moved the exploited region, not removed it, and you have made the signal sparser at the same time.
  • Does a run like this tell you anything about the target at all?
    Almost nothing. It measured the rater. You can only say that this search, with this signal and this budget, produced no verified hits, and note the rater failure that invalidated the signal.
  • How can the run itself flag the problem early rather than at the end?
    Sample a small share of high-scoring candidates for independent review every few generations. A widening gap between the rater and the reviewer is the early warning.

Paying a search to maximise one examiner's marks for thousands of rounds does not produce better essays; it produces a very precise map of that examiner's habits. The candidates that come out on top are the ones that fooled the marker, not the ones that did the work.

saying these in an interview costs you the question

  • Concluding the target is robust because the run produced no confirmed hits.
  • Raising the rating threshold and re-running as the whole fix.
  • Using the same automated rater for both selection and final verification.
  • Reporting the peak rating as a result without any human or independent confirmation.

context