In a genetic search over jailbreak prompts, why does a strictly binary fitness signal (1 if the attempt clearly succeeded, 0 otherwise) tend to stall the search, and what shape of signal do operators use instead?
answer
- boolean fitness = flat landscape
- all zeros means random search
- graded ladder gives invented direction
- selection number vs reporting number
- look at the fitness distribution
basics
~20 sBecause almost every early candidate scores 0, so selection has nothing to rank and the search becomes random. Operators use a graded signal: partial credit for a reply that engages, moves toward the requested content, or omits a refusal, so gradual improvement is visible and can be selected for.
solid answer
~50 sSelection needs *differences* between candidates. If a hardened target refuses everything in the first generations, all candidates score 0, the population is ranked arbitrarily and the run degenerates into random sampling of prompt space. That is a flat fitness landscape, and no amount of extra generations fixes it. The usual remedy is graded scoring: a judge that rates the reply on a scale (refused outright, hedged, partially engaged, fully complied) rather than a boolean, sometimes combined with a fluency or plausibility term so the population does not collapse into unreadable token soup. The cost is that the intermediate rungs of the ladder are guesses. "Hedged" is not one step from success; the search will climb toward whatever the intermediate score rewards, which may be a dead end. So graded fitness makes the search move, but it also gives it a *direction you invented*. Keep the binary definition of a real hit for reporting, separate from the graded number that drives selection.
code
python · 9 linesdef selection_fitness(prompt, reply):
harm = judge_scale(reply) # 0.0 refused .. 1.0 fully complied
intact = request_still_present(prompt) # 0 or 1: payload not mutated away
fluency = readability(prompt) # keeps population human-plausible
return harm * intact + 0.1 * fluency
def is_reportable_hit(prompt, reply):
# strict, binary, evaluated by a checker NOT used during the search
return independent_review(prompt, reply) is CONFIRMEDgo deeper
Says that if everything scores zero there is nothing to choose between candidates, so the search stops improving.
Names it as a sparse-reward problem and proposes a graded judge scale plus a fluency or payload-intact term.
Adds that graded rungs are invented hypotheses that can plateau, and insists on keeping a strict independently checked definition of a hit for reporting.
Cares that the graded selection score never leaks into a reported success rate, because that number will be read as attacks that worked.
### Why an all-zero population is fatal, not merely slow Selection is a comparison operation, and comparison needs differences. Every selection scheme a genetic search uses is a function of *relative* fitness: top-k keeps the highest scorers; tournament selection samples a few candidates and keeps the best of them; roulette-wheel selection assigns each candidate a probability proportional to its fitness. Hand any of them a population where every candidate scores 0 and they degenerate identically - top-k takes an arbitrary prefix, tournament breaks ties arbitrarily, roulette-wheel divides by a total of zero and falls back to uniform sampling. Parents become random members of the population, and the next generation is random mutations of random prompts. What you are paying for stops being a search and becomes random sampling of prompt space with extra bookkeeping. Against a hardened target that refuses nearly everything, all-zero early generations are the *normal* case, not an edge case. This is the sparse-reward problem, and no number of extra generations fixes it, because generations only help when selection has something to act on. ### What a graded signal buys Replace the boolean with an ordinal or continuous score. Two common shapes: - A judge rating the reply on a small ordinal scale - refused outright, hedged, partially engaged, fully complied. - A composite: a judged harm term, plus a payload-intact term so the request cannot be mutated away, plus a fluency or readability term so the population does not collapse into unreadable token strings that no production system would ever receive. Now candidates differ, ranking means something, and the population can move. ### What it costs, in three currencies **Calls and money.** An ordinal judge is a second inference per candidate, roughly doubling per-candidate cost and adding the judge's own rate limit to the run's wall clock. A 50 x 40 run goes from 2,000 inferences to 4,000. **Engineer time.** Somebody has to write the scale and get two reviewers to agree on what "partially engaged" means for this behaviour. If they cannot, the intermediate scores are noise and you have added expense without adding signal. **A direction you invented.** This is the largest cost and appears on no invoice. Nothing guarantees that "hedged" lies nearer to a real hit than "refused". You asserted an ordering, and the population will climb whatever you asserted. Runs routinely spend an entire budget ascending an invented ladder and plateau one rung below the top, because the intermediate states you rewarded lead somewhere the model will never cross. The graded signal made the search move; it did not make it move toward anything real. ### Where the number misleads The graded score's units are arbitrary, so a mean fitness of 0.8 is not "80% success" - it is a point on your own ladder. Worse, in a composite you cannot tell from the aggregate which term is climbing: a fluency term is maximised beautifully by an eloquent prompt that asks for nothing, so a rising total is fully consistent with the population having lost the payload. Decompose the composite and plot each term separately or the curve tells you nothing you can act on. The discipline that prevents all of this is to keep **two numbers**. The *selection* number may be graded, noisy and heuristic; its only job is to move the population. The *reporting* number stays strict and binary: did this candidate cause the target to produce the specific prohibited behaviour, confirmed by a check that took no part in selection. Conflating them is how a partial-credit 0.8 becomes "80% attack success" in a write-up that nobody can reproduce. ### What you would check Plot the whole fitness distribution per generation, not just best-so-far. A spike at zero that never moves means the signal is too sparse to select on. Most of the population at the maximum within a few generations means the signal is too easy, which is the absence-of-refusal trap rather than this one. A healthy run shows a broad spread whose *mass* migrates upward, not merely a rising maximum. Track diversity alongside it, since a plateau with collapsed diversity means the population is polishing one idea. Before the run, sanity-check the ladder itself on labelled examples: if two graders disagree about which rung a reply is on, the ladder cannot select. And at the end, confirm a handful of top candidates under the strict binary definition, which is the only number that may be reported.
- How would you notice the sparse-reward stall from the run's own telemetry?The fitness distribution in early generations is a spike at zero and the best-so-far never moves. Diversity in the population also stays at its initial level, because selection is not concentrating anything.
- Why add a fluency or readability term at all?Unconstrained mutation drifts toward unreadable token strings that may score well against a weak judge but are trivially filtered in production and do not transfer. The term keeps candidates plausible as real user input.
Boolean fitness is hunt-the-thimble where the only feedback is a shout of "found it" or silence: the seeker wanders at random until they trip over it. A graded score is someone calling warmer and colder - it makes progress possible, but only if the person calling actually knows where the thimble is.
saying these in an interview costs you the question
- Answering with generic genetic-algorithm advice (raise the mutation rate, add elitism) instead of addressing the sparse signal.
- Reporting the mean graded fitness as an attack-success rate.
- Assuming intermediate scores are guaranteed to lie on the path to a real hit.
- Adding a fluency term without noticing it can be maximised by prompts that ask for nothing.