skip to content

Genetic Prompt Search

A population search is only as good as its fitness signal: reward the absence of a refusal phrase and it converges on prompts that score well and produce nothing. Interviewers ask what you scored.

on this pageshow

explore

questions

5

In a population-based (genetic) prompt search that mutates and crosses over candidate jailbreak prompts against a target model, what job does the fitness function do, and why is choosing it the operator's decision rather than the algorithm's?

level: juniorimportance: must knowfreq 72%

answer

  1. fitness = the objective, written as a number
  2. mutation and crossover are blind
  3. search maximises what you hand it
  4. cost = population x generations x per-call
  5. high fitness is a lead, not a result

basics

~20 s

The fitness function scores each candidate prompt's result and decides which prompts survive to be mutated and crossed over. The search is indifferent to meaning: it maximises whatever number you hand it. So the operator's choice of that number defines what the run can possibly discover.

solid answer

~50 s

A population search keeps a pool of candidate prompts, sends each to the target, converts the target's reply into a **single number**, and lets the highest-numbered candidates parent the next generation through mutation and crossover. Mutation and crossover only shuffle text; the fitness function is the only place where the operator's definition of *success* enters the loop. That is why picking it is a red-team decision, not a maths decision. The algorithm will happily drive fitness to its ceiling whether the number means "the model produced the prohibited content" or merely "the reply did not start with an apology". Two runs with identical mutation operators and identical seeds produce completely different populations if their fitness signals differ. Practically: write down what a hit means, decide who or what decides it, and only then start burning target queries.

go deeper

for a junior

Says the fitness function scores candidates and drives which ones survive, and that the operator chooses what it measures.

for a middle

Adds that the fitness signal is the only place attack semantics enter the loop, and that its per-candidate cost sets the whole run's budget.

for a senior

Frames fitness as the run's success criterion and insists that top candidates are independently re-verified before anything is reported.

for a principal

Treats the fitness definition as an organisational decision: it determines what the programme can claim it tested, and it must match the policy the deployed system is held to.

### The loop, and the one place meaning enters it A population search over prompts holds a *population*: a fixed number of candidate prompt strings, typically tens to low hundreds. One *generation* is four steps. 1. **Evaluation.** Every candidate is sent to the *target* - the model or endpoint under test - and its reply is collected. 2. **Scoring.** Each (prompt, reply) pair is converted into a single number, the *fitness*. 3. **Selection.** Candidates are ranked by fitness and a top slice is kept as parents. 4. **Variation.** The next population is built by *mutation* (perturbing one parent: swapping a word, inserting or deleting a clause, changing punctuation) and *crossover* (splicing text from two parents). Steps 1, 3 and 4 are identical no matter what you are searching for; they are string manipulation and sorting, and they carry no security semantics whatsoever. Step 2 is the only step in which a human's definition of "the attack worked" appears anywhere in the loop. That is the whole answer to the question: the fitness function is the attack objective written as a number, and the algorithm has no other channel to your intent. One vocabulary note, because this tree uses four words for the same job. The component that turns a reply into a number is a *scorer* in PyRIT, a *detector* in garak, a *grader* in promptfoo, and a *judge* in most papers. In a genetic search it occupies the fitness slot. Say which one you mean, because their contracts differ and "the scorer said it worked" is not a statement anyone can check. ### What a generation costs Per candidate you pay at least one call to the target, plus whatever the scoring itself costs: | fitness signal | cost per candidate beyond the target call | how fast it can run | |---|---|---| | substring or regex refusal check | effectively zero | unlimited | | judge model call | a second inference, often the same order as the target's | bounded by the judge's own quota | | a human reading the reply | minutes of an expert | cannot run per generation at all | Total target calls = population x generations, multiplied again by turns per candidate if each evaluation is a conversation rather than a single exchange. A modest run - 50 candidates, 40 generations, one exchange each - is 2,000 target calls; adding a model judge makes it 2,000 more inferences. On a metered chat API that is usually tens of dollars, not thousands, which is exactly why people under-plan it. The costs that actually bite sit elsewhere: the endpoint's requests-per-minute limit, which turns 4,000 calls into hours of wall clock; GPU hours when the target is self-hosted and each call is seconds of generation; and the human hours needed to triage whatever comes out. It is entirely normal for a run to be cheap in dollars and unaffordable in reviewer time. ### Where the number misleads Fitness is a selection device, and it is systematically the wrong thing to report. Three specific misreadings, all common: - **Mean or peak fitness read as an attack-success rate.** A composite score of 0.8 is not "80% of attempts succeeded". It is a point on a scale you invented, with no external referent. - **Fitness compared across runs.** Two runs' numbers are incomparable unless their signals are byte-identical, because each signal defines its own scale. "Higher than last quarter" can be entirely a change of scoring. - **Fitness treated as evidence for an individual candidate.** By the end of a run the survivors are precisely the candidates the scorer rates highest, which - if the scorer has any error at all - is exactly where its false positives concentrate. The score was produced by the same component the search spent its whole budget pushing against. A high-fitness candidate is a lead to verify, never a confirmed finding. ### What you would check State the fitness function in one sentence: what is fed to it, what it returns, and what a maximum means. If nobody on the team can write that sentence, the run has no success criterion and its output cannot be triaged afterwards, no matter how much compute it consumed. Then ask the decisive question: can this signal be driven to its ceiling without the behaviour you care about ever occurring? For an absence-of-refusal check the answer is trivially yes, and that is the whole reason the next question in this leaf exists. Run a pilot before the main spend - score a handful of known-successful examples and a held-out set of clearly benign prompts with the proposed signal, and look at whether the two distributions overlap. Multiply population x generations x per-candidate cost and check it against the endpoint's rate limit, not only its price. Finally, name the component that will *verify* the survivors and confirm it is not the component that selected them.

  • If mutation and crossover are unchanged, can two runs of the same search produce different kinds of prompts?
    Yes. The variation operators only propose text; the fitness signal decides what survives, so changing it changes the population the run converges on.
  • Where does the cost of a generation actually go?
    One call to the target per candidate evaluated, plus the cost of scoring the reply. A model-based scorer roughly doubles the per-candidate cost; a string check adds nothing.

saying these in an interview costs you the question

  • Describing genetic algorithm mechanics (selection, elitism, mutation rate) without ever saying what is being scored.
  • Assuming a library ships one obvious correct fitness function for jailbreak search.
  • Treating the final fitness score as an attack-success rate that can be reported as-is.
  • Not knowing that every fitness evaluation costs at least one call to the target.

context

open as a page

A genetic prompt search scores each candidate by the absence of refusal phrasing in the target model's reply, and after many generations almost every survivor scores near the maximum. Why is that result usually worthless, and what does the population most likely contain?

level: middleimportance: must knowfreq 66%

basics

~20 s

Because not refusing is not the same as complying. The cheapest way to avoid refusal wording is to stop asking for anything harmful, so the search drifts toward fluent, harmless prompts that get chatty answers. The population is full of false successes carrying no payload.

open as a page

A genetic prompt search uses an automated judge model to rate whether the target model's reply is harmful, and uses that rating as fitness. Over a long run the rating climbs steadily, but manual spot-checks of the top candidates find nothing actually harmful. What is happening, and how do you change the setup?

level: seniorimportance: must knowfreq 52%

basics

~20 s

The search is optimising the judge, not the target. Any judge has blind spots, and selection pressure finds them: replies that look harmful to the rater but are not. Fix it by verifying top candidates with a different checker that never drove selection, and by holding out judged examples to measure the rater.

open as a page

In a genetic search over jailbreak prompts, why does a strictly binary fitness signal (1 if the attempt clearly succeeded, 0 otherwise) tend to stall the search, and what shape of signal do operators use instead?

level: middleimportance: should knowfreq 48%

basics

~20 s

Because almost every early candidate scores 0, so selection has nothing to rank and the search becomes random. Operators use a graded signal: partial credit for a reply that engages, moves toward the requested content, or omits a refusal, so gradual improvement is visible and can be selected for.

open as a page

A team asks you to approve several days of GPU time plus a large metered-API query allowance for a genetic search that mutates and crosses over jailbreak prompts against a deployed assistant. What must they settle about the fitness signal before you approve, and what result would you accept as a return on that spend?

level: principalimportance: should knowfreq 34%

basics

~20 s

Require a written success criterion: what behaviour counts as a hit, what scores it during the search, and what independently confirms it afterwards. Require a cheap pilot showing the signal separates known hits from known-benign prompts. Accept confirmed, reproducible, verified prompts, not a peak fitness curve.

open as a page