In a population-based (genetic) prompt search that mutates and crosses over candidate jailbreak prompts against a target model, what job does the fitness function do, and why is choosing it the operator's decision rather than the algorithm's?
answer
- fitness = the objective, written as a number
- mutation and crossover are blind
- search maximises what you hand it
- cost = population x generations x per-call
- high fitness is a lead, not a result
basics
~20 sThe fitness function scores each candidate prompt's result and decides which prompts survive to be mutated and crossed over. The search is indifferent to meaning: it maximises whatever number you hand it. So the operator's choice of that number defines what the run can possibly discover.
solid answer
~50 sA population search keeps a pool of candidate prompts, sends each to the target, converts the target's reply into a **single number**, and lets the highest-numbered candidates parent the next generation through mutation and crossover. Mutation and crossover only shuffle text; the fitness function is the only place where the operator's definition of *success* enters the loop. That is why picking it is a red-team decision, not a maths decision. The algorithm will happily drive fitness to its ceiling whether the number means "the model produced the prohibited content" or merely "the reply did not start with an apology". Two runs with identical mutation operators and identical seeds produce completely different populations if their fitness signals differ. Practically: write down what a hit means, decide who or what decides it, and only then start burning target queries.
go deeper
Says the fitness function scores candidates and drives which ones survive, and that the operator chooses what it measures.
Adds that the fitness signal is the only place attack semantics enter the loop, and that its per-candidate cost sets the whole run's budget.
Frames fitness as the run's success criterion and insists that top candidates are independently re-verified before anything is reported.
Treats the fitness definition as an organisational decision: it determines what the programme can claim it tested, and it must match the policy the deployed system is held to.
### The loop, and the one place meaning enters it A population search over prompts holds a *population*: a fixed number of candidate prompt strings, typically tens to low hundreds. One *generation* is four steps. 1. **Evaluation.** Every candidate is sent to the *target* - the model or endpoint under test - and its reply is collected. 2. **Scoring.** Each (prompt, reply) pair is converted into a single number, the *fitness*. 3. **Selection.** Candidates are ranked by fitness and a top slice is kept as parents. 4. **Variation.** The next population is built by *mutation* (perturbing one parent: swapping a word, inserting or deleting a clause, changing punctuation) and *crossover* (splicing text from two parents). Steps 1, 3 and 4 are identical no matter what you are searching for; they are string manipulation and sorting, and they carry no security semantics whatsoever. Step 2 is the only step in which a human's definition of "the attack worked" appears anywhere in the loop. That is the whole answer to the question: the fitness function is the attack objective written as a number, and the algorithm has no other channel to your intent. One vocabulary note, because this tree uses four words for the same job. The component that turns a reply into a number is a *scorer* in PyRIT, a *detector* in garak, a *grader* in promptfoo, and a *judge* in most papers. In a genetic search it occupies the fitness slot. Say which one you mean, because their contracts differ and "the scorer said it worked" is not a statement anyone can check. ### What a generation costs Per candidate you pay at least one call to the target, plus whatever the scoring itself costs: | fitness signal | cost per candidate beyond the target call | how fast it can run | |---|---|---| | substring or regex refusal check | effectively zero | unlimited | | judge model call | a second inference, often the same order as the target's | bounded by the judge's own quota | | a human reading the reply | minutes of an expert | cannot run per generation at all | Total target calls = population x generations, multiplied again by turns per candidate if each evaluation is a conversation rather than a single exchange. A modest run - 50 candidates, 40 generations, one exchange each - is 2,000 target calls; adding a model judge makes it 2,000 more inferences. On a metered chat API that is usually tens of dollars, not thousands, which is exactly why people under-plan it. The costs that actually bite sit elsewhere: the endpoint's requests-per-minute limit, which turns 4,000 calls into hours of wall clock; GPU hours when the target is self-hosted and each call is seconds of generation; and the human hours needed to triage whatever comes out. It is entirely normal for a run to be cheap in dollars and unaffordable in reviewer time. ### Where the number misleads Fitness is a selection device, and it is systematically the wrong thing to report. Three specific misreadings, all common: - **Mean or peak fitness read as an attack-success rate.** A composite score of 0.8 is not "80% of attempts succeeded". It is a point on a scale you invented, with no external referent. - **Fitness compared across runs.** Two runs' numbers are incomparable unless their signals are byte-identical, because each signal defines its own scale. "Higher than last quarter" can be entirely a change of scoring. - **Fitness treated as evidence for an individual candidate.** By the end of a run the survivors are precisely the candidates the scorer rates highest, which - if the scorer has any error at all - is exactly where its false positives concentrate. The score was produced by the same component the search spent its whole budget pushing against. A high-fitness candidate is a lead to verify, never a confirmed finding. ### What you would check State the fitness function in one sentence: what is fed to it, what it returns, and what a maximum means. If nobody on the team can write that sentence, the run has no success criterion and its output cannot be triaged afterwards, no matter how much compute it consumed. Then ask the decisive question: can this signal be driven to its ceiling without the behaviour you care about ever occurring? For an absence-of-refusal check the answer is trivially yes, and that is the whole reason the next question in this leaf exists. Run a pilot before the main spend - score a handful of known-successful examples and a held-out set of clearly benign prompts with the proposed signal, and look at whether the two distributions overlap. Multiply population x generations x per-candidate cost and check it against the endpoint's rate limit, not only its price. Finally, name the component that will *verify* the survivors and confirm it is not the component that selected them.
- If mutation and crossover are unchanged, can two runs of the same search produce different kinds of prompts?Yes. The variation operators only propose text; the fitness signal decides what survives, so changing it changes the population the run converges on.
- Where does the cost of a generation actually go?One call to the target per candidate evaluated, plus the cost of scoring the reply. A model-based scorer roughly doubles the per-candidate cost; a string check adds nothing.
saying these in an interview costs you the question
- Describing genetic algorithm mechanics (selection, elitism, mutation rate) without ever saying what is being scored.
- Assuming a library ships one obvious correct fitness function for jailbreak search.
- Treating the final fitness score as an attack-success rate that can be reported as-is.
- Not knowing that every fitness evaluation costs at least one call to the target.