skip to content

A team asks you to approve several days of GPU time plus a large metered-API query allowance for a genetic search that mutates and crosses over jailbreak prompts against a deployed assistant. What must they settle about the fitness signal before you approve, and what result would you accept as a return on that spend?

level: principalimportance: should knowfreq 34%

answer

  1. the spend buys whatever fitness defines
  2. written success criterion first
  3. selector must not be the verifier
  4. pilot: known hits vs benign set
  5. deduplicate the converged population

basics

~20 s

Require a written success criterion: what behaviour counts as a hit, what scores it during the search, and what independently confirms it afterwards. Require a cheap pilot showing the signal separates known hits from known-benign prompts. Accept confirmed, reproducible, verified prompts, not a peak fitness curve.

solid answer

~60 s

The spend is not really buying compute, it is buying whatever the fitness signal defines, so that definition is the thing to review. **Before approval, require in writing:** - The behaviour under test, stated concretely enough that two reviewers agree on a given reply. - The signal that drives selection, and its per-candidate cost against the target. - A separate, stricter verification step whose component takes no part in selection. - A short pilot: does the signal score known successful examples high and a held-out benign set low? If it cannot separate those, no amount of compute will help. - Kill criteria: what pattern of results ends the run early. **As the return, accept** a small number of independently confirmed, reproducible prompts with the replies attached, plus an honest statement of what the run covered. Do not accept a fitness curve, a count of high-scoring candidates, or a percentage derived from the selection score. Also accept a well-argued negative: the signal held up on the pilot, the run spent its budget, nothing confirmed. That is a real, if uncomfortable, result.

go deeper

for a junior

Would ask what counts as success and who checks the results before the run starts.

for a middle

Adds the cost model (population times generations times per-candidate cost) and an independent verification step.

for a senior

Requires a pilot separating known hits from benign controls, written kill criteria, and deduplication of survivors into families before anything is counted.

for a principal

Fixes the acceptance criterion in advance, budgets the reviewer hours as tightly as the GPU hours, accepts a well-argued negative result, and refuses selection-score headline numbers that cannot be compared across runs.

### What the money is actually buying Compute buys generations; generations buy convergence toward whatever the fitness signal rewards. So the request in front of you is not "days of GPU plus an API allowance", it is "several days of convergence toward this definition of success" - and the definition is the only part still cheap to change. Once the run is over, every number will be interpreted generously by the people who spent the budget, so the acceptance criterion has to be fixed in writing before the first call is made. ### The checklist, and what each item costs to satisfy 1. **Definition.** One sentence naming what the target must do for a candidate to count, concrete enough that two reviewers reading the same reply agree. Costs an hour of argument. If it cannot be written, cancel: the run has no objective and its output cannot be triaged. 2. **Signal and per-candidate cost.** What converts a reply into the selection number, and what that costs. A string check adds nothing; a judge model roughly doubles the inference bill and adds its own rate limit to the wall clock. Ask for the arithmetic on the request itself: population x generations x turns-per-candidate x per-call cost, checked against the endpoint's requests-per-minute ceiling as well as its price. Rate limit, not price, is usually what turns a run into days. 3. **Separation of roles.** The component that selects must not be the component that confirms. This single line prevents the most common way these runs end with output nobody can report. 4. **Pilot.** Before the main spend, score a handful of known-successful examples and a held-out benign set with the proposed signal. Overlapping distributions mean the run is pre-doomed. A pilot costs a few dozen calls and an afternoon, and it is the highest-leverage item on this list. 5. **Verification capacity.** Someone has to read the survivors. Reviewer hours, not GPU hours, are usually the binding constraint: a run that hands back 400 candidates nobody can triage has produced nothing, and one whose output is triaged by the same automated judge that selected it has produced worse than nothing. 6. **Kill criteria, written up front.** Fitness pinned at zero for N generations (signal too sparse to select on); most of the population at the ceiling within a few generations (signal too easy); high scorers failing independent verification above a set rate (the judge is being optimised rather than the target). Any of those ends the run instead of extending it. ### Where the numbers will mislead you afterwards Four to refuse in advance, in writing: - **Peak or mean selection score as a headline.** Arbitrary units, produced by the very component the search was optimising against. - **A count of high-fitness candidates presented as jailbreaks found.** That is a count of things the scorer liked, before deduplication and before verification. - **Cross-run or cross-team comparison of fitness.** Different signals define different scales with no common referent, so "better than last quarter" can be entirely a change of scoring. - **Raw survivor counts.** Selection concentrates a population, so the end state is typically many small edits of one prompt. Deduplicate into families before anything is counted, or one result gets reported as hundreds and whoever has to fix it is badly misdirected about the size of the problem. ### What to accept as the return A small number of independently confirmed, reproducible prompts with their replies attached and reproduction steps for your own environment; a coverage statement in the run's own terms - which behaviour, which target configuration and model version, how many candidates evaluated over how many generations; and the deduplicated family count reported alongside the raw survivor count. Also accept a well-argued negative: the pilot showed the signal separated known hits from benign controls, the fitness distribution moved during the run, the budget was spent, nothing was confirmed. That is a real result about that behaviour on that configuration, and it must be written exactly that narrowly - never as "the assistant is robust". ### The distinction that decides whether a null run was wasted If the signal was validated on the pilot and the distribution moved, you bought a bounded negative and you should say so without embarrassment. If fitness pinned at zero throughout, or sat at the ceiling from generation two, the run measured its own signal and you bought nothing at all - and note when that money was actually lost. It was lost at approval time, not during the run, which is precisely why the pilot and the kill criteria belong in the approval rather than in the retrospective.

  • The run finishes with nothing confirmed. How do you decide whether that is a real negative or a wasted run?
    Look at the pilot and the run telemetry. If the signal separated known hits from benign controls and the fitness distribution moved, it is a real negative for that behaviour and configuration. If fitness pinned at zero or at the ceiling throughout, the signal failed and the run measured nothing.
  • Why insist on deduplication into families before counting results?
    Selection concentrates the population, so the survivors are typically many small edits of one prompt. Counting them individually inflates a single result into a large number and misdirects whoever has to fix it.

saying these in an interview costs you the question

  • Approving on the strength of the algorithm's sophistication without asking what is scored.
  • Budgeting compute while ignoring the reviewer capacity needed to triage the output.
  • Accepting a peak or mean selection score as the deliverable.
  • Comparing two runs' scores when their fitness signals were defined differently.
  • Counting a converged population's near-identical survivors as separate results.

context