skip to content

A white-box token search appends a tuned suffix to a request and marks a trial successful when the model's reply begins with a preselected affirmative phrase. Why is that success criterion only a proxy, and what do you check before recording the trial as a real result?

level: juniorimportance: must knowfreq 60%

answer

  1. prefix match is not compliance
  2. loss covers only the target tokens
  3. Goodhart on the in-loop objective
  4. re-judge the full continuation
  5. greedy versus sampled decoding

basics

~20 s

Because the check only reads the reply's first few tokens. A model can open with the agreed phrase and then refuse, stall, or produce useless text. The search optimises exactly that prefix, so it overfits it. Before recording a hit, read or grade the whole continuation.

solid answer

~50 s

The optimiser's loss is defined over the target completion's tokens only, so it is rewarded for producing that opening and for nothing that follows. That is a textbook Goodhart setup: the measured thing and the thing you care about come apart. Prefix matching fails in both directions. **False positives**: the reply opens with the phrase and then refuses, repeats the prompt, or emits filler — common, because the optimiser pushed only the first tokens' probabilities. **False negatives**: the model does comply but words its opening differently, so a strict string match misses it. The working practice is to treat prefix match as a cheap gate during the search and to re-verify every gated hit with a separate pass over the full completion — a human reader or a grading model that judges the whole reply, not its first line. Re-verification also has to fix the decoding settings, because a suffix tuned under greedy decoding may not reproduce when the reply is sampled.

go deeper

for a junior

Should say that the check only looks at the start of the reply and that the rest must be read before calling it a jailbreak.

for a middle

Adds that the loss is defined over the target tokens, so the optimiser overfits the prefix, and that the in-loop gate must be followed by a judgement over the full completion.

for a senior

Also fixes the decoding settings, reports success as a rate over samples, and insists the matcher definition travels with any quoted number.

for a principal

Frames it as a measurement-definition problem: prefix-match rates are not comparable across teams, so the engagement standard must pin target string, matcher and decoding before anyone quotes a figure.

**What the search is actually minimising.** A white-box token search — the family the published GCG method belongs to — holds the request fixed, appends a short run of tokens (the *suffix*), and edits those tokens to drive down a loss. "White-box" means the operator has the weights on hardware they control and can run a backward pass through the model; "token search" means the variable being edited is a sequence of vocabulary ids, not free-form English. The loss is the cross-entropy of a *target completion* — a string the operator fixes in advance, classically a short affirmative opening — given prompt-plus-suffix. Concretely: the summed negative log-probability the model assigns to those exact target tokens, in that exact order, at those exact positions. Read that definition twice, because the entire problem lives inside it. The loss mentions the target tokens and nothing else. If the target is N tokens long, token N+1 of the reply contributes exactly zero to the number being minimised. The optimiser is not neglecting the rest of the answer out of carelessness; the objective is literally silent about it. **Why prefix matching became the in-loop success test.** The quantities the test needs are already computed. A single search step is a backward pass for the ranking signal plus a batch of candidate forward passes — often hundreds — and a run is hundreds of steps. Calling a separate grading model on every candidate would multiply that by the judge's own cost and latency, which nobody pays. So the in-loop gate greedily generates a few tokens and string-matches them against the target. It is free, and it is perfectly aligned with the loss — which is exactly why it inherits the loss's blind spot rather than correcting it. **The two directions it is wrong.** *Over-counting.* The reply opens with the agreed phrase and then refuses, restates the request, drifts into filler or boilerplate, or produces fluent text that is not the requested content at all. This is not an exotic edge case; it is the expected behaviour of an optimiser paid for N tokens. Generation leaves the optimised region and the model's refusal behaviour — never disturbed by the search — reasserts itself. *Under-counting.* The model complies fully but opens in different words, or with a leading newline, or in different case. A strict match scores that a miss, so the true rate is understated in a way that varies by target string. **What verification costs.** The prefix gate is free; believing it is not. Every gated hit needs a pass over the whole reply, and under sampled decoding a single reply is one draw, so you need several samples per trial before "success" is a rate rather than an anecdote. A sweep over 100 behaviours at 10 samples each is 1,000 full generations plus 1,000 grader calls — small money against the GPU hours of the search itself, but real wall-clock, and the grader still has to be validated against human reads on a sample, which is engineer hours, not compute. Budget the triage explicitly: teams routinely fund the search and forget the verification, then quote the free number. **Where the number misleads.** "87% attack success rate" is not one metric. It can mean exact prefix match on the chosen phrase, case-insensitive match, match anywhere in the reply, or a grader's verdict on the full answer — four different numbers from the same run, and the spread between them is often tens of points. Because matcher strictness is a free parameter that nobody standardises, two write-ups can rank the same two checkpoints in opposite orders without either model differing at all. The second misleading move is decoding: a suffix tuned so the target tokens win the argmax under greedy decoding is a claim about greedy decoding only. Served at a nonzero temperature, a refusal token can still be drawn at the first position, and the rate falls — sometimes to near zero — with no change to the weights. **What to check before recording a hit.** - Read or grade the *full* continuation, and confirm the content is genuinely the disallowed thing rather than a confident hallucination the grader liked. - Re-run at the decoding settings the deployment actually uses, over several samples, and report a rate. - Record the target string and the exact matcher next to every figure; a rate without them is not comparable to anyone else's. - Sample-check the grader against a human read, so you know its own false-positive rate before you inherit it. ``` prefix_match(first_tokens(reply), TARGET) -> free in-loop gate; not a result judge_or_human(full_reply) over N samples -> the only thing you may report ``` A candidate who treats those two lines as the same measurement has read about the method rather than run one.

  • Why can the same suffix succeed under greedy decoding and fail when replies are sampled at a nonzero temperature?
    The optimisation only has to make the target tokens win the argmax. Sampling can still draw a refusal token, so success has to be measured as a rate over several samples at the deployment's decoding settings.
  • How can prefix matching also under-count successes?
    The model may comply while opening in different words. A strict match on the chosen phrase scores that as a miss, which is one more reason the reported number should come from a grader over the whole reply.
  • Two reports quote very different success rates for the same search on the same weights. What is the first thing to compare?
    The target string and the matcher — exact versus loose versus graded — followed by the decoding settings and the number of samples per trial.

Marking a jailbreak by its opening line is like judging a contractor by the fact that they said "sure, no problem" on the phone. The words you optimised for are the words you got; whether the work happened is a separate question you have to go and look at.

saying these in an interview costs you the question

  • Calls a prefix match a confirmed jailbreak without reading the continuation.
  • Quotes a success rate without saying it is a prefix-match rate.
  • Treats one greedy sample as proof the suffix works.
  • Assumes a longer target phrase automatically makes the number more honest without checking the optimisation got there.

context