In a gradient-guided adversarial suffix search against a model whose weights you hold locally, why must the optimisation objective be a fixed target completion string rather than a compliance verdict from a judge model, and how does the choice of that string change the run?
answer
- discrete tokens, differentiable loss
- gradient ranks candidates, forwards decide
- judge verdict has no gradient path
- longer target, slower convergence
- improbable target, flat surface
basics
~20 sGradients need a differentiable number. Cross-entropy on a fixed target completion gives one; a judge's yes or no is a discrete label from another model with no usable gradient. The string you pick sets the difficulty: a short generic opener is easy but weak evidence, a long specific one is slower and rarer.
solid answer
~50 sThe search descends the loss of an exact continuation given the prompt plus the candidate suffix. That loss is differentiable with respect to the suffix's token embeddings, which is what lets you rank candidate token substitutions instead of trying the whole vocabulary at every position. A judge verdict is a discrete output of a separate network — no gradient path back to your suffix — so it can only be an outer check, not the thing you optimise. The target string is therefore a design decision with real consequences: - **Short and generic** (a bare agreement to help): easiest to reach, fewest steps, weakest evidence. - **Longer, content-bearing**: a harder optimisation, more steps and GPU hours, but a hit means much more. - **Something the model would essentially never emit**: a near-flat loss surface where the greedy search stalls. You are trading optimisation difficulty against how much a hit proves, so state which target you used with any rate you report.
code
text · 8 linesloss(suffix) = cross_entropy( model(prompt + suffix), TARGET_TOKENS )
for step in range(step_budget):
g = d loss / d onehot(suffix_positions) # ranking signal only
candidates = top_k_substitutions_from(g) # shrink the search set
suffix = argmin_over(candidates, loss) # one forward pass each
if full_reply_judgement(model(prompt + suffix)): # outer check, not the loss
breakgo deeper
Should know the search optimises towards a fixed target text rather than towards an abstract idea of compliance.
Explains that the loss must be differentiable, that a judge model gives no gradient, and that target length and plausibility set how hard the run is.
Adds that the gradient only ranks candidate substitutions while forward passes decide, reads per-token loss curves to diagnose stalls, and treats the target string as part of any reported number.
Treats target-string selection as a standard the team fixes in advance, so rates from different runs and different operators mean the same thing.
**The obstacle the design is working around.** The thing being edited is a suffix of discrete vocabulary ids. You cannot nudge token id 4127 a little way toward token id 4128 — there is no "a little way". Yet the only cheap directional information a neural network gives you is a gradient, which is defined over continuous quantities. The standard resolution is a two-stage step: differentiate a loss with respect to the *continuous* representation of each suffix position — its one-hot vector or, equivalently, its embedding row — and use the resulting vector as a **ranking heuristic** over the vocabulary at that position. Take the top-k candidates that ranking suggests, substitute each into the suffix, and evaluate them with genuine forward passes. Keep the best. The gradient never decides; it only shrinks the candidate set from a full vocabulary of tens or hundreds of thousands of entries down to a batch you can afford to score exactly. That two-stage structure is what forces the objective to be differentiable, and it is what rules out the judge. **Why a judge verdict cannot be the loss.** A grading model that reads a reply and returns "complied" or "refused" produces a discrete label, emitted by a *different* network, from text that was already sampled. Three separate breaks in the chain: sampling is non-differentiable, the label is non-differentiable, and the judge's parameters have no connection to your suffix positions in the first place. There is no path along which "the judge said no" becomes "position 7 should prefer this token". Even a free, instant judge would not help — latency is not the blocker, the missing derivative is. So a judge can only be an *outer* check: run occasionally, on full replies, to decide what you are allowed to report. **What the objective is instead.** A fixed target completion — a specific string, written out in advance — scored by cross-entropy: the summed negative log-probability the model assigns to those tokens given prompt-plus-suffix. That is smooth, cheap, already computed by a forward pass, and differentiable back to the suffix embeddings. It is also, unavoidably, a proxy for the thing you care about, which is a compliant and usable *whole* answer. You accept the mismatch and pay for it with the outer verification pass. **How the target string reshapes the run.** | Target choice | Effect on the run | Effect on the evidence | |---|---|---| | Short, generic affirmation | Fewest steps, cheapest | Weak: proves an opening, not compliance | | Long, content-bearing | More steps, more GPU hours, may not converge | Strong: a hit is close to real compliance | | Text the model would essentially never emit | Near-flat loss surface, greedy search stalls | None — you bought a non-result | | One suffix over several request/target pairs | Higher per-step cost, slower convergence | Suffix less glued to a single request; reusable | Each extra target token is another term in the sum and another constraint the same short suffix must satisfy simultaneously. Naturalness matters as much as length: if the model already assigns non-trivial probability to the target, there is a gradient to follow; if it assigns near-zero probability everywhere in the neighbourhood, the surface is flat and the search wanders. **What it costs.** Per-step cost is one backward pass plus the candidate batch of forward passes, and the batch usually dominates. Doubling target length does not double cost per step, but it typically multiplies the *number* of steps needed, and it raises the fraction of runs that finish the budget with nothing. Since the artefact is bound to the checkpoint it was optimised against, that spend repeats at every fine-tune, merge or quantisation change. **Where the number misleads.** A success rate is only meaningful against the target that produced it. A team quoting a high rate on a bare "Sure, here is" opener and a team quoting a low rate on a long content-specific target are not disagreeing about the model; they are reporting two different experiments. Because target choice is invisible in the headline figure, this is the single easiest way for a suffix-search comparison to be silently wrong. **What to check.** Log per-target-token loss, not just the total — a stalled run is very often one stubborn token, which tells you the target is wrong for this model rather than that the budget is short. Confirm the ranking heuristic is doing work by comparing against random substitutions on a short pilot. Re-verify every hit against the full reply. And put the target string in the artefact, next to the rate.
- If the gradient does not choose the substitution, what is it actually used for?To rank which token swaps at which positions look promising, so the run evaluates a small candidate batch with forward passes instead of the whole vocabulary at every position.
- Why might a run stall even with plenty of steps left?The target may sit in a region the model assigns almost no probability, leaving a near-flat surface, or the greedy candidate step is stuck in a local minimum. Per-token loss curves usually show which target token is the blocker; a restart from a different initial suffix is often cheaper than more steps.
- What changes if you optimise one suffix against several request/target pairs at once?Each step costs more and convergence is slower, but the resulting suffix is less tied to a single request, which is what you want if the artefact has to be reused.
saying these in an interview costs you the question
- Claims the search backpropagates through a judge model's verdict.
- Thinks the gradient directly produces the new token rather than ranking candidates.
- Believes a longer, more specific target is free — no acknowledgement of slower convergence or non-convergence.
- Reports a success rate without stating the target completion used.