skip to content

Before spending GPU hours optimising a jailbreak string against a locally run open-weights model in the hope it fires at a hosted endpoint you cannot inspect, how do you choose which local model to optimise against, and why do practitioners often optimise against several local models at once?

level: middleimportance: must knowfreq 48%

answer

  1. surrogate = guess at a model you cannot see
  2. tokenizer, chat wrapper, refusal style
  3. ensemble kills single-model quirks
  4. GPU cost scales with ensemble size
  5. held-out local model as a cheap pre-test

basics

~20 s

Pick a local model as close as you can guess to the hidden one: similar tokenizer, similar chat formatting, similar safety tuning. The string is optimised for those exact tokens and that exact refusal behaviour. Optimising against several local models at once forces the string onto behaviour they share instead of one model's quirks, which usually survives the jump better.

solid answer

~50 s

You are choosing a **stand-in for a model you cannot see**, so the choice is a guess informed by behaviour, not by weights. What matters: the **tokenizer**, because a string optimised as a token sequence can retokenize into something else under a different vocabulary; the **chat formatting and role markers**, because the search bakes in whatever wrapper you trained against; and the **kind of safety tuning**, because the objective is defined against a refusal you are trying to suppress. A model with a very different refusal style is a poor proxy. Optimising against an **ensemble** — averaging the objective across several local models — is the standard mitigation. A string that satisfies all of them cannot lean on one model's idiosyncrasies. The costs are real: GPU time scales with ensemble size, strings come out longer and more conspicuous, convergence is slower. The cheap discipline before spending query budget: hold out a local model you did **not** optimise against and test there first.

go deeper

for a junior

Should say you optimise on a local model that stands in for the hidden one, and that a closer stand-in transfers better.

for a middle

Should name the concrete mismatches — tokenizer, chat formatting, refusal behaviour — and explain the ensemble as a defence against overfitting to one model.

for a senior

Should add held-out validation before spending query budget, and should price the ensemble against GPU hours and against how conspicuous the resulting string becomes.

for a principal

Should treat surrogate choice as a bet whose payoff you can track across engagements, and set a policy on how much local compute is worth committing before any hosted evidence exists.

**The one decision you make completely blind.** Every other step in this workflow produces a number you can look at. Surrogate choice does not: you commit GPU hours to a stand-in for a model whose weights, tokenizer, system prompt and safety training you will never see, and you commit them *before* any evidence exists. The choice dominates the outcome, so it deserves more thought than "pick the biggest open-weights model that fits". **Fingerprinting instead of inspecting.** Since you cannot compare internals, you compare behaviour from the outside: how the hidden endpoint formats answers, how it words a refusal, how it handles long or malformed context, which topics it declines and how it declines them. That gives you a *family guess* — a plausible lineage — and the family guess is what you optimise against. It is a guess, and it should be written down as one, because when transfer fails the first hypothesis to test is that the guess was wrong. **The three mismatches that destroy transfer.** 1. *Tokenization.* The search operates over token ids, and its output is a token sequence, not really a string of characters. Under a different vocabulary the same characters split differently and the structure the search found is simply gone. This is the most brutal mismatch and the least visible, which is why a shared tokenizer lineage is worth more than a similar capability score. 2. *The conversation wrapper.* You optimised inside role markers and a system prompt you chose. The hosted endpoint imposes its own wrapper, including instruction text you never see, and vendors differ enormously in how much they inject and how strongly it anchors behaviour. 3. *Alignment behaviour.* The objective is the suppression of a refusal. Refusals are learned, so two equally capable models can have completely different decision surfaces for the same request; a string fitted to one refusal geometry has no particular reason to move the other. **Ensembles: what they buy and exactly what they cost.** Averaging the objective across several local models penalises any solution that leans on one model's idiosyncrasy, so what survives is closer to behaviour the members share — which is the closest proxy available for behaviour a *fourth*, unseen model might also share. The costs are concrete and rarely stated. GPU hours scale roughly linearly with the number of members, because each member must be evaluated on every step, so an overnight single-model run becomes a weekend. All members must be resident, so accelerator memory caps the ensemble size before the schedule does. Convergence is slower and the resulting string tends to be longer and stranger — which makes it *more* conspicuous to exactly the input classifier that will decide whether it ever reaches a hosted model. There is a genuine tension here: the thing that makes a string generalise is also the thing that makes it easy to spot. **Held-out validation, and where its number misleads.** The cheap discipline is to exclude one local model from the optimisation and test there first. It costs local inference only — minutes, no metered queries, no vendor traffic — and it separates "this string generalises" from "this string fits the surrogate". But be careful how you read the number it gives you. The held-out model is *bare*: no hidden system prompt, no input classifier, no output classifier, and it is probably drawn from the same open-weights population as the ensemble. Its success rate is therefore systematically optimistic as a predictor of hosted success. Use it as a filter — a string that fails here almost never justifies spending metered queries — not as an estimator of what will happen at a deployment. **What you would check before committing the compute.** Confirm the surrogate's tokenizer family and record it. Confirm that the chat template you are optimising inside is at least plausible for the target, rather than whatever your local serving stack defaulted to. Log GPU hours per usable string, so the "cheap because we own the hardware" claim can be tested against opportunity cost. And when you report, keep the held-out number and any hosted number in separate columns with raw counts, so a reader can see how much of the claim is generalisation and how much is fit to a guess.

  • Why is a shared tokenizer lineage worth more than a similar capability score when picking the local model to optimise against?
    Because the search produces a token sequence. Under a different vocabulary the same characters split differently and the optimised structure is destroyed, whereas capability similarity does not preserve anything about the token-level solution.
  • What is the cheapest useful pre-test before firing a locally optimised string at a metered hosted endpoint?
    Run it against a local model you deliberately excluded from the optimisation. It costs only local inference and it separates 'this string generalises' from 'this string fits the surrogate'.

saying these in an interview costs you the question

  • Choosing the surrogate purely by benchmark capability, ignoring tokenizer and refusal behaviour.
  • Claiming the surrogate 'is basically the same model' as a closed endpoint with no basis for the claim.
  • Optimising on one model and going straight to a metered hosted endpoint with no held-out test.
  • Assuming an ensemble is free — not accounting for the multiplied GPU cost or the longer, more conspicuous string.

context