The system you are contracted to attack is a hosted chat product you can only send requests to, but you hold open weights of a similar model plus GPU hours. Under what conditions is running a white-box suffix search on your local copy a defensible plan, and what must you reserve request allowance for?
answer
- surrogate produces candidates, not findings
- ensemble optimisation for transfer
- product is a stack, not a model
- two ledgers: GPU and requests
- report the transfer rate with its denominator
basics
~20 sOnly when you treat the result as a candidate, not a finding: a suffix optimised against local weights is tuned to that copy's tokenizer and parameters. Reserve enough requests against the real product to test transfer, and expect most candidates to fail there, especially through an unseen system prompt and output filter.
solid answer
~50 sIt is defensible when three things hold. First, **the local model is a plausible surrogate** — same tokenizer family and a similar training lineage, not merely a model you happened to have. Second, **you can afford the transfer tests**, because a candidate that is never sent to the real product proves nothing about it. Third, **the deliverable tolerates a candidate pipeline**: you are producing strings to try, and the report counts only those that survived against the deployed system. The usual improvement is to optimise across an **ensemble of local models** rather than one, which trades GPU-hours for a string less overfitted to a single parameter set. The plan stops being defensible when the deployed product differs structurally, not just numerically: a system prompt, retrieval, an input classifier and an output filter are invisible to a local checkpoint. Against such a stack, a black-box loop that adapts to the deployment's own refusals often finds more per unit spend.
go deeper
Should recognise that a local model is not the hosted product and that anything found locally must still be tried against the real endpoint.
Should explain transfer, name ensemble optimisation as the lever, and note the tokenizer must match.
Should split the budget into a GPU ledger and a request ledger, reserve the request one first, and describe the deployment stack — system prompt, input classifier, output filter — as the reason candidates die.
Should set a stopping rule: after a representative sample of candidates fails to transfer, redirect spend to an adaptive black-box loop rather than generating more from the same surrogate.
### What is actually being proposed A suffix optimised against local weights is a string fitted to a very specific interaction between token ids and one parameter set. The premise of the surrogate plan is that some of what it exploits is a property models trained on similar data with similar objectives *share*, so a string that works on one may work on another. That premise is neither obviously true nor obviously false; it is an empirical rate, and your job is to treat it as one. Three levers raise the odds, and each has a price: 1. **Ensemble optimisation.** Optimise one string across several local models simultaneously, so it cannot lean on the quirks of any single parameter set. Costs roughly a forward and backward pass per model per step, so GPU-hours and VRAM scale with the ensemble size. 2. **Multiple target behaviours.** Optimise toward several continuations rather than one, so the string encodes a general effect rather than a single memorised path. Costs more steps to converge, and often converges worse. 3. **Tokenizer alignment.** Ensure the string tokenizes the same way where it lands. Costs nothing, and skipping it silently invalidates everything above. ### What blocks it regardless of how well you optimise The hosted product is **not a model, it is a stack**. A hidden system prompt reframes every input before the model sees it. An input classifier may reject the request before generation happens at all. Retrieval may inject content you never modelled. An output filter may suppress a compliant answer, in which case the *model* was jailbroken and the *product* was not — a distinction the report has to make, because they are different findings with different remediations. None of these components exist in the checkpoint you optimised against, and none can be inferred from it. ### The two ledgers Split the budget explicitly, and reserve in this order: | ledger | buys | unit | |---|---|---| | GPU | **candidates** | hours per suffix x suffixes x ensemble multiplier | | requests | **evidence** | candidates x attempts each, plus retries, plus paraphrase variants when the raw string fails | The request ledger is the binding one, it is rate-limited, and it is the only one that produces something reportable. A plan that spends the whole GPU allowance and leaves no request allowance has produced a folder of untested strings, which is not a result about the product under test. ### Where the number misleads - **A single working string has no denominator.** "We achieved a jailbreak" is unfalsifiable and unusable. "Forty candidates generated, forty tested, three survived" is a transfer rate a reader can act on, and it is also a fact about your *method* that will let you price the next engagement. - **Held-out local transfer is not product transfer.** A string that transfers between two open models in your ensemble tells you about model similarity. It says nothing about a system prompt and an output filter you have never seen, so do not quote the first number as if it predicted the second. - **An instant generic refusal is ambiguous.** It may be an input classifier firing before generation rather than the model declining, and the two demand different next moves — obfuscating past a classifier versus reworking the objective for the model. Latency and response shape usually separate them; the raw failure count does not. - **A zero transfer rate is itself evidence.** After a representative sample of candidates dies, generating more from the same surrogate rarely changes the outcome. That is the point to redirect spend to an adaptive black-box loop that reads the deployment's own refusals, because it optimises against the thing you are attacking rather than a proxy for it. ### What to check Is the local model a plausible surrogate — same tokenizer family, similar lineage — or merely the model you happened to have? Is the request allowance reserved *before* the GPU run starts? Did you record, for the report, which model the search ran against, how many candidates were produced, how many were tested, and how many survived? And is there a written stopping rule, agreed in advance, for when the surrogate has been shown not to transfer?
- Why does optimising one string across several local models usually transfer better?The objective can no longer be satisfied by quirks of a single parameter set, so the surviving string relies on an effect the models share — which is more likely to be present in a fourth, unseen one.
- A candidate works locally, and the hosted product returns a generic refusal instantly. What are you actually seeing?Very possibly an input classifier rejecting before generation, not the model refusing. Distinguishing those changes what you try next — obfuscating for the classifier versus reworking the objective for the model.
- How should the report present candidates that never transferred?As a denominator. Candidates generated, candidates tested, candidates that survived — the transfer rate characterises the method and prevents a reader from over-reading a single success.
Optimising on a local copy is cutting keys against a lock you bought that looks like the one on the client's door. Until you walk over and try them, you have a bag of metal, not an entry; and the report's honest unit is how many of the bag turned, not that one did.
saying these in an interview costs you the question
- Reports a suffix that worked on the local copy as a finding against the hosted product.
- Budgets GPU-hours with no reserved request allowance for transfer testing.
- Assumes a same-named local checkpoint is the served model.
- Keeps generating candidates from one surrogate after transfer testing has repeatedly failed.