skip to content

Transferring an Attack

An attack optimised on a model you hold may or may not fire at one you cannot see, and whatever transfers decays with the next vendor update. Interviewers ask what a transferred success rate is worth.

on this pageshow

explore

questions

5

You ran a gradient-based search against an open-weights model on your own GPUs to produce an adversarial suffix, then sent that frozen string unchanged to a hosted chat endpoint whose weights you cannot inspect. What does it mean to say the attack "transferred", and why is a failure at the hosted endpoint less informative than a failure on the local model?

level: juniorimportance: must knowfreq 55%

answer

  1. frozen string, no re-optimisation
  2. no gradients back from a hosted endpoint
  3. reply text is the only channel
  4. endpoint is a pipeline, not a bare model
  5. failure to transfer is not robustness

basics

~20 s

Transfer means the string you optimised on the local model still produces the disallowed behaviour on the hosted endpoint, with no re-optimisation. Failure there is uninformative because you get no gradients and no internals back, only the reply text. On the local model a failure still shows you the loss and lets you keep searching.

solid answer

~50 s

**Transfer** is the claim that an artefact optimised against one model still works, unchanged, against a different model you did not optimise against. The search itself happened entirely on the local copy: the objective was that model's own next-token loss, and every step needed its weights. Once the string is frozen and fired at a hosted endpoint, the feedback channel collapses to whatever text comes back. A refusal could mean the string never generalised, or that a filter in front of the model rejected it, or that a hidden system prompt changed the framing, or that the endpoint samples differently on this attempt. You cannot separate those from one reply. So a local failure is a *search* result you can act on — adjust the objective, spend more steps. A hosted failure is a single observation about a system you cannot instrument, and it does not license the sentence "this endpoint is not vulnerable".

go deeper

for a junior

Should say transfer means the same unchanged string works on a model it was not optimised against, and that you only see the reply from a hosted endpoint.

for a middle

Should list the confounders behind a hosted refusal: different build, hidden system prompt, filter in front or behind the model, sampling variance.

for a senior

Should insist a transfer attempt is a measurement — repeat count, decoding settings, how a hit was judged — and should refuse to write 'not vulnerable' from one failure.

for a principal

Should frame it as what the deliverable claims: a transferred result is a dated observation about a deployed system, not a durable property of a model.

## What the word "transfer" is actually claiming **Transfer** is the claim that an artefact — here a fixed string of characters produced by an optimisation run — still produces the behaviour it was tuned for on a *different* model, with no further optimisation. Both halves of that sentence are load-bearing: - **"Unchanged"** matters because if you adjust the string against the new target you are running a fresh search, not observing transfer. - **"A model you did not optimise against"** matters because success on the model you searched on is not a result at all; it is the search terminating. ## Phase one: the search, on a model you hold A **gradient-based search** needs weights because its objective is defined on the model's own output distribution — it scores candidate token substitutions by how far each moves the model toward a target continuation and away from a refusal. Every evaluation is a local forward and backward pass. The bill is **GPU hours**: in practice hours of a large-memory accelerator per target model per run, sometimes most of a day for a stubborn objective, plus the engineer time to define the objective, choose the surrogate and babysit the job. What it does *not* cost is API calls. Nothing leaves your network, no vendor's monitoring sees it, no rate limit applies, and you can run it before the engagement window even opens. ## Phase two: the transfer attempt, at a model you do not hold The string is now frozen and you are an ordinary customer sending prompts. Each attempt is one API call — cents, seconds. The entire feedback channel is the reply text plus whatever metadata the response carries. No loss, no logits, no attention, no gradient of any kind. That **asymmetry** is the trap the whole question turns on: the expensive half of the workflow yields something you can iterate on, and the cheap half yields roughly one bit per attempt. Because it is cheap, people run it once and write down a result. ## Why a failure at the hosted endpoint proves so little A hosted endpoint is a **pipeline**, not a bare model. At minimum it is a model plus a system prompt you never see; commonly there is also an input classifier that can reject a request before generation happens and an output classifier that can suppress a completion after it. At least five distinct causes produce the identical blank refusal in front of you: - the model behind the product name is not the build you guessed; - the hidden system prompt reframed the conversation; - an input guard rejected the unusual surface form of an optimised string; - an output guard suppressed the completion; - or the endpoint simply sampled unluckily this time. One reply cannot separate them, and no amount of staring at it will. ## Where the number misleads Two failure modes dominate reports. - **The first is the missing denominator:** one attempt is not a rate, and hosted endpoints sample, so the same frozen string can succeed and fail on consecutive calls. Writing "it works" or "it doesn't work" from a single call is the most common defect in this workflow. - **The second is promotion** — turning "did not transfer" into "the target is not vulnerable to this attack class". That sentence is a claim about a model; your evidence concerns one frozen artefact, against one deployment, on one day, judged by one criterion. The local failure and the hosted failure look alike on the terminal and are epistemically nothing alike: the local one is a search result with a loss attached that tells you where to go next, and the hosted one is a single uninstrumented observation. ## What you would check before believing either outcome - Fire the string enough times to have a **denominator** worth quoting, and record hits and attempts as raw counts rather than a percentage. - Record the **decoding settings** you sent, and note explicitly which ones you could not control. - Capture any **build or version identifier** the response carries. - Store **reply text verbatim** for successes *and* failures: a uniform canned policy sentence, returned faster than a normal answer with nothing streamed, is the signature of a guard in front of the model, while a varied, in-voice refusal is the model itself — and that distinction is unrecoverable later from a tally. - Send the same underlying request in ordinary language as a **control** so you know whether anything at all about the topic gets through. Then write the finding at the resolution your evidence supports, and no further.

  • You fire the frozen string ten times at the hosted endpoint and it works twice. What can you write down?
    That on this date, against this endpoint, two of ten attempts produced the behaviour, judged by whatever criterion you used. Not a model property, and not a number you may re-quote later without re-testing.
  • The hosted endpoint returns an identical canned policy message every time, with no partial completion. What does that hint at?
    That something in front of the model rejected the request rather than the model refusing it — the uniform wording and absence of any streamed content point at a separate classifier, not at the model's own refusal behaviour.

Optimising on a model you hold is like practising on a lock you can take apart to watch the pins move; firing the frozen string at a hosted endpoint is posting it through a letterbox and learning only whether the door opened. The same silence could mean the wrong key, or a second door behind the first that you never knew was there.

saying these in an interview costs you the question

  • Saying the hosted endpoint's refusal proves the model is safe against this attack class.
  • Believing you can keep optimising against the hosted endpoint using gradients from its replies.
  • Reporting a single attempt as a success rate.
  • Treating the hosted endpoint as a bare model with no surrounding system prompt or filters.

context

open as a page

Before spending GPU hours optimising a jailbreak string against a locally run open-weights model in the hope it fires at a hosted endpoint you cannot inspect, how do you choose which local model to optimise against, and why do practitioners often optimise against several local models at once?

level: middleimportance: must knowfreq 48%

basics

~20 s

Pick a local model as close as you can guess to the hidden one: similar tokenizer, similar chat formatting, similar safety tuning. The string is optimised for those exact tokens and that exact refusal behaviour. Optimising against several local models at once forces the string onto behaviour they share instead of one model's quirks, which usually survives the jump better.

open as a page

A quarter ago your report recorded that a gradient-optimised suffix produced the disallowed behaviour on roughly two of five attempts against a hosted chat endpoint. Re-running the same frozen string today, it almost never works. What are the plausible causes, and what should the original entry have recorded so this is diagnosable at all?

level: seniorimportance: must knowfreq 42%

basics

~20 s

The endpoint changed under you: a new model build, an altered system prompt, or a filter added in front or behind it. Your string did not decay. The entry should have stamped the date, the endpoint and any build identifier returned, decoding settings, the number of attempts, and how a hit was judged. Without those, nothing is attributable.

open as a page

An adversarial suffix optimised against a locally run open-weights model fires reliably at one vendor's hosted chat endpoint but does nothing at another vendor's hosted endpoint of comparable capability. What mechanisms explain the difference, and how would you work out which one is responsible without any access to either system's internals?

level: middleimportance: should knowfreq 38%

basics

~20 s

Comparable capability does not mean comparable internals. Different tokenizers re-split your string, different safety tuning gives a different refusal to suppress, hidden system prompts differ, and one endpoint may be a pipeline with classifiers around the model. You separate them from the outside by comparing the shape of the failures: wording, variation, latency, and whether any output streamed at all.

open as a page

You lead a red-team program with in-house GPUs for optimising attack strings against local models, and a set of hosted endpoints you can only send prompts to. Transferred strings tend to stop working after vendor changes. How do you decide how much of a program's effort goes into surrogate-optimised transfer attacks versus attacks discovered by querying the hosted endpoints directly, and what evidence would make you cut transfer work entirely?

level: principalimportance: should knowfreq 28%

basics

~20 s

Decide from your own history, not from the literature: track what fraction of locally optimised strings ever fired at a hosted endpoint and how long they kept working. If that half-life is shorter than your reporting cycle, transfer buys exhibits that are stale on delivery. Keep it where it produces cheap candidates; cut it when the endpoints are guard-dominated.

open as a page