You ran a gradient-based search against an open-weights model on your own GPUs to produce an adversarial suffix, then sent that frozen string unchanged to a hosted chat endpoint whose weights you cannot inspect. What does it mean to say the attack "transferred", and why is a failure at the hosted endpoint less informative than a failure on the local model?
answer
- frozen string, no re-optimisation
- no gradients back from a hosted endpoint
- reply text is the only channel
- endpoint is a pipeline, not a bare model
- failure to transfer is not robustness
basics
~20 sTransfer means the string you optimised on the local model still produces the disallowed behaviour on the hosted endpoint, with no re-optimisation. Failure there is uninformative because you get no gradients and no internals back, only the reply text. On the local model a failure still shows you the loss and lets you keep searching.
solid answer
~50 s**Transfer** is the claim that an artefact optimised against one model still works, unchanged, against a different model you did not optimise against. The search itself happened entirely on the local copy: the objective was that model's own next-token loss, and every step needed its weights. Once the string is frozen and fired at a hosted endpoint, the feedback channel collapses to whatever text comes back. A refusal could mean the string never generalised, or that a filter in front of the model rejected it, or that a hidden system prompt changed the framing, or that the endpoint samples differently on this attempt. You cannot separate those from one reply. So a local failure is a *search* result you can act on — adjust the objective, spend more steps. A hosted failure is a single observation about a system you cannot instrument, and it does not license the sentence "this endpoint is not vulnerable".
go deeper
Should say transfer means the same unchanged string works on a model it was not optimised against, and that you only see the reply from a hosted endpoint.
Should list the confounders behind a hosted refusal: different build, hidden system prompt, filter in front or behind the model, sampling variance.
Should insist a transfer attempt is a measurement — repeat count, decoding settings, how a hit was judged — and should refuse to write 'not vulnerable' from one failure.
Should frame it as what the deliverable claims: a transferred result is a dated observation about a deployed system, not a durable property of a model.
## What the word "transfer" is actually claiming **Transfer** is the claim that an artefact — here a fixed string of characters produced by an optimisation run — still produces the behaviour it was tuned for on a *different* model, with no further optimisation. Both halves of that sentence are load-bearing: - **"Unchanged"** matters because if you adjust the string against the new target you are running a fresh search, not observing transfer. - **"A model you did not optimise against"** matters because success on the model you searched on is not a result at all; it is the search terminating. ## Phase one: the search, on a model you hold A **gradient-based search** needs weights because its objective is defined on the model's own output distribution — it scores candidate token substitutions by how far each moves the model toward a target continuation and away from a refusal. Every evaluation is a local forward and backward pass. The bill is **GPU hours**: in practice hours of a large-memory accelerator per target model per run, sometimes most of a day for a stubborn objective, plus the engineer time to define the objective, choose the surrogate and babysit the job. What it does *not* cost is API calls. Nothing leaves your network, no vendor's monitoring sees it, no rate limit applies, and you can run it before the engagement window even opens. ## Phase two: the transfer attempt, at a model you do not hold The string is now frozen and you are an ordinary customer sending prompts. Each attempt is one API call — cents, seconds. The entire feedback channel is the reply text plus whatever metadata the response carries. No loss, no logits, no attention, no gradient of any kind. That **asymmetry** is the trap the whole question turns on: the expensive half of the workflow yields something you can iterate on, and the cheap half yields roughly one bit per attempt. Because it is cheap, people run it once and write down a result. ## Why a failure at the hosted endpoint proves so little A hosted endpoint is a **pipeline**, not a bare model. At minimum it is a model plus a system prompt you never see; commonly there is also an input classifier that can reject a request before generation happens and an output classifier that can suppress a completion after it. At least five distinct causes produce the identical blank refusal in front of you: - the model behind the product name is not the build you guessed; - the hidden system prompt reframed the conversation; - an input guard rejected the unusual surface form of an optimised string; - an output guard suppressed the completion; - or the endpoint simply sampled unluckily this time. One reply cannot separate them, and no amount of staring at it will. ## Where the number misleads Two failure modes dominate reports. - **The first is the missing denominator:** one attempt is not a rate, and hosted endpoints sample, so the same frozen string can succeed and fail on consecutive calls. Writing "it works" or "it doesn't work" from a single call is the most common defect in this workflow. - **The second is promotion** — turning "did not transfer" into "the target is not vulnerable to this attack class". That sentence is a claim about a model; your evidence concerns one frozen artefact, against one deployment, on one day, judged by one criterion. The local failure and the hosted failure look alike on the terminal and are epistemically nothing alike: the local one is a search result with a loss attached that tells you where to go next, and the hosted one is a single uninstrumented observation. ## What you would check before believing either outcome - Fire the string enough times to have a **denominator** worth quoting, and record hits and attempts as raw counts rather than a percentage. - Record the **decoding settings** you sent, and note explicitly which ones you could not control. - Capture any **build or version identifier** the response carries. - Store **reply text verbatim** for successes *and* failures: a uniform canned policy sentence, returned faster than a normal answer with nothing streamed, is the signature of a guard in front of the model, while a varied, in-voice refusal is the model itself — and that distinction is unrecoverable later from a tally. - Send the same underlying request in ordinary language as a **control** so you know whether anything at all about the topic gets through. Then write the finding at the resolution your evidence supports, and no further.
- You fire the frozen string ten times at the hosted endpoint and it works twice. What can you write down?That on this date, against this endpoint, two of ten attempts produced the behaviour, judged by whatever criterion you used. Not a model property, and not a number you may re-quote later without re-testing.
- The hosted endpoint returns an identical canned policy message every time, with no partial completion. What does that hint at?That something in front of the model rejected the request rather than the model refusing it — the uniform wording and absence of any streamed content point at a separate classifier, not at the model's own refusal behaviour.
Optimising on a model you hold is like practising on a lock you can take apart to watch the pins move; firing the frozen string at a hosted endpoint is posting it through a letterbox and learning only whether the door opened. The same silence could mean the wrong key, or a second door behind the first that you never knew was there.
saying these in an interview costs you the question
- Saying the hosted endpoint's refusal proves the model is safe against this attack class.
- Believing you can keep optimising against the hosted endpoint using gradients from its replies.
- Reporting a single attempt as a success rate.
- Treating the hosted endpoint as a bare model with no surrounding system prompt or filters.