skip to content

An adversarial suffix optimised against a locally run open-weights model fires reliably at one vendor's hosted chat endpoint but does nothing at another vendor's hosted endpoint of comparable capability. What mechanisms explain the difference, and how would you work out which one is responsible without any access to either system's internals?

level: middleimportance: should knowfreq 38%

answer

  1. capability rank is not the transfer axis
  2. tokenizer, wrapper, alignment, pipeline
  3. model difference vs. system difference
  4. uniform canned reply = guard in front
  5. benign control sharing the surface form

basics

~20 s

Comparable capability does not mean comparable internals. Different tokenizers re-split your string, different safety tuning gives a different refusal to suppress, hidden system prompts differ, and one endpoint may be a pipeline with classifiers around the model. You separate them from the outside by comparing the shape of the failures: wording, variation, latency, and whether any output streamed at all.

solid answer

~50 s

Four mechanisms, all invisible from outside: - **Tokenization.** Your string was optimised as a token sequence. Under a different vocabulary the same characters become a different sequence and the structure is gone. - **Alignment behaviour.** The search suppressed one model's learned refusal; another model refuses for different learned reasons. - **The hidden wrapper.** Each endpoint imposes its own system prompt and role formatting, and your string was fitted to a wrapper you controlled locally. - **The pipeline.** An endpoint may run classifiers around the model, so the model never sees the request or the completion is suppressed afterwards. **Telling them apart from outside.** Compare failure text across many attempts. One canned sentence, identical every time, returned fast with nothing streamed, is the signature of something in front of the model; a varied, in-voice refusal is the model. If a plainly benign message sharing the string's unusual surface form is also blocked, the filter keys on surface form, not intent.

go deeper

for a junior

Should say the two endpoints are different models with different setups, so the same string need not work on both.

for a middle

Should name tokenizer, hidden system prompt, alignment differences and a surrounding classifier pipeline as distinct causes.

for a senior

Should attribute from the outside using failure-text uniformity, streaming and latency, and benign and paraphrase controls, and should distinguish a model difference from a deployment difference in the report.

for a principal

Should note the reporting consequence: if the difference is a guard, the two vendors' models may be equally susceptible, and the recommendation is about deployment rather than about model choice.

**"Comparable capability" is the wrong axis entirely.** Capability rankings describe what a model can produce when it is cooperating. Transfer depends on how a model *represents text* and how it was taught to decline — two properties that leaderboards do not measure and that vary wildly between vendors sitting next to each other on a benchmark. Reasoning from capability similarity to transfer likelihood is the single most common error in this comparison. **The mechanisms, roughly in the order they bite.** 1. *Tokenization.* The optimised object is a sequence of token ids. Re-tokenized under a different vocabulary it is a different object, and whatever structure the search found is not preserved by the characters. This can fail transfer completely and silently. 2. *Conversation formatting and the hidden system prompt.* You optimised inside a wrapper you controlled. Each vendor imposes its own role markers and injects its own instruction text, and they differ in how much they inject and how strongly it anchors behaviour. 3. *Alignment.* The objective was "do not emit this refusal". Refusals are learned behaviour; a different post-training regime produces a different decision surface for the same request. 4. *The surrounding pipeline.* Operationally the most important distinction: is this a **model** difference or a **system** difference? If vendor B rejects at a classifier before the model, B's underlying model may be exactly as susceptible as A's, and what differs is the deployment. That changes the finding, the recommendation, and who owns the fix. **Attribution from the outside, using only replies.** You have no internals for either system, so you design controls and read shapes. - *Uniformity.* A model's refusal is generated text: wording drifts across attempts and with decoding settings. A pre-model classifier typically returns one fixed message, byte-identical every time. - *Latency and streaming.* If nothing streams and the response arrives faster than a normal answer, generation probably never happened. If tokens start and then stop, suspect an output-side check. - *Benign control.* Send a plainly harmless message that shares the unusual surface characteristics of your string. If that is blocked too, the block keys on surface form, not on the request's intent. - *Paraphrase control.* Send the same underlying request in ordinary language. If it is answered while the optimised string is blocked, the string itself is what is being caught, not the topic. - *Build identifiers.* Record whatever version metadata each response carries, so the attribution is pinned to a specific deployment on a specific date. **What this costs.** Individually the calls are trivial — pennies each. The multiplication is what people forget: two endpoints, four control conditions, and enough repeats per cell to see through sampling variance turns into a few hundred calls before you have anything defensible, and against a metered enterprise endpoint with per-account rate limits that is a scheduling problem as much as a budget one. The larger cost is engineer time: designing controls that isolate one variable each, and resisting the urge to conclude from the first pair of replies. And the expensive branch is the one you should think hardest about — re-optimising against a surrogate closer to vendor B is another block of GPU hours and another day of wall clock, spent on a guess, and it is only worth it if you have already established that B's *model* is what refused. **Where the number misleads.** "Fires at A, not at B" is a denominator swap waiting to happen: it is a statement about two deployments, and it gets written up as a statement about two models. If B blocks at a guard, "B is more robust" is false in the way that matters — B's model may be identical in susceptibility, and any surface at B that bypasses the guard inherits the exposure. The mirror error is also live: a success at A does not mean A has no guard, only that this artefact got past whatever A has. **What to report.** Not "endpoint B is safe". Report that on this date the frozen artefact did not reach the behaviour at endpoint B, give your best-supported attribution with the control evidence behind it, state plainly whether the evidence points at the model or at the deployment, and note that a fresh artefact optimised against a closer surrogate remains untested.

  • Why does it matter for the report whether the block came from the model or from a classifier in front of it?
    Because they imply different remediations and different generality. A guard means the model itself may be equally susceptible and the defence is a deployment property that other surfaces may lack; a model refusal is a property that travels with the model.
  • What does it tell you if a plainly benign message sharing your string's unusual surface form is also blocked?
    That the block keys on the surface form rather than on the request's intent — evidence for an input classifier reacting to anomalous text, not for the model judging the underlying ask.

A refusal generated by the model is the host declining to answer you; a rejection by a classifier in front of it is a bouncer turning you away at the door. From the street both look like the same 'no', but only one of them tells you anything about the host.

saying these in an interview costs you the question

  • Concluding the second vendor's model is inherently more robust with no attribution evidence.
  • Ignoring that the endpoint may be a pipeline and treating every refusal as the model's own.
  • Ranking transfer likelihood by capability benchmarks.
  • Running no benign or paraphrase control before asserting what caused the block.

context