skip to content

Why does a jailbreak prompt that works on one vendor's model usually fail on another's?

level: juniorimportance: must knowfreq 72%

answer

  1. one prompt carries two separable layers
  2. mechanism versus wording
  3. the wording was fitted to one model
  4. families ride shared pretraining data

basics

~20 s

A jailbreak prompt bundles two things: the family, meaning the general mechanism it leans on, and the exact wording, refined against one model. The mechanism often crosses to another vendor's model; the fitted wording usually does not.

solid answer

~50 s

Any jailbreak you can point at is really two layers welded together. The **family** is the mechanism it exploits, a general property of instruction-following models, such as escalating across many turns so each step reads as a small continuation of the last, or filling the context with prior compliance so compliance becomes the likely continuation. The **instantiation** is the wording that was refined by trial and error until it landed on one particular model. Refusal is a trained preference over token sequences, not a rule, and its boundary is shaped by one vendor's post-training run; a different vendor tokenizes differently, formats turns differently and tuned on different safety data, so the fitted wording lands somewhere else. The mechanism travels because the models are pretrained on heavily overlapping data and aligned with similar recipes. So a failed replay of somebody's exact prompt tells you the string did not land, not that the family is absent.

go deeper

for a junior

Be ready to state that a jailbreak has a mechanism and a wording, and that the wording was tuned against one particular model. Say plainly that a failed copy-paste proves nothing about the underlying weakness.

for a middle

Explain why the wording is fitted: different tokenizers, different turn formats, and a refusal boundary produced by one vendor's own tuning run. Then explain why the mechanism travels anyway.

for a senior

Show the habit of re-fitting rather than replaying when you move to a new model, and of running repeats before you claim a result. Interviewers want the reflex, not the vocabulary.

for a principal

Own the consequence for what a programme invests in: durable value sits in the catalogue of families, not in a library of strings that the next release invalidates.

## What a jailbreak is aimed at A jailbreak is a construction aimed at a model's *trained refusal* — the preference, installed by post-training, to decline certain requests. It is not the same thing as a prompt injection, which aims at the *application's* own instructions using data the application feeds to the model. That distinction matters here, because a refusal lives in the weights. It changes when the weights change: on a new tune, on a point release, and above all on a different vendor's model. ## Two separable layers in one prompt - **The family** is the mechanism — the general property of an instruction-following model that the construction leans on. Examples of the kind of property meant: a long history of compliance in the context makes further compliance the likely continuation; a plausible reason that the task is not what it appears to be shifts how the request is read; splitting a request across turns means no single turn looks like the thing being asked for. - **The instantiation** is the wording — the particular turns, phrasing, ordering and formatting that somebody refined, usually by trial and error, until they landed on one model on one day. A person who has only ever copied a working prompt cannot tell these apart, and that is exactly what an interviewer is probing for. ## Why the wording is fitted, and to what Refusal is a learned preference over token sequences, not an if-statement. The line between answered and declined is a surface shaped by one vendor's post-training data and one preference-optimisation run. The wording that works is the wording that happened to land on the permissive side of *that* surface. Several things that differ between vendors move it, and none of them is the mechanism: | Differs between vendors | Effect on a copied prompt | | --- | --- | | Tokenizer | The same characters segment into different tokens | | Chat format and role markers | Turn structure copied from one vendor reads oddly on another | | Safety tuning data and objectives | The refusal boundary sits in a different place | | Deployment defaults around the model | A different system prompt and sampling settings wrap the request | ## Why the family crosses anyway Models from different vendors are more alike than their positioning suggests. They are pretrained on heavily overlapping corpora — the public web, code, books — so they have absorbed the same conversational patterns. Their architectures are similar. They are aligned with similar recipes against similar objectives. A mechanism that exploits something all instruction-following models do is therefore a hypothesis worth testing anywhere. The string is a fact about one tune. ## Reading results correctly The direction of the claim is where candidates slip. - A **failed replay** of an exact prompt on a new model shows that this instantiation did not land. It does not show the family is absent, and it does not show anything was fixed. - A **successful replay** shows the model produced the output once, on that version, in that deployment. Decoding is stochastic, so a construction that fires some of the time gives a different verdict on different runs; one replay is one sample. - Someone who says "we tried the known jailbreaks and they all failed" has tested strings, not families. ## What practitioners actually do when crossing They expect to re-fit. The mechanism is kept; the wording is rebuilt for the target's formatting conventions and its refusal behaviour, and repeats are run before anything is claimed. That is why the *family* is the durable asset and the string is the perishable one: the family survives a retune and often survives a move to another vendor, while the wording is fitted to a single boundary that the next release moves.

  • Is a jailbreak the same thing as a prompt injection?
    No. A jailbreak aims at the model's trained refusal — what it was tuned to decline. An injection aims at the application's own instructions, using data the application itself feeds into the model, such as a retrieved chunk or a fetched page. They target different things, and a defence or a tune that moves one need not move the other.
  • If the exact prompt fails on the second model, how would you tell whether the family still works there?
    Re-instantiate the family: keep the mechanism, rebuild the wording for that model's turn format and refusal behaviour, and run repeats. A single replay of a frozen string is one sample of one instantiation, so it cannot separate 'the wording missed' from 'the mechanism is gone'.
  • If the mechanism is what works, why does the wording matter at all?
    Because the refusal is a trained preference over token sequences, and the wording decides which side of that boundary a given attempt lands on. The mechanism explains why the attempt is plausible at all; the wording is what makes a particular attempt land on a particular tune, and small changes move it.

Picking a lock is a technique; a key cut for one lock is a fitted object. The technique travels to the next door, the cut key does not.

saying these in an interview costs you the question

  • Claims one working prompt proves every model is vulnerable
  • Reads a failed replay as proof the weakness is gone
  • Treats jailbreaking and prompt injection as the same attack
  • Assumes copied wording transfers because the models are similar
  • Cannot separate the mechanism from the wording that instantiates it

context