What does a jailbreak working on one vendor's model establish about other vendors' models?
answer
- one success, one tune, one sample
- separate the family from the fitting
- shared pretraining makes families plausible
- vary the wording, then run repeats
- scope the claim to route and date
basics
~20 sVery little on its own. One success is a single sample from one model version in one deployment. It makes the family worth testing elsewhere, but the fitted wording is evidence about that one tuning run and nothing more.
solid answer
~50 sIt establishes that the construction produced the output once, on that version, in that deployment — no more. To turn that into a claim about other models you have to separate the two layers. The **family**, the mechanism being exploited, is plausibly shared: vendors pretrain on heavily overlapping corpora and align with similar recipes, so a mechanism that leans on a general property of instruction-following is worth testing anywhere. The **fitting** — the wording, turn structure and formatting refined until it landed — is evidence about one tokenizer and one preference-tuning run, and it is the part the next point release invalidates. A single trial also cannot separate a weak real effect from sampling noise, because decoding is stochastic. The defensible report is scoped: this construction fired *n* of *m* trials, on this route, on this date, and here is the family it instantiates.
go deeper
Know that one success is one observation on one model, and that the honest summary names which model and how many attempts were made. Avoid sweeping statements about models in general.
Explain why the mechanism is plausibly shared across vendors while the wording is not, and describe varying the wording and repeating trials as the way to tell which layer produced the result.
Show that you scope a finding to route, version and trial count, and that you would ablate the fitting before writing anything about a family. Be ready to say what your result does not support.
Own how findings are worded across the organisation: a report that names a family and gives a measured rate stays useful after an upgrade, while one that names a string expires with it.
## The wrong answer this question is built to catch The answer an interviewer is listening for you *not* to give is "it worked on one model, so the class is vulnerable." It sounds like a reasonable generalisation and it is the standard mistake, because it collapses two different things — the mechanism and its fitting — into one result and then generalises the whole bundle. ## What one success actually licenses A success is an observation with a very specific scope: on this model version, behind this deployment, with these sampling settings, this construction produced the output on this attempt. Four things limit how far it travels. 1. **It is one sample.** Generation is stochastic. A construction that fires roughly one attempt in five will look like a solid finding or like nothing at all depending on the draw. Without repeats you cannot say which you have. 2. **It is one tune.** The refusal you got past is a preference installed by one vendor's post-training run. The next point release re-runs that process, and the boundary moves. 3. **It is one deployment.** The model behind a route usually sits inside default instructions, a turn format, and sampling settings chosen by whoever operates it. Any of those can be why the attempt landed. 4. **It bundles the mechanism with the wording.** The result cannot tell you which layer did the work unless you varied one of them. ## What genuinely carries between vendors The reason cross-vendor transfer is a real phenomenon rather than a coincidence is that the models are alike where it counts. Pretraining corpora overlap heavily — the same public text, code and books. Architectures are similar. Alignment recipes are similar, and they optimise similar objectives against similar categories of disallowed request. So a mechanism that exploits a general property of instruction-following models — that context is persuasive, that a long run of compliance biases the next continuation, that a plausible reframing of the task changes how a request is read — is exploiting something every one of them has. What does not carry is everything fitted: tokenization, turn markers, the exact ordering that happened to land, and the specific position of one vendor's refusal boundary. ## Turning one result into a claim you can defend The move a strong candidate describes is *ablation of the fitting*. Keep the mechanism, rebuild the wording several different ways, and run repeats. If several independent instantiations of the same family fire above baseline, you have evidence about the family. If only the original wording fires, you have evidence about a string, and you should say so — that finding has a short half-life and will not survive a retune. The same discipline applies to the report. "This class of model is vulnerable to X" is a claim about many systems from one measurement. "This family fired on 12 of 20 trials against route A and 4 of 20 against route B, using three different instantiations" is a claim you can stand behind, and it is also the claim that stays useful after the next upgrade, because it names the family and gives a baseline to re-measure against. ## The direction of every inference here - A model producing the content proves the answer was generated, not that the vendor's tuning is generally ineffective. - A construction working on two vendors' models is real evidence about the family, and still says nothing about a third until you test it. - A construction failing elsewhere is evidence about that wording, not about the mechanism. - A high success rate on one route proves that route's configuration is permissive to this construction, not that every deployment of the same model is. Holding those directions straight is most of what separates a candidate who has run these tests from one who has read about them.
- What would make you willing to call the family, rather than the instance, confirmed?Several instantiations of the same mechanism, worded and structured differently, firing above baseline across repeated trials, and ideally on more than one vendor's model. That rules out the possibility that one particular string happened to land on one boundary, which is the only thing a single fitted prompt can ever show.
- Why do repeated trials matter even when you only care about a single model?Because decoding is stochastic and many constructions fire probabilistically. A construction that succeeds one time in five yields a positive or a negative depending on the draw, so one trial cannot distinguish a weak but real effect from noise, and it cannot support a before-and-after comparison across a version change.
- Which part of a one-model result should you expect to be dead after the vendor's next point release?The fitting. The wording was refined against a refusal boundary that a new tuning run moves, so the exact prompt is the first casualty. The family usually survives, which is why the durable part of a finding is the description of the mechanism rather than the string that demonstrated it.
saying these in an interview costs you the question
- Reports the whole class as vulnerable from one success
- Runs a single trial and calls it a result
- Treats the fitted wording as the finding itself
- Assumes similar vendors share the same refusal boundary
- Ignores that the deployment around the model may explain the result