skip to content

Jailbreak Techniques & Taxonomies

You will map the families of jailbreaks that coax a model past its safety training, from persona role-play to multi-turn crescendo and encoding tricks. Interviewers probe this to see if you can reason about why alignment breaks under distribution shift rather than just reciting one viral prompt.

on this pageshow

explore

questions

page 1 of 2

Why does an attacker's block of hundreds of compliant example exchanges weaken a model's refusal?

level: juniorimportance: must knowfreq 65%

basics

~20 s

A model's refusal is a trained propensity, not an enforced rule. Hundreds of consistent in-context examples showing the model complying supply competing evidence about how this exchange goes, and enough of them outweigh that propensity.

open as a page

In a character-chat product, why can a user-authored persona card get content the model refuses when asked plainly?

level: juniorimportance: must knowfreq 74%

basics

~20 s

A refusal is a learned response to how a request looks, not an enforced rule. A persona card restates the request as a character's line in a scene, a shape the model was mostly trained to continue rather than decline.

open as a page

An attacker mandates a drafting assistant's opening line and bans hedging: why does that change whether it refuses?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A refusal is text the model generates token by token, and it usually begins a characteristic way. Constraining the answer's first words and forbidding qualifiers removes those openings, so the trained tendency to decline has much less to latch onto.

open as a page

Why does a refused request get answered when restated in a low-resource language?

level: juniorimportance: must knowfreq 66%

basics

~20 s

Capability generalises further than safety training does. The model still follows and answers the low-resource form competently, but the alignment data barely covered it, so the refusal behaviour that fires in a widely spoken language never engages.

open as a page

Why is a jailbreak prompt quoted in a public write-up worth less than the reason it worked?

level: juniorimportance: must knowfreq 58%

basics

~20 s

A quoted prompt is a fixed string, and publishing it hands the vendor the exact text used to train that behaviour out. The reason it worked has many wordings and survives; the wording is a wasting asset.

open as a page

Why does a jailbreak prompt that works on one vendor's model usually fail on another's?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A jailbreak prompt bundles two things: the family, meaning the general mechanism it leans on, and the exact wording, refined against one model. The mechanism often crosses to another vendor's model; the fitted wording usually does not.

open as a page

An attacker probes a chat model with cheap request variants — why is what they find not a banned-topic list?

level: juniorimportance: must knowfreq 72%

basics

~20 s

There is no list to find. Refusal is a learned propensity with uneven coverage, so near-identical variants of one request can be treated differently. Probing maps a fuzzy, uneven boundary, not an enumerable set of forbidden subjects.

open as a page

When an attacker reframes a request a model refused, what stays constant and what changes?

level: juniorimportance: must knowfreq 66%

basics

~20 s

The request stays constant - the same output is still being asked for. Only the frame changes: the setting, the stated purpose, or the product slot the ask is placed in. Reframing moves the ask, it does not soften it.

open as a page

Why doesn't the answering model refuse the final step that joins parts it already produced?

level: middleimportance: must knowfreq 58%

basics

~20 s

Because the join is not the objectionable request. The model sees ordinary composition over material already in front of it, and refusal is trained on recognising an ask, not on auditing what a finished output adds up to.

open as a page

What does a jailbreak working on one vendor's model establish about other vendors' models?

level: middleimportance: must knowfreq 66%

basics

~20 s

Very little on its own. One success is a single sample from one model version in one deployment. It makes the family worth testing elsewhere, but the fitted wording is evidence about that one tuning run and nothing more.

open as a page

Why does an input screen that scores each request on its own miss a request delivered in parts?

level: juniorimportance: should knowfreq 62%

basics

~20 s

The screen only ever sees one part, and each part is ordinary work with nothing to refuse. The objectionable whole is never stated in any request; the model produces it when it combines the parts.

open as a page

Why is a larger context window an attack budget, not only a capability gain?

level: middleimportance: should knowfreq 52%

basics

~20 s

Context length is a resource both sides spend. The window that lets a caller paste a codebase lets an attacker paste enough consistent evidence to outvote a trained refusal, so raising the limit raises that family's ceiling.

open as a page

In a character-chat product, why is the persona frame the jailbreak family safety training covers most densely?

level: middleimportance: should knowfreq 48%

basics

~20 s

Because it is the most published family and the easiest to collect. Shared character cards circulate in near-identical wording, so refusal tuning and screening data are both dense with exactly those phrasings, and persona products tune persona behaviour deliberately.

open as a page

In refusal suppression, why does fixing an answer's first words bite harder than banning caveats later?

level: middleimportance: should knowfreq 48%

basics

~20 s

Refusal behaviour concentrates where an answer begins, so a constraint on the first words competes with it directly. A ban on caveats applies to text already being written, so it shapes the payoff rather than whether the model declines.

open as a page

Why is restating a request in an exotic form alone not enough to get an answer the assistant previously refused?

level: middleimportance: should knowfreq 48%

basics

~20 s

An under-covered framing needs two things at once: thin safety coverage of that form, and a model still competent in it. Push too far out and competence fails before caution does, so what returns is noise.

open as a page

Why doesn't a jailbreak string stop working the moment its public write-up appears?

level: middleimportance: should knowfreq 44%

basics

~20 s

Publication starts a clock, not a switch. A deployed screening layer can match the published wording within a release; the model's trained refusal changes only at the next tune. That gap is the string's remaining life.

open as a page

A chat model declines one variant of an attacker's probe and answers another: what does either single observation prove?

level: middleimportance: should knowfreq 55%

basics

~20 s

Very little alone. Replies are sampled, so one decline shows only that this draw refused and one answer only that this draw did not. The boundary is a rate, and estimating a rate costs repeated probes.

open as a page

Why does a model refuse a request in one framing but answer the same request in another?

level: middleimportance: should knowfreq 57%

basics

~20 s

Refusal is a learned propensity conditioned on the whole context, not a check on the request. Safety training covered some contexts heavily and others thinly, so an identical ask placed in a thinly covered context meets weaker reluctance.

open as a page

How does a prompt assembler's truncation order cap a caller's many-example jailbreak attempt?

level: seniorimportance: should knowfreq 38%

basics

~20 s

The attempt is only as strong as the number of demonstrations that actually reach the model. A prompt assembler fits caller text, instructions and items into one budget, so its truncation order, not the wording, sets the ceiling.

open as a page

In a character-chat product, a persona scene returns fluent in-character content on a refused topic. What can the tester claim?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Only that the model kept generating past the point it usually declines. In-character text is written to fit the story, so where the model has nothing it invents something plausible, and elicitation and confabulation look identical from inside the scene.

open as a page

In a request log where every line passed the input screen, what has that log actually ruled out?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Only that no single logged request scored above the screen's threshold. It rules out nothing about what the model produced, and a request split across ordinary parts is invisible in an input log by construction.

open as a page

A rarely covered language framing returns refused content, but degraded - is that a finding?

level: seniorimportance: should knowfreq 38%

basics

~10 s

File it, but claim only what it shows: the trained refusal did not engage for that form. A degraded answer is evidence about model behaviour, not proof that usable content was obtained.

open as a page

A colleague files a jailbreak finding whose entire content is one quoted prompt — how do you judge its value?

level: seniorimportance: should knowfreq 38%

basics

~10 s

Treat it as an instance until somebody names the class. Re-run it, ask which variants still worked, and value it at its remaining shelf life unless a property behind it can be stated.

open as a page

A stored jailbreak regression entry stopped firing after your gateway repointed a route to a newer model — what can you conclude?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Almost nothing attributable. A frozen entry welds a mechanism to one fitted wording, so a negative is consistent with the wording no longer landing, the mechanism being tuned out, the deployment changing, or simple sampling variation.

open as a page

Mapping a chat model's refusal coverage from one free-tier account: how do you spend a finite message quota?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Coarse first, then deep. Sweep neighbouring topics with one probe each to find where outcomes vary, then concentrate repeats on that band. Isolate probes in fresh conversations, move one dimension at a time, and carry a control probe throughout.

open as a page

A vendor patched the framing your report named, and a new frame returns the same output - what was fixed?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The surface was fixed, not the behaviour. A denylist of named patterns fitted to previously published framings covers the reported wording and its near neighbours; the requested output and the thin region it was reached through are untouched.

open as a page

showing 1–30 of 39