Jailbreak Techniques & Taxonomies
You will map the families of jailbreaks that coax a model past its safety training, from persona role-play to multi-turn crescendo and encoding tricks. Interviewers probe this to see if you can reason about why alignment breaks under distribution shift rather than just reciting one viral prompt.
on this pageshowhide
explore
- Anatomy of a Refusal8 questions
- Mapping the Reluctance4 questions
- Reframing the Same Ask4 questions
- Families That Get Through23 questions
- Fiction as Cover4 questions
- Building Consent by Degrees4 questions
- Crowding Out the Training4 questions
- Splitting and Reassembly4 questions
- Under-Covered Framings4 questions
- No Room to Decline3 questions
- Half-Life of a Method8 questions
- Burned on Publication4 questions
- Transfer Across Versions4 questions
questions
page 2 of 2After a persona-card finding, an owner blocks that card's wording and adds a system-prompt refusal line. What does that fix?
basics
~20 sIt removes one wording, already the weakest part of the construction. The framing that carried the behaviour is a shape, not a phrase, and a product whose feature is a free-text persona field lets any user rewrite it.
A form constraint made a CRM drafting assistant emit a confident, caveat-free recommendation with nothing disallowed: is that a finding?
basics
~20 sYes, but not as a content-policy bypass. Nothing disallowed was produced; what was removed are the qualifiers a reader and the downstream workflow use to decide whether to act, which makes it a misinformation and overreliance finding.
In a split-and-reassemble jailbreak, why is the joining request the part that usually breaks?
basics
~10 sBecause the join must be specifiable without naming the result. Name it and you have re-created the single refusable request; leave it vague and the model returns an inert pile, not the artefact.
How do you re-test an inherited library of framings in languages no reviewer reads?
basics
~20 sRe-run them rather than re-read them. Each entry is a dated measurement against coverage that moves, so batch re-tests scored by rate beat review, and any entry whose success nobody can evaluate should be retired rather than carried.
Is a caller-supplied exemplar block that outvotes refusal a bug to file or a product limit?
basics
~20 sNo wording is at fault, so there is nothing to patch. What is on the table is how much untrusted text the product sells the right to send, which is a product owner's call, written down as a stated limit.
Leadership wants your team's only working jailbreak shown verbatim in a public talk — what do you argue?
basics
~20 sArgue to publish the mechanism and keep the wording out. The class-level description earns nearly all the credit and stays true after the next tune, while the quoted string spends the asset to prove a point the description already makes.
After a model upgrade, half your frozen jailbreak regression entries stop firing — what do you report the suite is worth?
basics
~20 sReport it split: entries still tied to a live family, entries that now test nothing, and a claim you refuse to make. A suite of frozen strings decays with every release, and pricing the recut is the actual decision.
Why fund refusal mapping that produces no finding, in a fixed-hours red-team engagement?
basics
~20 sRecon decides where the remaining hours go. Refusal coverage is uneven, so unmapped effort spreads evenly over ground that is mostly covered and returns anecdotes. Cap it, though: the map is a perishable, dated snapshot.
Twenty external submissions name twenty techniques for one refused output - how many findings do you record?
basics
~20 sRecord one class, keyed on the requested output and the behaviour reached, with twenty instances under it. The frames measure how wide the reachable region is - but your counting rule is a payment rule, and reporters optimise against it.
showing 31–39 of 39