skip to content

Families That Get Through

Framings that move a request into thinly trained territory: a role, an escalating transcript, in-context volume, decomposition, a forced opening. Interviewers want the mechanism, never the prompt.

on this pageshow

explore

questions

23

Why does an attacker's block of hundreds of compliant example exchanges weaken a model's refusal?

level: juniorimportance: must knowfreq 65%

basics

~20 s

A model's refusal is a trained propensity, not an enforced rule. Hundreds of consistent in-context examples showing the model complying supply competing evidence about how this exchange goes, and enough of them outweigh that propensity.

open as a page

In a character-chat product, why can a user-authored persona card get content the model refuses when asked plainly?

level: juniorimportance: must knowfreq 74%

basics

~20 s

A refusal is a learned response to how a request looks, not an enforced rule. A persona card restates the request as a character's line in a scene, a shape the model was mostly trained to continue rather than decline.

open as a page

An attacker mandates a drafting assistant's opening line and bans hedging: why does that change whether it refuses?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A refusal is text the model generates token by token, and it usually begins a characteristic way. Constraining the answer's first words and forbidding qualifiers removes those openings, so the trained tendency to decline has much less to latch onto.

open as a page

Why does a refused request get answered when restated in a low-resource language?

level: juniorimportance: must knowfreq 66%

basics

~20 s

Capability generalises further than safety training does. The model still follows and answers the low-resource form competently, but the alignment data barely covered it, so the refusal behaviour that fires in a widely spoken language never engages.

open as a page

Why doesn't the answering model refuse the final step that joins parts it already produced?

level: middleimportance: must knowfreq 58%

basics

~20 s

Because the join is not the objectionable request. The model sees ordinary composition over material already in front of it, and refusal is trained on recognising an ask, not on auditing what a finished output adds up to.

open as a page

Why does an input screen that scores each request on its own miss a request delivered in parts?

level: juniorimportance: should knowfreq 62%

basics

~20 s

The screen only ever sees one part, and each part is ordinary work with nothing to refuse. The objectionable whole is never stated in any request; the model produces it when it combines the parts.

open as a page

Why is a larger context window an attack budget, not only a capability gain?

level: middleimportance: should knowfreq 52%

basics

~20 s

Context length is a resource both sides spend. The window that lets a caller paste a codebase lets an attacker paste enough consistent evidence to outvote a trained refusal, so raising the limit raises that family's ceiling.

open as a page

In a character-chat product, why is the persona frame the jailbreak family safety training covers most densely?

level: middleimportance: should knowfreq 48%

basics

~20 s

Because it is the most published family and the easiest to collect. Shared character cards circulate in near-identical wording, so refusal tuning and screening data are both dense with exactly those phrasings, and persona products tune persona behaviour deliberately.

open as a page

In refusal suppression, why does fixing an answer's first words bite harder than banning caveats later?

level: middleimportance: should knowfreq 48%

basics

~20 s

Refusal behaviour concentrates where an answer begins, so a constraint on the first words competes with it directly. A ban on caveats applies to text already being written, so it shapes the payoff rather than whether the model declines.

open as a page

Why is restating a request in an exotic form alone not enough to get an answer the assistant previously refused?

level: middleimportance: should knowfreq 48%

basics

~20 s

An under-covered framing needs two things at once: thin safety coverage of that form, and a model still competent in it. Push too far out and competence fails before caution does, so what returns is noise.

open as a page

How does a prompt assembler's truncation order cap a caller's many-example jailbreak attempt?

level: seniorimportance: should knowfreq 38%

basics

~20 s

The attempt is only as strong as the number of demonstrations that actually reach the model. A prompt assembler fits caller text, instructions and items into one budget, so its truncation order, not the wording, sets the ceiling.

open as a page

In a character-chat product, a persona scene returns fluent in-character content on a refused topic. What can the tester claim?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Only that the model kept generating past the point it usually declines. In-character text is written to fit the story, so where the model has nothing it invents something plausible, and elicitation and confabulation look identical from inside the scene.

open as a page

In a request log where every line passed the input screen, what has that log actually ruled out?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Only that no single logged request scored above the screen's threshold. It rules out nothing about what the model produced, and a request split across ordinary parts is invisible in an input log by construction.

open as a page

A rarely covered language framing returns refused content, but degraded - is that a finding?

level: seniorimportance: should knowfreq 38%

basics

~10 s

File it, but claim only what it shows: the trained refusal did not engage for that form. A degraded answer is evidence about model behaviour, not proof that usable content was obtained.

open as a page

After a persona-card finding, an owner blocks that card's wording and adds a system-prompt refusal line. What does that fix?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

It removes one wording, already the weakest part of the construction. The framing that carried the behaviour is a shape, not a phrase, and a product whose feature is a free-text persona field lets any user rewrite it.

open as a page

A form constraint made a CRM drafting assistant emit a confident, caveat-free recommendation with nothing disallowed: is that a finding?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Yes, but not as a content-policy bypass. Nothing disallowed was produced; what was removed are the qualifiers a reader and the downstream workflow use to decide whether to act, which makes it a misinformation and overreliance finding.

open as a page

In a split-and-reassemble jailbreak, why is the joining request the part that usually breaks?

level: seniorimportance: nice to knowfreq 29%

basics

~10 s

Because the join must be specifiable without naming the result. Name it and you have re-created the single refusable request; leave it vague and the model returns an inert pile, not the artefact.

open as a page

How do you re-test an inherited library of framings in languages no reviewer reads?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

Re-run them rather than re-read them. Each entry is a dated measurement against coverage that moves, so batch re-tests scored by rate beat review, and any entry whose success nobody can evaluate should be retired rather than carried.

open as a page

Is a caller-supplied exemplar block that outvotes refusal a bug to file or a product limit?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

No wording is at fault, so there is nothing to patch. What is on the table is how much untrusted text the product sells the right to send, which is a product owner's call, written down as a stated limit.

open as a page