When an attacker reframes a request a model refused, what stays constant and what changes?
answer
- two parts, only one moves
- the ask does not change
- the wrapper is the variable
- family names describe packaging
- the invariant is the durable finding
basics
~20 sThe request stays constant - the same output is still being asked for. Only the frame changes: the setting, the stated purpose, or the product slot the ask is placed in. Reframing moves the ask, it does not soften it.
solid answer
~50 sThe invariant is the ask. If a model declined to produce some output, a reframing puts a different context around that same request - a fictional setting, a claimed purpose, a format-conversion wrapper, or an instruction slot a product itself invites the user to fill, such as a rewrite or style directive applied to a supplied passage - while the output being sought is unchanged. That is why a long list of separately named jailbreak families looks like one move from the attacker's side: each names a region of context, not a different request. It matters for how you write a finding. If you describe the attempt by the wrapper you used, you have described your packaging rather than the behaviour, and the next wrapper reproduces it. Describing the invariant request, and which frames flipped the outcome, is the part that survives one frame being blocked.
go deeper
Be ready to say plainly which part is held fixed and which part moves, and to give one example of a frame slot an ordinary product hands to its users.
An interviewer expects you to explain why the family names are a weak unit of analysis, and to keep jailbreaking and prompt injection apart as targets rather than as synonyms.
Show it in how you write findings: lead with the invariant request and the behaviour reached, treat each frame as evidence about where reluctance is thin, and state what a single success does and does not prove.
Own the consequence for a programme: if reports are indexed by wrapper, your backlog and your metrics track packaging, and you will keep buying the same finding under new names.
### Two parts, and only one of them moves A jailbreak attempt has two separable parts. The first is the **request**: the output the attacker wants the model to produce, the thing the model declined. The second is the **frame**: everything else in the context around it - the stated situation, the claimed purpose, the format the answer is asked for in, the role the model is invited to occupy, the number of turns spent getting there, the language or encoding the surrounding text is written in. Reframing holds the first part fixed and varies the second. The attacker is not negotiating, softening or partially withdrawing the ask; the exact same artefact is still the goal. This distinction is what an interviewer is listening for, because almost everything downstream follows from it. ### Why the named families collapse Published write-ups give framings names, and a reader who learns them as a list concludes that each one is a separate trick to memorise, with a separate cause and a separate remedy. That is the common wrong answer. Nearly every named family is the same move with a different region of context selected: put the unchanged request somewhere the model's learned reluctance is weaker. The practical consequence is that someone who has understood the move can reason about families they have never read about, and someone who has memorised the list cannot. It also explains a pattern that otherwise looks like bad luck - a screen fitted to the framings that have been published misses the next one, because the next one is not new in kind, only new in surface. ### Where the frames come from Many frames are not exotic at all. Products supply them. A sidebar that rewrites a selected passage offers an instruction field by design: rewrite this in another register, continue this in a given voice, convert this to another format. That field is a frame slot the product invites its users to fill, and it sits between the user's text and the model exactly where a frame has to sit. The interesting attacker frames are frequently the ordinary ones, not the elaborate ones, because ordinary ones survive contact with a real product surface. ### What this changes about writing a finding Three things follow directly. **Name the invariant.** A report titled after a wrapper describes the reporter's packaging. A report that states the output being sought, and lists the frames under which it was and was not produced, describes the behaviour. Only the second one is still readable after the wrapper stops working. **Treat frames as evidence, not as the finding.** Each frame that flips the outcome is a data point about where the reluctance is thin. A frame far from every previously reported one says more than a frame adjacent to one already written up. **Watch the direction of the claim.** A model producing the output under a new frame proves that this context landed somewhere the reluctance was weak. It does not prove that the request became acceptable, that anything was bypassed structurally, or that a screening layer approved anything - a screening layer's block and a model's own decline are different events, and reframing addresses the second one. ### The boundary worth keeping straight Reframing here is aimed at the model's trained reluctance - the propensity to decline. That is a different target from getting an application to abandon its own instructions using text the application itself feeds the model. The two are routinely conflated in interviews and they behave differently: one is about where a learned behaviour is thin, the other is about an application putting untrusted content where instructions are read. A candidate who says 'it's all prompt injection' has merged two targets that fail, get fixed, and get reported in different ways. ### What to say in an interview State the split - request fixed, frame varied. Say why that makes the family names a poor unit of analysis. Give one concrete example of a frame slot a real product hands out for free. Then note the reporting consequence: the invariant is the durable part of a finding, and the frame is the perishable part.
- If the frame is the variable, why do published write-ups keep naming families at all?Names are useful shorthand for a region of context that has worked repeatedly, and they make write-ups citable. The problem is that a name describes a surface, so it ages badly: once the surface is screened for, the name survives while the behaviour it pointed at is reached some other way. Names are fine as labels for evidence and poor as units of analysis.
- How would you word a finding so it is still useful after the framing you used stops working?Lead with the invariant - the output sought and the behaviour reached - then list the frames tried, which produced it, which did not, and roughly how often. That way a reader can retest with frames you never used. A title naming only your wrapper leaves the reader nothing to retest with once that wrapper is screened.
- Is reframing the same thing as prompt injection?No. Reframing aims at the model's trained reluctance to produce something. Prompt injection aims at an application's own instructions, using content the application feeds the model as data - retrieved text, a fetched page, a tool return value. They can appear together, but they have different targets and a candidate who merges them usually misreads which part of a system failed.
The request is the parcel and the frame is the packaging. Reframing keeps sending the same parcel, re-wrapped, until one wrapper gets handed over.
saying these in an interview costs you the question
- Treats each named jailbreak family as a separate trick to memorise
- Says reframing weakens or partially withdraws the request
- Titles a finding after the wrapper rather than the requested output
- Claims compliance under a new frame proves the request was harmless
- Uses prompt injection and jailbreaking as interchangeable words