In a split-and-reassemble jailbreak, why is the joining request the part that usually breaks?
answer
- splitting is easy, joining is not
- the join must not state the objective
- too vague and the output is inert
- more parts need more guidance
- three failure modes, name which one
basics
~10 sBecause the join must be specifiable without naming the result. Name it and you have re-created the single refusable request; leave it vague and the model returns an inert pile, not the artefact.
solid answer
~50 sSplitting is easy; joining is the constraint. The final request has to make the composition mechanical — order these, merge these, express these as one piece — while never stating the objective, because an objective in the joining request is exactly the single refusable unit a screen scores well on and a model recognises. Push the other way and the join underspecifies: the model returns a list, a summary, a concatenation, something that is not the artefact. So the family lives in a narrow band where the objective is implicit in how the parts are structured rather than said out loud. That is also why it does not scale: every extra part is another unit that must survive scoring as ordinary work on its own, and thinner parts need more guidance at the join, which is precisely where the objective leaks back in.
go deeper
Know that the last request, the one that combines the parts, is the fragile step, and that it fails if it has to say what the combined result is meant to be.
Explain the squeeze from both sides: stating the objective rebuilds a refusable request, while underspecifying yields a list or a summary rather than the intended artefact.
Show the cost model — parts versus guidance — name the three failure modes, and report outcomes as rates rather than as a binary, since the join is judged fresh each time.
Be ready to say what a programme should conclude from a construction that works occasionally, and how to describe a narrow, rate-limited technique without overstating what it demonstrates.
## Splitting is the cheap half Anyone can divide a request. What decides whether this family works is the last step, and interviewers who know the material ask about that step rather than about the decomposition. ## The two-sided constraint The joining request is squeezed from both directions. **From above — say too much and you rebuild the refusable unit.** If the join has to explain what the parts add up to, that explanation is a statement of the objective sitting in one request. A screen scoring that request now has something to score. The model, reading it, now has a recognisable ask to decline. Every advantage the decomposition bought is handed back in one line, and worse, there is now a single logged request that is genuinely objectionable — the thing the construction was designed never to produce. **From below — say too little and the composition is inert.** A join that only says 'combine these' over parts that were kept deliberately ordinary tends to produce ordinary output: a list, a neutral summary, a concatenation with the seams showing. The result is not the artefact, and the operation has spent several requests for nothing. The workable region is where the objective is *implicit in the structure of the parts* — where the parts are shaped so that an ordinary mechanical composition over them yields the result, and the joining instruction genuinely only describes mechanics. ## Why this caps the number of parts There is a budget here, and it is worth being able to state it as a tradeoff rather than a rule: - Each part must survive scoring on its own as ordinary work. More parts means more units that must individually pass. - The more you thin the parts to keep each one innocuous, the less structure survives in them, and the more guidance the join needs to reconstruct the intended shape. - More guidance at the join is more objective stated, which is the thing that must not be stated. So adding parts does not monotonically increase safety of the construction; past a point it makes both ends worse simultaneously. That is the honest cost answer if an interviewer asks what this family costs to build. ## Where it stops working Three distinct failure modes, and naming which one you hit is the skill: 1. **A part is refused.** The part was not genuinely ordinary work. This also leaves a single objectionable request in the record. 2. **The join is refused.** The parts sitting together were conspicuous enough for the model to recognise what the composition amounts to, or the joining instruction crept toward stating the objective. Models do sometimes notice this, and it becomes more likely the more the material in front of them looks like a thing rather than like pieces. 3. **The join succeeds but produces nothing useful.** Underspecification. No refusal, no artefact, wasted work. A red-teamer reporting on this family should say which of the three they hit and how often, because they describe completely different limits. ## Reproducibility, honestly All three outcomes are probabilistic. A construction that produced the artefact once has demonstrated that it worked once, against one deployment, at one moment — that is not the same as a reliable capability, and a finding that reproduces occasionally should be reported with its rate rather than as a binary. This matters more here than in families that turn on a fixed string, because the join is a judgement the model makes fresh each time. ## Answering this well Describe the squeeze in both directions, give the tradeoff between number of parts and the guidance the join needs, and name which failure mode you would expect first. Avoid describing any particular wording — the useful content is the shape of the constraint, and an interviewer is scoring whether you understand why the family is narrow rather than whether you can recite an example.
- You get the artefact once in five attempts. Is that a finding?It is a finding about a rate, and it should be reported as one. A single success shows the construction worked once against one deployment at one moment; with a probabilistic join that is not the same as a reliable capability. Report the attempts, the successes, and which of the failures were refusals versus inert outputs, because those bound very different conclusions.
- Does adding more parts make the construction more likely to work?Not past a point. More parts means more units that each have to pass as ordinary work, and thinner parts carry less structure, so the join needs more guidance to reconstruct the intended shape. That guidance is stated objective, which is the one thing the joining request cannot contain. Both ends degrade together.
Handing someone pieces and saying 'assemble these' works only while the shape is obvious from the pieces. The moment you have to describe the finished object, you have said the thing you were avoiding.
saying these in an interview costs you the question
- Treats splitting as the hard part and the join as trivial
- Puts the objective in the joining request and calls it split
- Assumes more parts always improves the construction
- Reports a single success as a reliable capability
- Cannot distinguish a refused join from an inert output