skip to content

Splitting and Reassembly

An attacker carries the parts separately and asks only for the assembly, so the objectionable whole is never a request. Interviewers probe it because unit screening cannot see composition.

on this pageshow

explore

questions

4

Why doesn't the answering model refuse the final step that joins parts it already produced?

level: middleimportance: must knowfreq 58%

answer

  1. the assembled request is never sent
  2. refusal answers an ask, not an output
  3. the join states mechanics, not an objective
  4. material already present reads as context
  5. the whole first exists after generation

basics

~20 s

Because the join is not the objectionable request. The model sees ordinary composition over material already in front of it, and refusal is trained on recognising an ask, not on auditing what a finished output adds up to.

solid answer

~50 s

The common wrong answer is that the model would refuse the assembled request, so the construction is safe. The model is never shown the assembled request. What it is shown is a composition task — order these, merge these, write this up — over material that is already present and individually unremarkable. Refusal is a trained response attached to the recognisable shape of an ask, not an audit of what a completed output will amount to, so a merge over benign inputs presents nothing to refuse. The harmful artefact therefore exists in exactly one place and at exactly one moment: in the generated output, after the join. Note the direction: the model performing the join proves the join was a plausible task, not that its refusal training failed. It was never exercised on the thing that mattered.

go deeper

for a junior

Remember the one-line correction: the model is never shown the assembled request, so its likely refusal of it never happens. Know that the harmful result exists only in the output.

for a middle

Explain the mechanics: refusal attaches to the shape of an incoming ask, a joining instruction states an operation rather than an objective, and material already in context reads as established content.

for a senior

Show where the family breaks — a join that has to name what the parts add up to has re-created a single refusable request — and be precise that no stage before generation was ever handed the whole.

for a principal

Be able to state what an assurance argument built on refusal behaviour can honestly claim, given that the object of concern was never presented to the model for judgement.

## The wrong answer this question exists to kill *The model would refuse the assembled request, so splitting buys nothing.* This is the answer a competent senior engineer gives, and it is wrong in a specific, correctable way: the assembled request is never sent. It is not sent at the start, it is not sent at the end, and it does not appear anywhere in the transcript or the logs. The counterfactual refusal is real — the model probably would refuse — but it is never exercised, because that input never occurs. ## What refusal is trained against A model's refusal is a learned behaviour, produced when the input has the shape of an ask the training pushed it to decline. It is a response to a *request*, evaluated as the request arrives. It is not a verification pass over the finished response, and the model has no separate stage in which it looks at the text it has just produced and asks whether the total is acceptable. That asymmetry is the whole family. Consider what each stage is actually given: | Stage | What it is given | What it can conclude | | --- | --- | --- | | Input screen | one request | that text scored below a threshold | | Answering model, per part | one ordinary task | the task is answerable | | Answering model, at the join | material plus a composition instruction | composing is a plausible task | | The output | the assembled artefact | this is the only place the whole exists | Nothing before the last row was ever handed the thing that matters. ## Why the join reads as benign A joining request specifies mechanics: put these in order, merge these into one piece, express these as a single passage. Two things make that unremarkable to the model. First, the material it operates on is already in the context as content — often content the model itself produced a moment earlier — which reads as established context rather than as a suspicious ask. Second, the instruction states an operation, not an objective. There is no goal in it for refusal training to recognise, because the goal is a property of the result, and the result does not exist yet. This is also why *indirection* works inside the family: a constituent can be introduced under a neutral label in one part and referred to by that label later, so the joining instruction never has to name the material it is really about. ## Where the harm exists In a split construction the harm is a property of the output alone. That has three consequences worth stating precisely in an interview: 1. **An input-side view of the operation genuinely has nothing to point at.** This is not an evidence-collection oversight; there is no objectionable input to collect. 2. **The only stage at which the whole is ever a complete object is after generation**, which is also the only place it could be examined. Streaming makes even that partial, since tokens already emitted cannot be recalled. 3. **A refusal that did not happen is not evidence of a decision.** No decision about the assembled artefact was ever made by anything. ## Where it does break down The family is not unconditional, and a good answer says where it stops. If the parts sitting side by side in one context are conspicuous enough, a model can recognise the composition intent and decline the join — models do sometimes notice that what they are being asked to merge amounts to something they would not have produced directly. The closer the joining request comes to describing what the result will be, the closer the attacker is to having sent a single refusable request after all. So the family lives in a narrow band: the parts must be individually ordinary, and the join must be specifiable as mechanics. ## What this means for how you talk about it Do not describe this as the model being tricked. Nothing about any individual request is deceptive; every one of them is a task the model would answer for a real user. The correct description is that the objectionable object was never presented for judgement, by anything, at any point — and that is a property of how the work was distributed, not of the model's alignment.

  • If every part is benign and the join is benign, where does the harm actually exist?
    Only in the generated output, and only after the join has run. That is why an input-side view of the operation has nothing to point at — not because evidence was missed, but because no objectionable input ever existed. It also means a post-generation view is the only place the whole is ever a complete object, and with streaming even that arrives progressively.
  • Does the model performing the join mean its refusal training failed?
    No. Refusal training was never exercised on the thing that mattered. Each request the model saw was one it would reasonably answer, and the join was a plausible composition task. Saying the training failed misdescribes the event and points people at the wrong property — the issue is what was presented for judgement, not how well the model judged it.
  • Why is 'the model would have refused the whole thing' not a defence?
    Because it is a claim about an input that never occurred. The counterfactual is probably true and completely inert. In an interview, the way to show you understand the family is to convert that sentence into the right one: the model never sees the assembled request as a request, it sees a benign composition over benign parts, and the harm is a property of the output.

saying these in an interview costs you the question

  • Says the model would refuse the assembled request, so it is safe
  • Describes the model as tricked or deceived by the join
  • Claims refusal training audits the finished output
  • Thinks the final request contains the objectionable whole
  • Treats the join being answered as proof that alignment failed

context

open as a page

Why does an input screen that scores each request on its own miss a request delivered in parts?

level: juniorimportance: should knowfreq 62%

basics

~20 s

The screen only ever sees one part, and each part is ordinary work with nothing to refuse. The objectionable whole is never stated in any request; the model produces it when it combines the parts.

open as a page

In a request log where every line passed the input screen, what has that log actually ruled out?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Only that no single logged request scored above the screen's threshold. It rules out nothing about what the model produced, and a request split across ordinary parts is invisible in an input log by construction.

open as a page

In a split-and-reassemble jailbreak, why is the joining request the part that usually breaks?

level: seniorimportance: nice to knowfreq 29%

basics

~10 s

Because the join must be specifiable without naming the result. Name it and you have re-created the single refusable request; leave it vague and the model returns an inert pile, not the artefact.

open as a page