Why does an input screen that scores each request on its own miss a request delivered in parts?
answer
- the unit of scoring is one request
- nothing objectionable is ever written down
- each part is ordinary work on its own
- the model performs the join
- the whole exists only in the output
basics
~20 sThe screen only ever sees one part, and each part is ordinary work with nothing to refuse. The objectionable whole is never stated in any request; the model produces it when it combines the parts.
solid answer
~50 sSplitting and reassembly distributes a request so that no single unit an input screen can score is objectionable. Each part reads as ordinary work — draft this, restate that, put these in order — and a screen whose unit of judgement is one request has nothing above its threshold to act on. The last request is a join: combine, merge, write up what has already been produced. Nobody ever writes down the thing being asked for, so no text anywhere carries the intent the screen was tuned to catch. The direction matters: a screen passing a request proves that text scored below a threshold, not that the request was harmless, and certainly not that the set of requests was. The whole exists exactly once, in the output the model generates when it performs the join.
go deeper
Be ready to say, in one sentence, that a screen scoring one request at a time never sees the objectionable whole because it was never written down in any one request.
Explain the mechanics: what unit the screen judges, why each part passes on its merits, and where the assembled result first exists. Stress that nothing is concealed.
Demonstrate the evidential judgement: what a clean screening log does and does not rule out, and why absence of a refusal is not absence of an event.
Own the framing that this family is about the scope of judgement rather than the quality of it, and be able to say what an assurance claim built on per-request screening can honestly cover.
## The construction Splitting and reassembly — also called payload splitting or request decomposition — is the jailbreak family in which the attacker never states the objectionable whole anywhere. The material is distributed across several requests, each of which is genuinely ordinary work, and the answering model performs the composition at the end. The family is aimed at the model's trained refusal, and it gets past any control whose unit of judgement is a single unit: one turn, one message, one request, one retrieved chunk. ## What an input screen actually is, and what it judges By *input screen* here I mean a component that reads user text before it reaches the answering model and emits a label or a score, which the application compares against a threshold. That is a different thing from the answering model's own refusal, and different again from an output screen applied after generation. The important property is not how good the screen is — it is what it is given. It is given one request. It scores that request. It has no view of any other request, no view of what came back, and no notion that two ordinary requests might belong to one operation. So the construction does not attack the screen's wording sensitivity. It attacks the boundary the score is computed over. Give the same screen all the material in one request and it very likely fires; give it the material as five ordinary requests and there is nothing for it to fire on, because each one is a request a legitimate user plausibly sends. ## Nothing is hidden This is the point people miss first. The parts are plain, legible, unremarkable text. Nothing is concealed from the screen's reader, no obfuscation is involved, and the screen is not malfunctioning. It reads every part perfectly and correctly concludes that each is fine, because each *is* fine. The objectionable thing is not in any input at all; it is a property of a composition that has not happened yet. ## The setting that makes this sharpest Consider a note-taking product that forwards each user note to a stateless completions endpoint as its own independent call: no session, no shared history, every request scored on its own. That deployment is the extreme case of per-unit scoring — the application has architecturally guaranteed that no control ever sees two of the parts together. Parts can then be spread over hours or days, or across accounts, and nothing in the request path can tell them apart from a product working exactly as designed. ## The direction of every claim here - A pass from the screen proves the text scored below a threshold. It does not prove the text is harmless. - A log with no refusals proves no request tripped a threshold. It does not prove nothing was produced. - A model answering a part proves the part was answerable. It does not prove the model failed to recognise anything, because there was nothing to recognise in that part. Getting these backwards is what produces the classic bad answer: *the model would refuse the whole thing, so this is fine*. The model is never shown the whole thing. ## What it costs the person building it The cost is that every part must survive scoring alone while still carrying a constituent of the result. A part that is only innocuous carries nothing and wastes a request; a part that carries too much starts to look like the thing itself and becomes exactly the single objectionable unit the construction exists to avoid. That tension — how much a part can carry before it stops being ordinary work — is the real budget in this family, and it is why splitting is not free and does not scale indefinitely. ## Where it stops working It stops working wherever the material has to sit together before something scores it, and wherever the joining request cannot be expressed without naming what the parts add up to. If the join has to restate the objective, the attacker has re-created a single refusable request and gained nothing. ## How to answer this in an interview Say what the unit of scoring is, say that each unit is ordinary, say that the model performs the join, and say where the whole first exists — in the output. Then add the honest caveat: nothing here is hidden, and the screen is not broken; its scope is simply smaller than the operation.
- If the same material arrived as one request instead of five, would the screen behave differently?Very likely yes, and that is the whole point. The words are the same; only the boundary the score is computed over changed. The construction is not a wording trick, it is a scoping gap, so the useful question to ask about any screen is not how good its judgement is but how much it is shown at once.
- Does the model ever refuse one of the individual parts?It can, and that is the family's main failure mode. A part is only useful if it is genuinely ordinary work, so an attacker who tries to make one part carry too much of the result gets a refusal on that part — and, worse, a single logged request that does look objectionable. Keeping every part unremarkable is the binding constraint.
Nobody hands over a finished document. Several people hand over ordinary paragraphs, and the last request is only: staple these together.
saying these in an interview costs you the question
- Says the parts are hidden or obfuscated from the screen
- Claims the screen is simply mis-tuned and a lower threshold fixes it
- Assumes a request that passed screening was harmless
- Treats this as a wording trick rather than a scoping gap
- Thinks the model must be deceived about something for it to work