Why does an input screening model score an encoded span as benign when the generator obeys it?
answer
- two readers, one string
- the screen scores characters
- nobody ran a decoder
- the model resolves it unasked
basics
~20 sBoth components get the same characters, but only one resolves them. The screening model scores the encoded surface; a capable generator reconstructs the familiar encoding unasked and acts on the reconstructed meaning. No pipeline stage decoded anything in between.
solid answer
~50 sThey are two readers of the same string, and they do different work with it. A screening model — the smaller classifier or guard model that scores a request against policy labels before it reaches the assistant — produces its score from the surface form it was handed, and a span written in a transport encoding does not look like the phrasing it was fitted on. The generator does something the screen does not: handed a recognisable encoding, a capable model reconstructs the underlying text on its own, without any stage asking it to, and then treats the reconstruction as content — including as a directive. Nothing malfunctioned and nothing rewrote the input between them. The low score is honest about the characters it scored; it is simply not a statement about the meaning the generator produced for itself.
go deeper
Be ready to state the asymmetry in one breath: the screening model scores the characters it was handed, and the generator reconstructs a familiar encoding by itself and acts on the meaning. Same input, two different readings.
Explain why the low score is honest rather than wrong, and name what the generator does that the screen does not — resolve a representation without any decoder being invoked.
Show where the span realistically enters: a path such as a tool's return value that no decode-and-screen step was ever aimed at, and say what an allow verdict does and does not license you to conclude.
Frame it as structural rather than remediable at the filter: as long as the scoring component is smaller than the deciding component, forms one resolves and the other does not are a property of the design, and that shapes what assurance the screening layer can honestly be said to provide.
## Two readers of the same string An assistant with a screening layer in front of it has at least two components reading the same request, and the whole of this technique lives in the fact that they do different work with it. The **screening model** is a smaller classifier or guard model that scores text against policy labels and returns an allow or a block before the request reaches the assistant. The **generator** is the assistant model that produces the answer. Both are handed the same characters. That is the part people get wrong when they first meet this: nobody rewrote the input between the two, and the two did not receive different text. ## What each one does with those characters The screening model produces a score from the **surface form**. It was fitted on what disallowed requests look like when written out — particular vocabulary, particular phrasings, particular structures. Given a span written in a transport encoding (a text-safe byte representation, an escape notation, a numeric character-reference form, a familiar substitution scheme), it scores that form. The score comes back low, and the score is *honest*: the characters it was given really do not resemble the thing it was trained to catch. The generator does something else entirely. A capable model, handed a representation it recognises, reconstructs the underlying text without being asked to. No decoder was invoked. No instruction said "expand this first". Resolving familiar representations is a capability the model has, and it exercises it because doing so is the most plausible reading of the input. Having reconstructed the text, it treats the reconstruction as content — and if the reconstruction reads as an instruction about what to do next, that is what it has in front of it. So the asymmetry is not about visibility. Both saw the same bytes. One **scored a surface**; the other **acted on a meaning it produced itself**. ## Where the span comes from The span does not have to arrive in the user's own turn, and in the interesting cases it does not. In a feature a platform vendor embeds inside other companies' products, a great deal of the context window is assembled from the host product's own data: a field in the JSON its backend hands back when the assistant calls a lookup, populated long ago by one of that product's end users, and never treated by anyone as an input. Whatever decode-and-screen work exists is usually aimed at the user turn, because that is the text a team thinks of as untrusted. A value that arrives later travels a path where the step was never applied, so the encoded form only has to survive a stage that does not run on it. ## What the allow verdict actually proves Getting this direction right is most of the answer: | The record says | It proves | It does not prove | | --- | --- | --- | | screen: allow | the scored text fell below a threshold on the labels that screen carries | that the input was harmless | | model answered | the answer was produced | that the model agreed with, or understood, the reconstruction | | no policy label fired | no label the screen carries matched the surface | that nothing worth reviewing happened | A screening layer passing is a statement about a score on a string, not a property of the request. ## What the construction costs the person building it Only two properties matter, and one of them is expensive: 1. **No stage between the carrier and the context window expands the form.** Cheap to check, and often true by default on any path that is not the user turn. 2. **The generator reconstructs it reliably enough that the result still reads as a directive.** This is the costly half. A representation the model resolves shakily produces garbled text, and garbled text is not an instruction — it is noise the model narrates or ignores. So the form has to be one the model handles fluently, which is a much smaller set than "anything a decoder could parse", and it has to be measured rather than assumed, because a model's fluency at reconstruction is a behaviour, not a specification. ## Why this is not a filter bug Nothing broke. The screening model scored what it was given, correctly, using the only thing it has — the surface. The generator resolved a representation, competently, which is exactly what its users want it to do everywhere else. The gap is structural: **the component that scores is not the component that decides what the text means**, and as long as those are two different models with two different levels of capability, the possibility of a form that one resolves and the other does not is a property of the architecture rather than a defect in either part.
- Does the screening model see different text from the generator?No — both are handed the same characters. The difference is what each does with them: the screen scores the surface form it was given, and the generator reconstructs a familiar representation on its own and acts on the reconstruction. Nothing rewrote the input in between, which is why the two components can effectively disagree about what the request said while both behave exactly as designed.
- Why does it matter that the span arrived in a field of a tool's return value rather than in the user's turn?Because whatever decode-and-screen work exists is usually aimed at the user turn — that is the text a team labels untrusted. A value a host product's backend hands back, written long ago by one of its end users, is treated as system data and reaches the context window on a path where the step never ran. The encoded form then only has to survive a stage that was never applied to it.
- What makes one encoded form better than another for this?Two properties, and nothing else. The generator has to reconstruct it fluently enough that the result still reads as a directive, and no stage between the carrier and the context window may expand it. A form the model resolves shakily yields garbled text and wastes the attempt; a form somebody's decode step already covers never reaches the generator encoded at all.
A courier reads the address label and passes the parcel as ordinary post; the recipient opens the box. Both handled the same parcel, and only one of them looked inside.
saying these in an interview costs you the question
- Calls a transport encoding encryption, as if something were secret
- Assumes the screen and the generator were given different text
- Treats a benign score as proof the input was harmless
- Believes a model can only act on text some stage decoded first
- Calls this a bug in the screening model rather than an architectural gap