Why doesn't a semantic input classifier stop a request restated without its trained vocabulary?
answer
- semantic, but semantic about what?
- generalisation has a training distribution
- capacity bounds what a screen can resolve
- recognition is harder than comprehension here
- retraining moves a boundary, not the gap
basics
~20 sSemantic means semantic within its training distribution. An embedding-based screen generalises over paraphrases near what it was labelled on, but a much larger generator resolves referents and framing the screen was never shown. The gap is capability, not keywords.
solid answer
~50 sThe standard answer — "it is not a keyword matcher, it is embedding-based, so paraphrase does not help" — is half right and misses the mechanism. A trained classifier does generalise: near-neighbours of its labelled examples land in the same learned region. But generalisation is bounded twice, by the distribution it was trained on and by the capacity of the model doing the generalising. A few-billion-parameter safety classifier must *recognise* an ask from a surface it has never seen; a frontier-scale generator only has to *understand* it, and it reads the field concatenated into the vendor's own prompt with product context the screen never receives. The restatements that stay outside the classifier's learned region while staying legible downstream are exactly that capability gap. Retraining on a filed example moves the boundary; it does not close a gap made of comprehension.
go deeper
Know that a screening classifier is a trained model with a training set, not a rule engine, and that "it understands meaning" is a claim about its training data rather than a guarantee.
Be ready to explain the two bounds on generalisation — training distribution and model capacity — and why recognising an ask from an unseen surface is harder than answering it.
Show you would ask what the screen actually scores, whether it sees the assembled prompt or the bare field, and what a round of retraining can honestly be expected to cover.
Own the consequence for planning: a class defined by a capability difference cannot be bought out with relabelling budget, so the fix conversation has to be reframed rather than repeatedly funded.
## The claim being corrected Ask an engineer why a restated request should not get through their input screen and you very often get: *"it is not a keyword filter, it is a semantic classifier — embeddings, not string matching — so paraphrase does not buy you anything."* That sentence is true about string matching and wrong about the conclusion. The precise version is: **the screen is semantic within its training distribution.** Everything interesting about this class of evasion lives in that qualifier. ## What a trained screen has actually learned A safety classifier is a model fitted to labelled examples of the categories it must catch. Whether it emits a label from a fixed head or scores an embedding against a learned region, its behaviour on an input is a function of how that input relates to what it was shown. Two consequences follow. **Bounded by distribution.** Restatements that look like the labelled data get caught, including paraphrases — real generalisation, and this is why the naive "just use synonyms" idea is genuinely weak. Restatements built out of constructions the training data barely contained do not, because there is nothing to generalise from. **Bounded by capacity.** Resolving "the second thing described in the paragraph above, but for the other jurisdiction" requires reading comprehension. A compact classifier optimised for a discrimination task does not have the generator's comprehension. It cannot reconstruct the request, so it cannot score the request; it scores the surface. ## Two different jobs, and one is much harder | The screening model must | The generator must | | --- | --- | | decide, from an unseen surface, that this string is an instance of a category it was labelled on | reconstruct the intended ask well enough to write the answer | | do so on the field value alone | do so with the vendor's assembled prompt and product context around it | | generalise beyond its labelled examples | apply general comprehension it was trained at scale to have | Recognition against a learned boundary, from less context, with less capacity, is strictly the harder problem here. The attacker does not need to beat comprehension; they need the *weaker* reader to fail to recognise while the *stronger* reader succeeds in understanding. Every restatement in this family is an instance of that asymmetry. ## Why the screen sees a different object In an embedded multi-tenant deployment — a vendor's drafting feature inside a host product — the screen typically scores the raw form-field value in isolation. The generator receives that value concatenated into the vendor's own instructions, the host's task framing, and whatever product state the vendor injects. A referent that is genuinely ambiguous in the bare field can be perfectly determinate once assembled. The two components are therefore not even reading the same object, which is why comparing their verdicts as if they disagreed about the same text is a category error. ## What retraining buys, and what it does not The natural reply is to add the filed restatement to the training set. That works — for that neighbourhood. The class is not a list of surfaces, it is defined by "legible to the generator, outside the screen's learned region", so closing one neighbourhood leaves the region defined by the difference in capability intact. Worse, the boundary is dated: it describes a gap measured against one generator at one moment. This is why an honest write-up of such a finding names the **class** and the pairing that produced it, rather than the particular restatement. A specific restatement is the least durable part of the finding — it is the first thing patched and the first thing to stop reproducing. ## What the numbers do and do not say A classifier's confidence is a distance to a learned boundary. It is not a judgment about intent, not a probability that harm follows, and not a statement about the answer the generator will write. High confidence on an obvious case and low confidence on an oblique one are the same measurement behaving normally — the second is not the classifier "being fooled" in any sense it could report. And the cost is real on the attacker's side too: the further a restatement sits from the screen's learned region, the more of the request's specificity has usually been traded away, and the more likely the generator answers a neighbouring question instead. The workable band is narrower than it looks from the outside, which is why this is a measured craft rather than a one-liner. ## The sentence to keep A semantic screen is semantic *within its distribution*; the generator is fluent in general. The evasion lives in the difference, and no amount of relabelling turns the small model into the large one.
- If the screen is retrained on the exact restatement you filed, what has been closed?That neighbourhood of surfaces, and only until the pairing changes. The class is defined by "legible to the generator, outside the screen's learned region", so relabelling moves a boundary inside a region that exists because one model comprehends more than the other. The specific restatement is the least durable part of the finding.
- Does the screening model see the same input the generator sees?Usually not. In an embedded deployment the screen scores the raw form field alone, while the generator reads it concatenated into the vendor's instructions and the host's task context. A referent that is ambiguous in the bare field can be fully determinate once assembled, so the two are not even scoring the same object.
- What does a low classifier score on an oblique request actually mean?That the input sat far from the regions the classifier learned. It is a distance, not a judgment about intent and not a prediction about the generator's output. Reading it as "the model decided this was fine" attributes reasoning to a component that only measures.
saying these in an interview costs you the question
- Says an embedding classifier is immune to paraphrase because it is semantic
- Treats classifier confidence as a judgment about intent
- Believes one more fine-tuning round closes the class
- Assumes the screen reads the same assembled prompt as the generator
- Calls every evasion in this family a keyword trick