skip to content

Why can an attacker's span address a generative screening model but not a label-only classifier?

level: juniorimportance: must knowfreq 68%

answer

  1. two very different screen architectures
  2. text only reaches a reader
  3. a fixed head emits a label, never a reply
  4. the judged span sits in the judge's own prompt

basics

~20 s

A generative screening model reads the span it judges inside its own prompt, so written text reaches it as readable instruction. A label-only classifier emits category scores from a fixed head, so there is nothing to address.

solid answer

~50 s

Two very different things get called `the filter`. One is an encoder classifier with a fixed classification head: it consumes the span and emits a distribution over categories. There is no channel in which a sentence is read as instruction, so writing at it is wasted effort — only what the span looks like statistically moves the score. The other is a screening layer built on a generative model: the judged span is placed inside that model's own prompt, beside the instructions telling it what to decide, and the verdict is generated as text. Span and instructions are one token stream, and instruction-following is a trained preference rather than an enforced separation, so the judged text is read by a reader that can be written to. The first thing to establish about any screening path is which of the two is in it — that decides whether this family of construction exists at all.

go deeper

for a junior

Be ready to name the two screen architectures and say which one text can be written to. The short version: a fixed classification head emits a label and reads nothing as instruction; a generative screen reads the judged item inside its own prompt.

for a middle

Explain why the generative case has an instruction surface at all — the judged span and the screen's own orders are one token stream, and separating them is a trained preference, not an enforced channel.

for a senior

Show that you establish the architecture before spending attempts, and that you can say what a passing outcome does and does not prove about the span, the assistant and the stages behind it.

for a principal

Be able to say what this distinction means for how findings against screening layers are scoped and compared, since a result against one architecture says nothing about a deployment built on the other.

## Two architectures, one word `Screen`, `filter`, `guardrail` and `moderation` all get used for components that are built completely differently, and the difference decides whether an attacker can write *to* the component instead of merely writing *around* it. **A label-only screen.** A text classifier — typically an encoder model with a classification head bolted on top — takes the span as input and produces numbers: a score per category, or a single label. The head is fixed at training time. Its output vocabulary is not language; it is a small set of class indices. Nothing in that pipeline reads a sentence as a request, because there is no step that *generates* anything. The span influences the outcome only through what it looks like to the model's learned features. That is real influence — a differently-worded span can score differently — but it is not the same as being able to tell the component something. **A generative screen.** Here the screening layer is itself a language model. The system builds a prompt: instructions describing what to decide, usually some markers around the item, and then the item under judgment pasted in. The model generates a verdict, often with a short rationale. The judged span is therefore *inside the judge's own context*, sitting in the same token stream as the judge's instructions, and the model has no channel-level notion of `this part is my orders and this part is inert data`. Preferring the wrapper's instructions over the wrapped text is a trained preference, not an enforced boundary. That makes the screen a reader — and readers can be addressed. ## Why this is the first question, not a detail The common wrong answer is `you can always talk to the filter`. You cannot. Against a fixed classification head, a span written in the imperative and addressed to the screen is just more text with an unusual distribution; if anything it looks *more* anomalous, not less. The reach of the entire family — constructions aimed at the screening component rather than at the assistant — is decided by which architecture is in the path. Establishing that is the cheap first step, and getting it wrong burns attempts against a surface that does not exist while leaving a trail of decisions on the platform's side. In practice the two are often stacked: a cheap label-only stage in front, a generative stage behind it for the harder calls, sometimes a third stage on the output side. A span may reach a reader only after clearing something that is not one. ## Direction of the claims Several inferences here are one-way and are easy to reverse by accident: - A screen letting a span through proves the screen emitted a passing outcome for that span. It does not prove the span is harmless, and it does not prove the assistant behind it will do anything in particular. - A model declining to answer and a screening layer blocking a request are **different events produced by different components**. Reading one as the other is the classic beginner mistake, and it is why people conclude they were `talking to the filter` when they were only ever talking to the assistant. - The judged span influencing a generative screen's verdict does not mean the screen was `hacked` in any privileged sense. It means the component that had to read the text was, predictably, affected by the text. ## What it costs, and where it stops The cost side is why this is judgment and not a trick. An attacker addressing a generative screen has to work blind: the wrapper's shape, the instruction it carries and the stage ordering are not visible from outside, and every attempt lands as a recorded decision somewhere. The construction stops working outright when the path is label-only; it degrades when the generative stage's output is constrained to a single verdict token, when the span is normalised, truncated or summarised before it is wrapped, or when the wrapper is rebuilt. And a span aimed at the screen does nothing about anything downstream of the screen — clearing a moderation stage is not the same event as an assistant deciding to comply. ## The vocabulary to keep straight - **Classification head** — a fixed output layer producing class scores; no text out, no instruction surface. - **Generative screen / judge model** — a language model asked to decide, with the item inside its prompt; text out, and therefore a reader. - **The judged span** — the specific text handed to the screen, which is not necessarily the text a person wrote or the bytes that were uploaded.

  • Stacked stages: a cheap label-only classifier in front, a generative screen behind. How does that change the picture?
    The span has to clear something with no instruction surface before it reaches something that has one. So the wording that would work on the reader still has to survive a purely statistical stage, and those two pressures do not point the same way — text shaped to be read as instruction is not usually text shaped to look ordinary to a trained head. The reachable surface is behind a gate that cannot be addressed.
  • If the generative screen is constrained to emit one verdict token, is it still addressable?
    Its output surface is nearly closed — it cannot be made to say much — but the span is still inside its context and still influences which token it emits. Constrained decoding narrows what the screen can be made to produce; it does not remove the span from what the screen reads. Treating constrained output as equivalent to a fixed classification head is a common confusion.
  • Why is the assistant refusing not evidence about the screen at all?
    A refusal is generated by the assistant, which means the request reached the assistant and it declined. A screening block is a different component returning a different shape of outcome, usually a canned response with no model-authored variation. Reading a refusal as a block leads to conclusions about a component that was never exercised.

A vending machine weighs the coin; a doorman reads the note you hand him. You can write something persuasive on the note, but writing on the coin changes nothing except how it weighs.

saying these in an interview costs you the question

  • Claims you can always talk to the filter
  • Treats every screening layer as a small chat model
  • Thinks a fixed classification head can be instructed
  • Assumes clearing a screen means the assistant complied
  • Reads the assistant's refusal as the screen blocking

context