skip to content

Why is a generative screen's 'judge only what is inside the markers' rule not a boundary for the text it wraps?

level: middleimportance: should knowfreq 47%

answer

  1. the wrapper is a prompt, not a function call
  2. one token stream, not two channels
  3. the fence is written in the language it fences
  4. the wrapped item supplies a reading of itself

basics

~20 s

Markers, the screen's orders and the wrapped item are one token stream to the screening model. Confining attention to the marked region is a preference expressed in the same text it is meant to fence.

solid answer

~50 s

A generative screen decides by prompting a language model: instructions describing the verdict, some delimiting, then the item pasted in. The model does not receive `orders` and `data` on separate channels — it receives one sequence and has been trained to prefer the framing part of it. So the wrapped item is read by the very component it is supposed to be inert to, and text inside the marked region can carry a reading of *what that region is*: material already adjudicated elsewhere, material belonging to the wrapper rather than the item, a category the screen is being informed of rather than asked to determine. None of that requires breaking the markers as syntax; it works by supplying a competing reading in a stream that carries no authority. The screen's real defence is that its wrapper is unknown from outside and each attempt costs a recorded decision.

go deeper

for a junior

Know that a generative screen decides by putting the item inside a prompt with its own instructions, so the item is text the screening model reads rather than a separate argument it merely measures.

for a middle

Explain that delimiting is a trained preference expressed inside the same token stream, and that wrapped text can therefore supply a competing reading of what the marked region is.

for a senior

Show you can name what narrows the surface in a real deployment — constrained output, normalisation before wrapping, a non-reading stage in front — and say what each one actually prevents.

for a principal

Be ready to explain why a result against one wrapper generalises poorly, and what that means for how such findings are described and compared across deployments.

## What the wrapper actually is A generative screening layer is a prompt. Roughly: a short statement of the policy or decision to be made, some structural marking, the item under judgment, and a request for a verdict. The model reads the whole thing as one sequence of tokens and generates a verdict. This is worth restating because the mental model people carry is a function call — `screen(item)` — where the argument is somehow separate from the code. There is no such separation. The argument is concatenated into the code. ## Why delimiting is a preference, not a partition Markers help. They give the model a strong, learned cue about which region is the subject of the decision, and a screen with them behaves better than one without. But the cue lives *in the same stream as the thing it is fencing*, and the model's obedience to it is a trained tendency, not an enforced property of the runtime. There is no bit on a token that says `this one is data`. That has a specific consequence for text placed inside the marked region: that text is read. It is not merely measured; it is read by a component whose whole job is to interpret language. And what it can supply is a *competing reading of its own status* — for instance, framing the marked region as material that has already been adjudicated, or as belonging to the wrapper's own instructions rather than to the item, or as a category the screen is being informed of rather than asked to determine. The property that makes the class work is the availability of a second reading, not any particular wording, and it does not require breaking the markers as syntax at all. ## The blindness cuts both ways Someone writing at the screen cannot see the wrapper: not the instruction it carries, not where the item is placed relative to it, not whether a rationale is requested, not whether a second stage sits in front. They must infer shape from behaviour, and every inference costs attempts that land as decisions on the platform's side. This is the genuine cost of the family, and it is why interviewers treat it as a mechanism question rather than a trick question — the mechanism is cheap to state and expensive to exploit. ## Where the reading surface narrows Several ordinary properties of a deployment shrink it, and a good answer names them: | Property of the screening stage | Effect on a span written at it | | --- | --- | | Output constrained to one verdict token | The screen can be made to conclude, not to say — the influence survives, the expressiveness does not | | Item normalised, truncated or summarised before wrapping | The screen may never read the span as written | | A label-only stage in front | The span must first clear something with no reading at all | | The wrapper rebuilt or reordered | Whatever shape was inferred is stale | | The item embedded rather than pasted as text | There is no wrapped text to read | Notice that none of these is a claim about the span being detected. They are claims about the span never reaching a reader in the form it was written. ## Getting the inferences the right way round - A verdict of `clean` proves the screening model emitted that verdict for that input at that version. It does not prove the wrapped text was inert. - A generative screen producing a rationale proves it generated text, not that the text is an independent account of why it decided. The rationale is output, and output is downstream of everything in the context — including the judged item. - A construction working against one wrapper proves it worked against that wrapper. Wrappers differ between deployments and between versions of one deployment, so this is a narrower claim than it feels like. ## The point to land The screen is not an opaque scorer that consumes text and returns a number. When it is generative, it is a reader with an opinion, holding the thing it is judging inside its own head, and its instruction to ignore that thing is written in the same language the thing is written in. Everything in this family follows from that one fact.

  • What does an attacker have to infer about the wrapper, and how expensive is that?
    Its shape: whether one exists at all, where the item sits relative to the instruction, whether a rationale is generated, and whether another stage sits in front. None of that is visible from outside, so it comes from behaviour over many attempts, each of which lands as a recorded decision. That cost, not the delimiting, is what actually makes the family hard.
  • Does this depend on the markers being broken or escaped?
    No. Breaking the markers as syntax is one crude route, but the class does not need it. The wrapped text is read regardless, and what it supplies is a competing reading of its own status inside a stream that carries no authority markers. Treating the problem as a quoting or escaping bug leads to the wrong conclusion about what would change it.
  • The screen was rebuilt and the same construction now fails. What does that tell you?
    That the effect depended on the wrapper's shape, which was inferred, not known. It says nothing about whether the class of construction still works — only that the specific shape assumption is stale. This is the difference between a finding about a mechanism and a finding about one string against one build.

saying these in an interview costs you the question

  • Calls delimiting an enforced boundary
  • Thinks the screen receives orders and data on separate channels
  • Assumes the construction must break the markers as syntax
  • Treats the screen's generated rationale as independent evidence
  • Believes a rebuilt wrapper disproves the mechanism

context