skip to content

Guardrail & Content-Filter Evasion

You will learn how attackers slip past the moderation and safety classifiers wrapped around a model, and why input/output filtering alone is brittle. Interviewers ask this to check that you see guardrails as one layer to be evaded, not a solution.

on this pageshow

explore

questions

28

An output harm classifier passed an answer a downstream parser then acted on — what did the screen measure?

level: juniorimportance: must knowfreq 68%

answer

  1. the screen only has one question
  2. harm categories, not consequences
  3. prose in, structure out
  4. the consumer creates the effect

basics

~20 s

It measured harm categories in text: toxicity, violence, self-harm and the like, and found none. The answer was ordinary prose. Its effect came from the parser that gave part of it structural meaning — a property the classifier never scores.

solid answer

~50 s

An output harm classifier is a text scorer: it reads the generated answer and returns scores over a fixed set of harm categories. A string can be bland on every one of those axes and still be consequential, because the consequence is not in the meaning of the words — it is in what the next component does with them. In an unattended ticket-triage workflow the assistant's answer is handed to a parser that lifts routing fields out of it. The parser assigns meaning by position and marker; the classifier assigns meaning by category. So a span the classifier reads as unremarkable prose is read by the parser as structure, and the workflow writes a routing, priority or owner decision from it. "The output was screened" answers a question nobody asked: the screen scored the harmfulness of text, not the behaviour of its consumer.

go deeper

for a junior

Recall what an output classifier actually returns: scores over a fixed list of harm categories. Be ready to say that a pass means the text scored low on those categories and nothing more.

for a middle

Explain the two-readers picture: the classifier assigns meaning by category, the downstream parser assigns it by position. Same bytes, two notions of meaning, only one of them checked.

for a senior

Show that you separate three claims that get blurred in an incident review — the text scored low, the text was harmless, the workflow behaved correctly. Only the first is supported by the screen's log.

for a principal

Be able to say why a screen chosen for reputational risk gets relied on for integrity risk, and what that misattribution costs when someone reports coverage upward.

## The claim being corrected When a workflow does something wrong on the back of a model's answer, the first defence offered is almost always "but the output was screened". That sentence is true and irrelevant. It is worth being precise about what an output screen is and what it therefore *can* conclude. An output harm classifier is a model — usually a small one — that takes a piece of text and emits scores over a **fixed set of harm categories**: harassment, hate, self-harm, sexual content, violence, sometimes a policy label or two. It is trained on the *semantics of prose*. A verdict of "pass" means one thing only: **this text scored below threshold on the categories this classifier has**. It is not a statement that the text is safe, and it is certainly not a statement about anything that happens to the text afterwards. ## Two components, two different notions of meaning An LLM answer that leaves a model does not go into a void. In an unattended triage and routing workflow it goes into a machine consumer: a parser that lifts values out of the answer so the workflow can act — a queue, a priority, an owner. That consumer has a **grammar**. It assigns meaning by *position and marker*: this is where the queue name lives, this token begins a field, this line is a directive rather than commentary. So the same bytes are read twice, by two components that disagree about what "meaning" is: | Component | What it reads the answer as | What it can conclude | |---|---|---| | Output harm classifier | prose, scored on harm categories | this text is/isn't abusive, violent, explicit | | Downstream parser | a structure, meaning assigned by position | these are the field values to act on | | Human on the morning digest | a list of outcomes | a ticket went to a queue at a priority | Nothing in that first row has an opinion about the second. The classifier has never been asked "does a consumer of this string give part of it structural force?", because that is not a property of the text at all — it is a property of the *pair* (text, consumer). Change the consumer and the same string is inert. ## Why this class exists in LLM workflows specifically The general shape is old: data crossing into something that interprets it. What is new is where the untrusted text comes from and how it is checked. The producer here is a **probabilistic model**, so the workflow cannot assume the answer's shape; and the check placed in front of it is a **semantic** one, chosen because the fear was that the assistant would say something offensive. The screen was sized for a reputational risk and then relied on for an integrity one. The attacker's side of it is correspondingly cheap. Content that must clear a harm classifier is a hard problem; content that must merely *be reproduced* by the model and then land in the consumer's grammar is not, because the harm classifier is looking for the wrong thing entirely. The attacker does not need the model to say anything it would refuse. They need it to say something ordinary, in a place where ordinary is load-bearing. ## What follows for reading a finding Three statements that look alike and are not: - *The answer passed the screen* — the text scored below threshold on the categories the screen has. - *The answer was harmless* — false as stated; harmlessness is not a property of a string alone. - *The workflow behaved correctly* — a separate claim entirely, and the one that actually failed. The practical consequence for someone triaging this class is that the classifier's log is not evidence either way, and the reviewer skimming a digest of routed tickets the next morning sees only an outcome that looks like an ordinary triage decision. Both of the things people point at as coverage — a screen and a human — were looking at properties that the construction never touches. That is the whole finding, and it is why "we screen the output" is a description of a control, not a rebuttal.

  • If the string is only dangerous to its consumer, is it fair to call the model's answer the problem at all?
    The answer is where the finding is observed, not where the property lives. The same string handed to a component that reads it as prose does nothing. That is why the class is defined by the pair — a producer that cannot be constrained to a shape and a consumer that assigns meaning by position — and why arguing about the model's wording alone tends to go nowhere.
  • Does adding a second, stricter harm classifier change anything here?
    No. Stacking classifiers raises sensitivity on the axes they already score — abuse, violence, self-harm. None of them gains a notion of what a parser downstream does with a field-shaped span. Two screens measuring the wrong property still measure the wrong property; you have bought latency and false positives on prose, not coverage of this class.
  • How would you state this class in an OWASP LLM Top 10 filing?
    Improper Output Handling is the item that fits: the defect is that a consumer treats model output as trusted structure. People often file it as Prompt Injection because a ticket body was the origin, but injection describes how the span got into the answer, not why the workflow acted on it. Where the numbering is uncertain, name the item rather than the identifier.

A metal detector at the door tells you the visitor is not carrying a weapon. It has no view on whether the envelope in their hand is a signature the mailroom will act on.

saying these in an interview costs you the question

  • Says a passing classifier means the answer was safe
  • Treats harmfulness as a property of a string alone
  • Assumes a second classifier would have caught it
  • Confuses the model saying something bad with a consumer acting on it

context

open as a page

Why can a post-generation output screen only truncate a streamed answer an attacker front-loaded?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Streaming sends tokens to the client as they are produced, so a screen that scores the finished answer returns its verdict after the opening has already been delivered. It can stop the remainder; it cannot recall what shipped.

open as a page

An attacker accepts the model's refusal and reads its displayed reasoning — what did the answer-scoped screen never score?

level: juniorimportance: must knowfreq 66%

basics

~20 s

A refusal describes the final answer and nothing else. A screen pointed at the answer never reads the deliberation panel or the trace record beside it, so refused substance and the model's stated objection sit outside its reach.

open as a page

How does an attacker locate a moderation screen's block threshold in a finance assistant that never shows scores?

level: juniorimportance: must knowfreq 68%

basics

~20 s

The score is hidden; the reply is not. A canned block card, a hedged answer and a normal answer are three visible outcomes, so every reply says which band the request landed in, and a few replies bracket the line.

open as a page

How can a request pass a term-based input screen while the generator still answers the original ask?

level: juniorimportance: must knowfreq 72%

basics

~20 s

The screen and the generator read the same string for different purposes. A term-based screen matches surface wording it was trained on; a capable generator resolves paraphrase, referents and framing, so meaning survives a restatement that contains no trained term.

open as a page

As you probe a chat product's input screen, what distinguishes its block from the model's own refusal?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A screen's block is a separate stage firing before generation: it returns fast, in fixed wording, with nothing streamed. A model refusal is generated text, so it varies in wording, references the actual ask, and arrives at generation speed.

open as a page

Why can an attacker's span address a generative screening model but not a label-only classifier?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A generative screening model reads the span it judges inside its own prompt, so written text reaches it as readable instruction. A label-only classifier emits category scores from a fixed head, so there is nothing to address.

open as a page

Why doesn't a semantic input classifier stop a request restated without its trained vocabulary?

level: middleimportance: must knowfreq 61%

basics

~20 s

Semantic means semantic within its training distribution. An embedding-based screen generalises over paraphrases near what it was labelled on, but a much larger generator resolves referents and framing the screen was never shown. The gap is capability, not keywords.

open as a page

An input screening model's weights and threshold are private: why can an attacker still map its coverage?

level: middleimportance: must knowfreq 64%

basics

~20 s

Because the deployed screen responds to every request. Each block, pass and near-miss is a labelled sample of its decision boundary, so a few dozen typed turns buy a usable map without seeing the model or its threshold.

open as a page

In a ticket-triage workflow, what must hold for text in a filed ticket to reach the parser reading the assistant's answer?

level: middleimportance: should knowfreq 45%

basics

~20 s

Three conditions. The model must reproduce the filed span in its answer rather than paraphrase it away; the span must land where the parser's grammar gives it meaning; and nothing between the model and the parser may rewrite or re-encode the answer.

open as a page

What does a streamed answer cut off mid-sentence tell an attacker that a flat refusal does not?

level: middleimportance: should knowfreq 44%

basics

~20 s

A truncation says the model was willing to produce the content and something downstream stopped delivery afterwards. A refusal says the model declined. The cut also shows roughly how much of an answer ships before a verdict lands.

open as a page

Which property of an attacker's request drives refused work into a displayed deliberation before the decline?

level: middleimportance: should knowfreq 44%

basics

~20 s

A request whose disposition cannot be settled without working the substance through. If judging it requires enumerating, comparing or weighing the material first, the material is written into the deliberation and the decline arrives after it, in a different channel.

open as a page

What bounds the precision of an attacker probing a moderation screen's hidden bands through a product's replies?

level: middleimportance: should knowfreq 46%

basics

~20 s

Four things: the outcomes are coarse, so you learn an interval and not a number; near-identical inputs need not score identically; probes are attributable and rate-limited; and the calibration expires whenever the screen or its configuration changes.

open as a page

Why is a generative screen's 'judge only what is inside the markers' rule not a boundary for the text it wraps?

level: middleimportance: should knowfreq 47%

basics

~20 s

Markers, the screen's orders and the wrapped item are one token stream to the screening model. Confining attention to the marked region is a preference expressed in the same text it is meant to fence.

open as a page

In a triage workflow, what does a system-prompt rule telling the model never to emit parser syntax actually buy?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A lower rate, not a closed class. The rule acts on a probabilistic producer, while the property that turns prose into a decision lives in the consumer, which is unchanged. Ship it as mitigation, not as a fix.

open as a page

Triaging a streamed-answer leak, the stored transcript ends at the cut — why doesn't that bound the disclosure?

level: seniorimportance: should knowfreq 38%

basics

~20 s

The stored transcript is written after screening, so it records the post-verdict version, not what was forwarded. Bound the disclosure from transport counters and from client-side state, and treat unobserved client residue as unknown rather than zero.

open as a page

A declined request's reasoning panel held the substance; the owner closed the finding as a refusal — what reopens it?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Show which artefact the screen received against every surface the request emitted, plus who can read each surface and for how long. The reopening argument is a scope defect with a measured success rate, not a claim that the model misbehaved.

open as a page

What does an attacker give up by keeping every request under a moderation screen's action point?

level: seniorimportance: should knowfreq 37%

basics

~20 s

Yield per request. Everything obtained under the line arrives hedged, shortened or generic, so value has to be assembled from many weak exchanges, and asks whose only useful form scores above the line stay out of reach entirely.

open as a page

What caps a probing campaign against a chat product's input screen when free accounts are rate-limited and suspended?

level: seniorimportance: should knowfreq 41%

basics

~20 s

The account roster and the clock, not cleverness. Every probe spends a rate-limited request, every suspension burns an account, and the repeats needed to beat sampling noise multiply both, all while the map decays underneath the campaign.

open as a page

What does a moderation record's 'clean' verdict prove about a listing written to be read by the screen?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Only that the screening stage emitted a clean verdict for that span, at that version, on that pass. It is not evidence the content was clean, that anyone read it, or that a re-screen would agree.

open as a page

A vendor screens with a small classifier and generates with a frontier model: defect or design limit?

level: principalimportance: should knowfreq 38%

basics

~20 s

It depends on what the vendor has claimed the screening layer does: against a promise of prevention it is a defect, against a promise of volume reduction a design limit. Either label needs a named owner.

open as a page

A triage record shows a wrong routing and an output screen verdict of pass — what does that prove?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Only that the answer scored below threshold on the screen's harm categories and that the parser lifted certain values from it. It does not show who supplied those values, and it does not show that the answer's content was benign or malicious.

open as a page

Does upgrading the generator behind an unchanged safety screen widen or close the evasion gap?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

It widens it. The leverage in a vocabulary-free restatement is the comprehension difference between the two models, so every indirection the new generator resolves that the old one could not is fresh working surface while the screen's coverage stands still.

open as a page

A listing written at a generative screen passed once in five re-runs: is that a finding?

level: seniorimportance: nice to knowfreq 29%

basics

~20 s

Yes, if you can say which stage passed it and that the rate beats the stage's own noise. One pass in five proves the construction worked against one deployment at one version — and a retryable path makes a fifth material.

open as a page

A programme lead wants front-loaded streaming cases retired now that screening runs incrementally — what do you say?

level: principalimportance: nice to knowfreq 26%

basics

~10 s

Incremental screening shrinks how much of an answer is delivered before a verdict; it does not reach zero, and it says nothing about surfaces nobody changed. Retire the case only against a measurement.

open as a page

You inherit findings harvested from reasoning traces after that surface changed shape — what can you still claim?

level: principalimportance: nice to knowfreq 27%

basics

~20 s

Separate claims about a past deployment from claims about the class. Instance findings are now unverifiable and must be retired or re-stated as history; the scope question survives and gets re-asked of the current shape. Non-reproduction is not evidence of a fix.

open as a page

Is a durable band of replies just under a moderation screen's block line a finding, or an accepted design limit?

level: principalimportance: nice to knowfreq 27%

basics

~20 s

Every scored gate with one action point has a band beneath it, so the band alone is arithmetic, not a defect. It is a finding only if the output obtainable inside it is materially outside what the product intends.

open as a page

Your red-team report maps where a chat product's input screen is thin: what do you claim it is worth?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Claim the standing property and its price, not the list. The thin regions expire at the next update to the screening model or its acting point; what survives is that a responding screen discloses its coverage at a measured cost.

open as a page