An output harm classifier passed an answer a downstream parser then acted on — what did the screen measure?
answer
- the screen only has one question
- harm categories, not consequences
- prose in, structure out
- the consumer creates the effect
basics
~20 sIt measured harm categories in text: toxicity, violence, self-harm and the like, and found none. The answer was ordinary prose. Its effect came from the parser that gave part of it structural meaning — a property the classifier never scores.
solid answer
~50 sAn output harm classifier is a text scorer: it reads the generated answer and returns scores over a fixed set of harm categories. A string can be bland on every one of those axes and still be consequential, because the consequence is not in the meaning of the words — it is in what the next component does with them. In an unattended ticket-triage workflow the assistant's answer is handed to a parser that lifts routing fields out of it. The parser assigns meaning by position and marker; the classifier assigns meaning by category. So a span the classifier reads as unremarkable prose is read by the parser as structure, and the workflow writes a routing, priority or owner decision from it. "The output was screened" answers a question nobody asked: the screen scored the harmfulness of text, not the behaviour of its consumer.
go deeper
Recall what an output classifier actually returns: scores over a fixed list of harm categories. Be ready to say that a pass means the text scored low on those categories and nothing more.
Explain the two-readers picture: the classifier assigns meaning by category, the downstream parser assigns it by position. Same bytes, two notions of meaning, only one of them checked.
Show that you separate three claims that get blurred in an incident review — the text scored low, the text was harmless, the workflow behaved correctly. Only the first is supported by the screen's log.
Be able to say why a screen chosen for reputational risk gets relied on for integrity risk, and what that misattribution costs when someone reports coverage upward.
## The claim being corrected When a workflow does something wrong on the back of a model's answer, the first defence offered is almost always "but the output was screened". That sentence is true and irrelevant. It is worth being precise about what an output screen is and what it therefore *can* conclude. An output harm classifier is a model — usually a small one — that takes a piece of text and emits scores over a **fixed set of harm categories**: harassment, hate, self-harm, sexual content, violence, sometimes a policy label or two. It is trained on the *semantics of prose*. A verdict of "pass" means one thing only: **this text scored below threshold on the categories this classifier has**. It is not a statement that the text is safe, and it is certainly not a statement about anything that happens to the text afterwards. ## Two components, two different notions of meaning An LLM answer that leaves a model does not go into a void. In an unattended triage and routing workflow it goes into a machine consumer: a parser that lifts values out of the answer so the workflow can act — a queue, a priority, an owner. That consumer has a **grammar**. It assigns meaning by *position and marker*: this is where the queue name lives, this token begins a field, this line is a directive rather than commentary. So the same bytes are read twice, by two components that disagree about what "meaning" is: | Component | What it reads the answer as | What it can conclude | |---|---|---| | Output harm classifier | prose, scored on harm categories | this text is/isn't abusive, violent, explicit | | Downstream parser | a structure, meaning assigned by position | these are the field values to act on | | Human on the morning digest | a list of outcomes | a ticket went to a queue at a priority | Nothing in that first row has an opinion about the second. The classifier has never been asked "does a consumer of this string give part of it structural force?", because that is not a property of the text at all — it is a property of the *pair* (text, consumer). Change the consumer and the same string is inert. ## Why this class exists in LLM workflows specifically The general shape is old: data crossing into something that interprets it. What is new is where the untrusted text comes from and how it is checked. The producer here is a **probabilistic model**, so the workflow cannot assume the answer's shape; and the check placed in front of it is a **semantic** one, chosen because the fear was that the assistant would say something offensive. The screen was sized for a reputational risk and then relied on for an integrity one. The attacker's side of it is correspondingly cheap. Content that must clear a harm classifier is a hard problem; content that must merely *be reproduced* by the model and then land in the consumer's grammar is not, because the harm classifier is looking for the wrong thing entirely. The attacker does not need the model to say anything it would refuse. They need it to say something ordinary, in a place where ordinary is load-bearing. ## What follows for reading a finding Three statements that look alike and are not: - *The answer passed the screen* — the text scored below threshold on the categories the screen has. - *The answer was harmless* — false as stated; harmlessness is not a property of a string alone. - *The workflow behaved correctly* — a separate claim entirely, and the one that actually failed. The practical consequence for someone triaging this class is that the classifier's log is not evidence either way, and the reviewer skimming a digest of routed tickets the next morning sees only an outcome that looks like an ordinary triage decision. Both of the things people point at as coverage — a screen and a human — were looking at properties that the construction never touches. That is the whole finding, and it is why "we screen the output" is a description of a control, not a rebuttal.
- If the string is only dangerous to its consumer, is it fair to call the model's answer the problem at all?The answer is where the finding is observed, not where the property lives. The same string handed to a component that reads it as prose does nothing. That is why the class is defined by the pair — a producer that cannot be constrained to a shape and a consumer that assigns meaning by position — and why arguing about the model's wording alone tends to go nowhere.
- Does adding a second, stricter harm classifier change anything here?No. Stacking classifiers raises sensitivity on the axes they already score — abuse, violence, self-harm. None of them gains a notion of what a parser downstream does with a field-shaped span. Two screens measuring the wrong property still measure the wrong property; you have bought latency and false positives on prose, not coverage of this class.
- How would you state this class in an OWASP LLM Top 10 filing?Improper Output Handling is the item that fits: the defect is that a consumer treats model output as trusted structure. People often file it as Prompt Injection because a ticket body was the origin, but injection describes how the span got into the answer, not why the workflow acted on it. Where the numbering is uncertain, name the item rather than the identifier.
A metal detector at the door tells you the visitor is not carrying a weapon. It has no view on whether the envelope in their hand is a signature the mailroom will act on.
saying these in an interview costs you the question
- Says a passing classifier means the answer was safe
- Treats harmfulness as a property of a string alone
- Assumes a second classifier would have caught it
- Confuses the model saying something bad with a consumer acting on it