In an LLM app, why screen model output when the user input already passed moderation?
answer
- two checkpoints, not one
- the prompt is not the only source
- retrieved and relayed text bypasses input
- output check versus streaming
- different thresholds per side
basics
~20 sInput screening only sees what the user typed. A model can still emit harmful text from a benign prompt, from retrieved documents, or from another user's content pulled into context, so the output is a separate surface that needs its own classifier.
solid answer
~50 sInput and output classifiers catch different failures. An input classifier scores what the user sent: abusive prompts, requests for disallowed content, self-harm disclosures you want to route to a crisis flow. It also saves you from paying to generate a response you would discard. But most harmful output is not traceable to harmful input. A model can drift in roleplay, sample something ugly, or faithfully repeat abusive text that entered its context from a retrieved document or another user's message — for example a game helpdesk assistant that quotes player-submitted bug reports back into a reply. Only an output classifier sees that. The cost is latency and placement: the input check delays the start of generation, while the output check either forces you to buffer the whole reply before showing it or to moderate streamed chunks and accept that some text may need retracting.
go deeper
Know that moderation happens on both what the user sends and what the model produces, and be able to say why one check is not enough.
Explain a concrete path where a clean prompt yields a harmful reply, and describe where each check sits in the request path and what latency it adds.
Show you have operated this: chunked versus buffered output moderation under streaming, asymmetric thresholds per side, and logging that lets you measure what got through, not just what got blocked.
Own the policy question of which surfaces may show text before it is fully classified, and defend that decision against both the latency cost of buffering and the reputational cost of a retraction.
## Two places to put a classifier A moderation classifier is a model that scores a piece of text against a set of harm categories and returns per-category scores. In an LLM application there are exactly two natural places to run it: on the **input**, before the user's text reaches the model, and on the **output**, before the model's text reaches a person, a UI, or a downstream system. They are not redundant. They catch different things, they cost latency in different places, and mature systems run both with independently tuned thresholds. ## What input screening catches The input classifier sees only what the user submitted. That is enough for a real set of cases: overtly abusive messages, explicit requests for disallowed content, and disclosures — a message mentioning self-harm often needs a crisis-resources flow rather than a generation at all. Input screening also has an economic benefit: you reject before you spend generation tokens and before an agent takes any action, so a blocked request is cheap. Its blind spot is everything the user did not type. ## What only output screening catches Three families of harmful output survive a clean input check: 1. **Model-originated harm from a benign prompt.** Ask for a punchy in-game taunt and the model may produce a slur. Nothing in the prompt was flaggable. 2. **Harm that entered the context from elsewhere.** Retrieved documents, uploaded files, tool results, and other users' text all enter the context window without passing your input classifier. An assistant that summarises player-submitted bug reports will happily reproduce abuse pasted into one of them. 3. **Benign-looking prompts that produce harmful completions.** Framing, roleplay, and gradual escalation can all yield a harmful reply from a request that scores near zero on every input category. The model's own trained refusal behaviour is a separate control from the classifier: refusals are a property of the model, classifiers are a property of your system, and one should not be treated as a substitute for the other. If the model refuses, the output classifier simply sees a refusal and passes it. ## The latency each side adds This is where interviewers press. The input check sits on the critical path *before* generation, so it adds its full latency to time-to-first-token. It is usually a small, fast classifier, so this is tens of milliseconds — but it is serial. Some systems hide it by starting generation optimistically in parallel and cancelling if the input verdict comes back positive; that trades wasted tokens for latency. The output check is harder. If you buffer the whole response and classify it once, you have destroyed streaming: the user stares at a spinner for the full generation. The common compromise is to moderate in chunks — classify each flushed chunk with a small lookahead buffer, and on a hit stop the stream and replace what was shown with a safe message. That means harmful text can briefly appear before being retracted, which is a product decision, not a technical accident. For high-stakes surfaces, buffering fully is the right call and the latency is the price. ## Asymmetric strictness The two sides usually get different thresholds, and this surprises people. Input screening should often be *lenient* on categories like self-harm, because you want to detect and route those messages, not block a user reaching out. Output screening is usually *stricter*, because the output is your product speaking. Treating both sides with one global cutoff is a common design smell. ## Failure modes worth naming - Screening only the input and calling it moderated. - Screening only the final assistant message while intermediate content the model produced — arguments it passes onward, text it writes into a shared surface — goes unchecked. - Buffering the entire stream on every request when only a small fraction of surfaces need it, then blaming the model for slow perceived latency. - Logging only blocks, so you can never measure what got through. ## What interviewers listen for They want to hear that you know moderation is a pipeline with two independent stages, that you can name a concrete harmful-output path that starts with a harmless prompt, and that you understand the streaming conflict on the output side rather than hand-waving "we moderate everything".
- How do you moderate an output that is being streamed token by token to the user?Either buffer the whole response and classify once, which removes streaming's latency benefit, or classify each flushed chunk with a small lookahead buffer and stop the stream on a hit. Chunked moderation means some harmful text can appear briefly before being retracted, so high-stakes surfaces usually buffer instead. The choice is a product decision about which is worse: a spinner or a retraction.
- Why would you set a more lenient input threshold than output threshold for the self-harm category?Because on the input side a self-harm signal is something you want to detect and act on, not suppress. Blocking the message hides a user who may need crisis resources routed to them. On the output side the same category means the system is about to say something harmful in your product's voice, which warrants a much stricter cutoff.
- An agent's moderation only checks the final assistant message. What does that miss?Everything the model emits that is not the final message: text it writes into shared surfaces, content it passes to other components, and intermediate turns a user can see. Screening only the last turn assumes the model's sole output channel is the reply, which stops being true the moment it can act rather than just answer.
saying these in an interview costs you the question
- Says input moderation alone is sufficient
- Assumes harmful output requires harmful input
- Ignores that retrieved or relayed text never passed the input check
- Treats the model's own refusals as a substitute for a classifier
- Claims output moderation is free of latency cost