skip to content

Content Moderation

Where moderation sits in a request path, screening input before generation and output after, and how policy-conditioned classifiers replaced fixed harm taxonomies. It is the easiest layer to misplace.

on this pageshow

questions

4

In an LLM feature, why moderate both user input and model output?

level: middleimportance: must knowfreq 68%

answer

  1. two checkpoints, not one gate
  2. cheap rejection before the generation
  3. harm can arrive through retrieval
  4. output check sees what input never did

basics

~20 s

Input screening rejects harmful requests before you pay for a generation and before they enter the context. Output screening catches harm the model produces anyway, including harm that arrived through retrieved documents or tool results the input check never saw.

solid answer

~50 s

Moderation is two checkpoints around one generation, not a single gate. The **input check** runs before the model is called: it rejects abusive or clearly out-of-policy requests cheaply, keeps them out of the context window, and records what was asked. The **output check** runs on what the model produced, and it is the only place that catches harm the request never contained — an unsafe completion from a benign prompt, or slurs quoted out of a retrieved document or tool result that nobody screened. In a retrieval or agent system, most tokens in context did not come from the user, so an input-only pipeline inspects a minority of the material. The two sides also need different handling: a flagged prompt is a user-behaviour signal, while a flagged completion is a product defect you must not ship.

go deeper

for a junior

Be able to say that moderation happens in two places — on what the user sends and on what the model returns — and that checking only one of them leaves an obvious gap.

for a middle

Explain the mechanics: input screening saves the generation cost and keeps material out of context, while output screening is the only check that sees harm introduced by retrieval, tools, or the model itself.

for a senior

Show production judgment about placement in a real pipeline — every tool result and retrieved document is a boundary, streaming forces a real latency-versus-safety tradeoff, and a flagged input and a flagged output demand different responses.

for a principal

Own the framing that moderation is a probabilistic filter rather than a security boundary, and be clear about which risks it genuinely reduces versus which ones need architectural containment, permissions, and isolation instead.

## The shape of the problem A content-moderation layer answers one question about a piece of content: does it violate the policy this product operates under? Around a generative feature there are two natural places to ask it — before the model runs, on what came in, and after the model runs, on what came out. Systems that place it in only one of those spots are the most common moderation defect, and which spot they chose tells you which failures they will have. ## What the input check buys you Screening the incoming request happens before any expensive work. Three distinct benefits follow. **Cost and latency.** A request that will be refused costs one cheap classifier call instead of a full generation, and on a high-volume product that difference is the entire moderation budget. Hosted moderation endpoints are deliberately cheap or free for exactly this reason. **Keeping the material out of context.** Content you never put in the prompt cannot influence the model, cannot be echoed back, and cannot end up in a transcript you later have to retain, log or hand to a support agent. **A behavioural record.** Repeated flagged requests from one account are a signal about the account, not about the model. That signal only exists if you scored the input. ## What the output check buys you The output check exists because the input check cannot see the future. Three classes of harm are invisible on the way in. **Benign prompt, harmful completion.** "Write a scene where the villain intimidates the shopkeeper" is an ordinary request; the completion may not be ordinary. Nothing in the input scores as a violation. **Harm injected by the retrieval or tool layer.** In a retrieval-augmented or agentic system, most tokens in the context window are not the user's. A retrieved support ticket, a scraped page, a tool result, a memory entry written in an earlier session — none of it passed a user-facing input check, and the model will happily quote it. Screening the assistant's final message is the one place that content is inspected. **Composition.** Individually innocuous facts assembled into an actionable whole read as a violation only in the finished text. ## Both, and they are not symmetric Running both is the default answer, but the two checkpoints are configured and handled differently. A flagged input is about a person: you decline the request, and you may rate-limit, warn or escalate the account. A flagged output is about your system: you must not emit it, and a repeated pattern is a bug to investigate, not a user to sanction. The categories that matter can differ too — self-harm signals on the input side are a duty-of-care routing decision, while the same category on the output side is a hard block. Output checking also has a mechanical cost the input side does not: it lands on the critical path after generation, adding latency to every response. Streaming makes this sharper, because tokens have already reached the user by the time a complete message can be scored. The usual answers are to score in windows as the stream progresses and retract, to buffer the response until it is scored and lose the streaming benefit, or to stream only for low-risk request classes. There is no free option; the choice is an explicit product tradeoff. ## What moderation is not Moderation classifiers are probabilistic filters, not security boundaries. As of 2026 the published consensus, after adaptive-attack studies broke a long list of detection-based defenses, is that a determined adversary will get past a classifier eventually. Moderation raises the cost of casual abuse and catches the overwhelmingly common accidental case; it is not what you rely on to contain an attacker with access to your tools and data — that is an architectural problem solved with permissions and isolation, not with a score. It is also distinct from the model's own trained refusals. A provider model may decline a request on its own, but you do not control that behaviour, it varies by model and version, and it encodes the vendor's policy rather than yours. Your moderation layer is where your policy — which may be stricter, looser or simply different in scope — is actually enforced. ## The failure to avoid The classic misplacement is a retrieval product that screens the user's question, ships the answer unscreened, and is surprised when the assistant reproduces abusive content from an indexed document. The second classic is an agent that screens the first user turn and nothing else, while ten tool results and three model turns flow through unexamined. Moderation belongs at every boundary where untrusted text enters or leaves the system, and in a multi-turn agent that is more than two places.

  • In an agent that makes ten tool calls per task, where exactly do the moderation checks go?
    At every boundary where untrusted text enters or leaves. That means the user's turn, each tool result before it is appended to context, and the assistant's final message to the user. Screening only the first user turn leaves the majority of the context unexamined, since tool output and retrieved documents are untrusted content the user never wrote.
  • How do you moderate a streamed response without destroying the streaming experience?
    Three options, all with costs. Score rolling windows as tokens emit and retract or halt the stream on a flag — fast, but users may glimpse flagged text. Buffer the full message, score it, then release — safe, but it forfeits time-to-first-token. Or stream only for request classes you judge low-risk and buffer the rest. Pick per surface; there is no free choice.
  • If the provider's model already refuses harmful requests, why add your own moderation layer?
    Because those refusals encode the vendor's policy, not yours, and they shift between model versions without notice. Your policy may be stricter, may cover categories the vendor ignores, and must be auditable. You also get no score, no category and no record from a refusal — nothing to log, escalate or appeal against.

saying these in an interview costs you the question

  • Says the provider model already refuses, so no moderation is needed
  • Screens the user prompt only and never the generated output
  • Forgets retrieved documents and tool results are unscreened untrusted content
  • Treats a moderation classifier as a security boundary against attackers
  • Applies identical handling to a flagged prompt and a flagged completion

context

open as a page

In multimodal moderation, how do you catch harm that lives only in the image?

level: middleimportance: should knowfreq 38%

basics

~20 s

Score the image, not just the words around it. Text-only classifiers pass a clean caption over an unsafe picture, so the pipeline needs image-capable moderation — either a joint text-plus-image call or a dedicated image classifier — plus sampled frames for video rather than one thumbnail.

open as a page

When does a policy-conditioned moderation classifier beat a fixed hazard taxonomy?

level: seniorimportance: should knowfreq 45%

basics

~20 s

When your policy is specific to your product and changes often. A policy-conditioned classifier reads your written policy at inference time, so a rule change is a text edit rather than a labelling-and-retraining cycle. Fixed taxonomies stay cheaper and faster for stable, universal harm categories.

open as a page

How do you design human escalation and appeals into a moderation pipeline?

level: principalimportance: should knowfreq 30%

basics

~20 s

Treat human review as a component of the system, not a safety net bolted on. Every automated decision records the policy clause and model version behind it, affected users are told which rule they broke, appeals go to a reviewer with that evidence, and reversals feed back into the policy.

open as a page