When does a policy-conditioned moderation classifier beat a fixed hazard taxonomy?
answer
- policy as input, not weights
- taxonomy is fixed at training time
- edit the document, skip the retraining
- rationale and policy cost tokens
- cheap first pass, expensive second pass
basics
~20 sWhen your policy is specific to your product and changes often. A policy-conditioned classifier reads your written policy at inference time, so a rule change is a text edit rather than a labelling-and-retraining cycle. Fixed taxonomies stay cheaper and faster for stable, universal harm categories.
solid answer
~50 sA fixed-taxonomy classifier — Llama Guard 4's hazard categories, ShieldGemma 2 for images, or a hosted endpoint like `omni-moderation-latest` — has its categories baked in at training time. That is a good fit for universal harms: violence, sexual content, self-harm, hate. It is a bad fit for the rules that are actually specific to your product, such as what counts as unlabelled financial advice or a prohibited resale listing. **Policy-conditioned** classifiers such as `gpt-oss-safeguard` invert the relationship: you pass your written policy alongside the content, and the model reasons against it and returns a decision plus a rationale. Changing a rule becomes editing a document and re-running your eval set, not gathering labels and retraining. The costs are real — the policy text and the rationale add tokens and latency per item, and the policy becomes a versioned artifact you must test. Most production systems run both: a cheap fixed pass at volume, a policy-conditioned pass on the borderline and appealed slice.
go deeper
Know that moderation classifiers come in two shapes — ones with a built-in category list, and ones you hand your own written policy to at request time — and that the second kind is newer.
Explain the mechanics of the swap: a fixed taxonomy encodes the policy in trained weights, while a policy-conditioned model reads it from the prompt, so one changes by retraining and the other by editing text.
Demonstrate the production judgment about where each belongs — cheap fixed classifier at volume, policy-conditioned reasoning on the borderline and appealed slice — and be honest about the token, latency, and evaluation costs you take on.
Own the organizational consequence: the policy document becomes an engineering artifact with version control, review, and an eval gate, which changes who writes policy and how fast trust-and-safety can ship a rule.
## Two ways to encode a policy Every moderation classifier encodes a policy somewhere. The design question is *where*. In a **fixed-taxonomy** classifier the policy lives in the weights: someone chose a category list, gathered labelled examples for each, and trained. In a **policy-conditioned** classifier the policy lives in the prompt: the model receives your written rules together with the content and reasons about whether one violates the other. ## The fixed-taxonomy world This is what moderation looked like for most of the field's history and it is still the volume workhorse. Open models such as Llama Guard 4 ship a hazard taxonomy (its S-numbered categories cover things like violent crimes, sexual content, self-harm, and specialised advice); ShieldGemma 2 does the same job for images. Hosted endpoints follow the pattern: `omni-moderation-latest` returns per-category scores across a fixed category set for text and images, and Azure AI Content Safety returns severity levels across its four categories of hate, sexual, violence and self-harm. The strengths are exactly what you would expect from a small purpose-built classifier: low latency, low cost, one forward pass, predictable output shape, easy to run at millions of items a day. For harms that are genuinely universal — the ones every platform prohibits and every regulator names — this is the right tool and there is no reason to reach past it. The weakness is that the taxonomy is a fixed vocabulary. Your product's actual policy document is not a list of four to fourteen universal harms; it is pages of specific rules about *your* domain. A trading community prohibits unlabelled financial advice. A marketplace prohibits certain resale categories. A creator platform has a nuanced line between edgy and prohibited that took its trust-and-safety team years to write. None of that maps onto a hazard category. The usual workaround — pick the nearest category and hope, or bolt on keyword rules — degrades quickly, and a genuine change to your policy means collecting labels and retraining, a cycle measured in weeks. ## The policy-conditioned inversion Policy-conditioned safeguard models take the policy as *input*. You supply your rules — a several-hundred-word document, with definitions, inclusions and explicit exclusions — plus the item to classify, and the model returns a verdict and a written rationale citing the reasoning. `gpt-oss-safeguard`, released in open weights under Apache-2.0 in two sizes, is the reference example of the pattern. Three consequences follow. **Iteration speed collapses.** A policy revision ships as a text edit. You change a clause, re-run your labelled eval set, and see the effect immediately. Compared with a labelling-and-retraining loop this is the difference between an afternoon and a quarter. **You can hold many policies at once.** One deployment serves per-surface or per-jurisdiction policies by swapping the document, rather than maintaining a fleet of separately trained models. **The rationale is auditable.** A verdict with reasoning that points at the clause it applied is something a human reviewer, an appeals process, or a regulator can actually inspect. A bare score is not. ## What it costs The tradeoffs are not subtle. Every item now carries the policy text plus generated reasoning tokens, so per-item cost and latency are far above a small classifier's single pass — which is why this pattern rarely sits inline on the highest-volume path. The written policy becomes a first-class engineering artifact: it needs version control, review, and an eval set that runs on every change, because an ambiguous clause is now a production bug. Nuanced natural-language policies also give the model room to interpret, which cuts both ways — it generalises to cases you never enumerated, and it can drift on the ones you thought were settled. And a policy-conditioned model is still just a classifier: a determined adversary rewriting content to evade it will succeed at some rate, exactly as with the fixed kind. ## The layered answer Most production systems do not choose. A cheap fixed-taxonomy or hosted classifier runs on everything and disposes of the clear cases at volume. The policy-conditioned model runs on the narrow slice that matters: items near the decision boundary, items in categories where your policy is genuinely product-specific, appeals, and audit sampling. That places the expensive reasoning exactly where nuance is worth paying for, and it gives your human reviewers a written rationale to work from on the items they see. ## Currency This is a live area and the landscape moves. Legacy text-only moderation endpoints have been retired in favour of multimodal ones; at least one major inference provider has swapped its default open safety model from a fixed-taxonomy Llama Guard build to a policy-conditioned safeguard model; and Google's long-standing Perspective API is being sunset, with service ending at the end of 2026 and no migration path offered. Naming the direction of travel — from fixed taxonomies toward policy-conditioned reasoning classifiers, with the fixed models surviving as the cheap first pass — matters more in an interview than naming any one model.
- If the policy is now a prompt, how do you stop a policy edit from silently regressing moderation quality?Treat the policy as code. It lives in version control with review, and every change runs against a labelled eval set covering each clause plus the borderline cases that motivated past edits. Track per-clause agreement rather than a single aggregate, because a reworded clause typically breaks one category while the overall number barely moves.
- Why is a written rationale from the classifier worth the extra tokens?Because it turns a verdict into something reviewable. A human moderator handling an appeal sees which clause was applied and why, rather than re-deciding from scratch. It also makes systematic errors legible — if fifty reversals all cite the same ambiguous clause, you have found the bug in the policy rather than in the model.
- Where does a fixed-taxonomy classifier still clearly win in 2026?On the high-volume path for universal harms. When you are scoring millions of items a day against categories that are stable and near-universal — sexual content, graphic violence, self-harm — a small single-pass classifier is an order of magnitude cheaper and faster, and the taxonomy never needed to change anyway. Reserve the reasoning model for what is genuinely yours.
saying these in an interview costs you the question
- Assumes every moderation policy change requires retraining a classifier
- Thinks a policy-conditioned model costs the same as a small classifier
- Stretches a fixed hazard taxonomy to cover product-specific rules
- Ignores that the written policy becomes a versioned, tested artifact
- Believes reasoning over a policy makes the classifier evasion-proof