skip to content

Moderation & Content Classifiers

You learn to screen inputs and outputs with a harm taxonomy and a classifier, then tune the threshold. Every setting is a precision-recall bet: too tight and you refuse legitimate users, too loose and harmful content ships.

on this pageshow

questions

5

Why do content moderation classifiers return per-category scores, not one flag?

level: juniorimportance: must knowfreq 55%

answer

  1. one flag, one action — too blunt
  2. categories exist to differ in response
  3. cost of a miss varies by harm
  4. self-harm needs routing, not blocking
  5. thresholds live per category

basics

~20 s

Different harms need different responses. Per-category scores let a system hard-block one category, route self-harm signals to a crisis flow, and only warn on mild insults. A single unsafe flag forces one action and one cutoff for every kind of harm.

solid answer

~50 s

A harm taxonomy is the list of categories a moderation system scores against — things like hate, harassment, sexual content, violence, and self-harm. Classifiers return a score per category rather than one boolean because the product's response differs by category. In a competitive shooter's party chat, a message scoring high on self-harm should page a human within minutes and surface crisis resources; a message scoring high on trash talk about the match should probably do nothing at all. One flag collapses those into a single action. Per-category scores also let you set a different threshold per category, which matters because the cost of a false positive is wildly uneven: wrongly muting a player for banter is annoying, wrongly missing a credible threat is a safety incident. The design test for any category is simple — if two harms trigger the same action at the same threshold, they do not need to be separate categories.

go deeper

for a junior

Be able to name a few harm categories and say plainly that different categories trigger different actions, so a single unsafe flag is not enough.

for a middle

Explain the response test for whether a category should exist, and give a concrete pair of harms where the correct actions differ.

for a senior

Show how per-category thresholds are set from the cost of each error type, and how you handle overlap precedence and dead categories in a live taxonomy.

for a principal

Own the taxonomy as a product artifact: who approves changes, how new categories are introduced without invalidating historical metrics, and how enforcement policy is documented and defended externally.

## What a harm taxonomy is A harm taxonomy is the enumerated set of categories your moderation system recognises: hate, harassment, sexual content, violence, self-harm, and so on, often with subcategories. A classifier scores a piece of text against each category independently and returns a score per category, not a single verdict. The taxonomy is the contract between the classifier and the product: it defines what you can even talk about detecting. ## Why one flag is not enough A single `unsafe: true` flag forces three things you almost never want: 1. **One action for every harm.** Blocking is right for some categories and actively wrong for others. Suppressing a message that scores high on self-harm hides a user who may need help. 2. **One threshold for every harm.** Categories have wildly different base rates and wildly different error costs. A cutoff calibrated for mild profanity is far too loose for credible threats of violence. 3. **One review path.** Operationally you need to know *why* something was flagged in order to route it — to an automated mute, to a general review queue, or to an urgent queue with a tight response time. A boolean throws away all the information needed to make those decisions. ## Designing categories: the response test The useful design rule is that **a category earns its existence when it maps to a distinct response**. Not when it feels conceptually different, and not because a vendor's default list happens to contain it. Apply the test: if category A and category B always produce the same action at the same threshold, merge them. If one broad category is producing two different actions depending on what a reviewer sees inside it, split it. ## A worked example: party chat in a competitive shooter Consider voice-to-text party chat in a competitive online shooter. A flat "harassment" category performs badly here, because the same words carry opposite meanings. "You're trash, get off my team" between friends who queue together every night is banter. The identical string aimed at a random matchmade stranger is abuse. And "this map is trash, whoever designed it should be fired" is neither — it is trash talk about the match. Splitting harassment into **targeted-at-a-player** and **trash-talk-about-the-match** passes the response test, because the responses genuinely differ: the first can lead to a chat mute or a report escalation, while the second is left alone at any score. That split also lets you set the targeted subcategory's threshold low (catch more, tolerate some false positives) and effectively disable enforcement on the other, which a merged category cannot express. Note what the split does *not* solve: relationship context. The classifier sees text, not the fact that these two players are friends. Systems handle that with signals outside the classifier — party membership, prior reports, whether the recipient blocked the sender — combined with the score. Naming that limitation is what separates a thoughtful answer from a naive one. ## Per-category thresholds Once you have categories, each gets its own cutoff. Rough shape in practice: - **Categories where a miss is catastrophic** (child safety, credible threats): low threshold, accept false positives, human review on everything flagged. - **Categories where a false positive damages the product** (mild profanity, edgy humour in an adult context): high threshold, or no automated enforcement at all. - **Categories that need routing rather than blocking** (self-harm): moderate threshold, action is escalation and resource surfacing, not suppression. ## Where taxonomies go wrong - **Categories that never fire.** Dead categories add classifier cost and reviewer confusion. Measure per-category volume and prune. - **Overlapping categories with no precedence rule.** One message scoring high on three categories needs a defined winner, or your action is nondeterministic. - **Adopting a default taxonomy unchanged.** Off-the-shelf category lists are a starting point; the categories that matter to a shooter's chat, a health forum, and a legal-research tool are not the same set. - **Assuming categories are mutually exclusive.** They are not — scores are independent and a message can be high on several at once. ## What interviewers listen for They want the response test articulated, at least one concrete example where two harms demand different actions, and awareness that thresholds are per category rather than global. A candidate who says "we call the moderation check and block if it says unsafe" has described a system that will both over-block and under-protect.

  • When should you merge two harm categories rather than keep them separate?
    When they always produce the same action at the same threshold. If nothing downstream branches on the distinction, the split adds classifier surface, reviewer ambiguity and dashboard noise for no behavioural difference. The reverse also holds: split a category the moment reviewers start applying two different outcomes to items inside it.
  • A message scores high on three categories at once. How should the system decide what to do?
    With an explicit precedence rule, defined in advance, usually ordered by severity of the required response. Self-harm routing outranks a chat mute, and a hard-block category outranks a warning. Without a stated rule the action depends on evaluation order in code, which makes enforcement nondeterministic and impossible to audit.
  • Why can't a text classifier alone distinguish banter between friends from abuse toward a stranger?
    Because the distinguishing information is not in the text. Identical words carry opposite intent depending on the relationship between sender and recipient. Production systems combine the classifier score with out-of-band signals — party membership, mutual history, prior reports, whether the recipient blocked the sender — and reserve enforcement for the combination, not the score alone.

saying these in an interview costs you the question

  • Treats moderation as a single unsafe boolean
  • Uses one global threshold across all harm categories
  • Assumes blocking is the right response to every category
  • Adopts a default category list without checking it fits the product
  • Assumes a message belongs to exactly one category

context

open as a page

In an LLM app, why screen model output when the user input already passed moderation?

level: middleimportance: must knowfreq 62%

basics

~20 s

Input screening only sees what the user typed. A model can still emit harmful text from a benign prompt, from retrieved documents, or from another user's content pulled into context, so the output is a separate surface that needs its own classifier.

open as a page

How do you pick the score threshold for a content-moderation classifier?

level: seniorimportance: must knowfreq 68%

basics

~20 s

Not from a default. Score a labelled sample, plot precision against recall per category, and choose the operating point where the cost of a false positive and the cost of a missed harm balance for that category — then check the resulting flag volume fits your review capacity.

open as a page

If your moderation classifier times out, should the request fail open or fail closed?

level: seniorimportance: should knowfreq 48%

basics

~20 s

It depends on the context, and it must be written down in advance. Fail closed where unchecked content is unacceptable, fail open where a classifier outage taking down the product is worse than a rare miss — and log every bypassed request for retrospective review either way.

open as a page

How do you size and operate the human review queue behind a moderation classifier?

level: principalimportance: should knowfreq 38%

basics

~20 s

Derive daily escalation volume from traffic and the chosen thresholds, staff to that number with headroom for spikes, tier by urgency with a response-time target on the most time-critical categories, and sample allowed traffic too — flagged-only review can never reveal what you missed.

open as a page