skip to content

Why do content moderation classifiers return per-category scores, not one flag?

level: juniorimportance: must knowfreq 55%

answer

  1. one flag, one action — too blunt
  2. categories exist to differ in response
  3. cost of a miss varies by harm
  4. self-harm needs routing, not blocking
  5. thresholds live per category

basics

~20 s

Different harms need different responses. Per-category scores let a system hard-block one category, route self-harm signals to a crisis flow, and only warn on mild insults. A single unsafe flag forces one action and one cutoff for every kind of harm.

solid answer

~50 s

A harm taxonomy is the list of categories a moderation system scores against — things like hate, harassment, sexual content, violence, and self-harm. Classifiers return a score per category rather than one boolean because the product's response differs by category. In a competitive shooter's party chat, a message scoring high on self-harm should page a human within minutes and surface crisis resources; a message scoring high on trash talk about the match should probably do nothing at all. One flag collapses those into a single action. Per-category scores also let you set a different threshold per category, which matters because the cost of a false positive is wildly uneven: wrongly muting a player for banter is annoying, wrongly missing a credible threat is a safety incident. The design test for any category is simple — if two harms trigger the same action at the same threshold, they do not need to be separate categories.

go deeper

for a junior

Be able to name a few harm categories and say plainly that different categories trigger different actions, so a single unsafe flag is not enough.

for a middle

Explain the response test for whether a category should exist, and give a concrete pair of harms where the correct actions differ.

for a senior

Show how per-category thresholds are set from the cost of each error type, and how you handle overlap precedence and dead categories in a live taxonomy.

for a principal

Own the taxonomy as a product artifact: who approves changes, how new categories are introduced without invalidating historical metrics, and how enforcement policy is documented and defended externally.

## What a harm taxonomy is A harm taxonomy is the enumerated set of categories your moderation system recognises: hate, harassment, sexual content, violence, self-harm, and so on, often with subcategories. A classifier scores a piece of text against each category independently and returns a score per category, not a single verdict. The taxonomy is the contract between the classifier and the product: it defines what you can even talk about detecting. ## Why one flag is not enough A single `unsafe: true` flag forces three things you almost never want: 1. **One action for every harm.** Blocking is right for some categories and actively wrong for others. Suppressing a message that scores high on self-harm hides a user who may need help. 2. **One threshold for every harm.** Categories have wildly different base rates and wildly different error costs. A cutoff calibrated for mild profanity is far too loose for credible threats of violence. 3. **One review path.** Operationally you need to know *why* something was flagged in order to route it — to an automated mute, to a general review queue, or to an urgent queue with a tight response time. A boolean throws away all the information needed to make those decisions. ## Designing categories: the response test The useful design rule is that **a category earns its existence when it maps to a distinct response**. Not when it feels conceptually different, and not because a vendor's default list happens to contain it. Apply the test: if category A and category B always produce the same action at the same threshold, merge them. If one broad category is producing two different actions depending on what a reviewer sees inside it, split it. ## A worked example: party chat in a competitive shooter Consider voice-to-text party chat in a competitive online shooter. A flat "harassment" category performs badly here, because the same words carry opposite meanings. "You're trash, get off my team" between friends who queue together every night is banter. The identical string aimed at a random matchmade stranger is abuse. And "this map is trash, whoever designed it should be fired" is neither — it is trash talk about the match. Splitting harassment into **targeted-at-a-player** and **trash-talk-about-the-match** passes the response test, because the responses genuinely differ: the first can lead to a chat mute or a report escalation, while the second is left alone at any score. That split also lets you set the targeted subcategory's threshold low (catch more, tolerate some false positives) and effectively disable enforcement on the other, which a merged category cannot express. Note what the split does *not* solve: relationship context. The classifier sees text, not the fact that these two players are friends. Systems handle that with signals outside the classifier — party membership, prior reports, whether the recipient blocked the sender — combined with the score. Naming that limitation is what separates a thoughtful answer from a naive one. ## Per-category thresholds Once you have categories, each gets its own cutoff. Rough shape in practice: - **Categories where a miss is catastrophic** (child safety, credible threats): low threshold, accept false positives, human review on everything flagged. - **Categories where a false positive damages the product** (mild profanity, edgy humour in an adult context): high threshold, or no automated enforcement at all. - **Categories that need routing rather than blocking** (self-harm): moderate threshold, action is escalation and resource surfacing, not suppression. ## Where taxonomies go wrong - **Categories that never fire.** Dead categories add classifier cost and reviewer confusion. Measure per-category volume and prune. - **Overlapping categories with no precedence rule.** One message scoring high on three categories needs a defined winner, or your action is nondeterministic. - **Adopting a default taxonomy unchanged.** Off-the-shelf category lists are a starting point; the categories that matter to a shooter's chat, a health forum, and a legal-research tool are not the same set. - **Assuming categories are mutually exclusive.** They are not — scores are independent and a message can be high on several at once. ## What interviewers listen for They want the response test articulated, at least one concrete example where two harms demand different actions, and awareness that thresholds are per category rather than global. A candidate who says "we call the moderation check and block if it says unsafe" has described a system that will both over-block and under-protect.

  • When should you merge two harm categories rather than keep them separate?
    When they always produce the same action at the same threshold. If nothing downstream branches on the distinction, the split adds classifier surface, reviewer ambiguity and dashboard noise for no behavioural difference. The reverse also holds: split a category the moment reviewers start applying two different outcomes to items inside it.
  • A message scores high on three categories at once. How should the system decide what to do?
    With an explicit precedence rule, defined in advance, usually ordered by severity of the required response. Self-harm routing outranks a chat mute, and a hard-block category outranks a warning. Without a stated rule the action depends on evaluation order in code, which makes enforcement nondeterministic and impossible to audit.
  • Why can't a text classifier alone distinguish banter between friends from abuse toward a stranger?
    Because the distinguishing information is not in the text. Identical words carry opposite intent depending on the relationship between sender and recipient. Production systems combine the classifier score with out-of-band signals — party membership, mutual history, prior reports, whether the recipient blocked the sender — and reserve enforcement for the combination, not the score alone.

saying these in an interview costs you the question

  • Treats moderation as a single unsafe boolean
  • Uses one global threshold across all harm categories
  • Assumes blocking is the right response to every category
  • Adopts a default category list without checking it fits the product
  • Assumes a message belongs to exactly one category

context