skip to content

In a content-moderation labeling operation, why is each reported post judged by three annotators instead of one?

level: juniorimportance: must knowfreq 60%

answer

  1. one verdict has nothing to check it
  2. disagreement is the useful output
  3. contested items need an escalation path
  4. agreement tracks guideline health per category
  5. the adjudicator settles a genuine split

basics

~20 s

One verdict on a moderated post carries no error signal of its own. Three verdicts make disagreement visible, route the contested post into an adjudication path, and turn agreement into a running health check on the written guideline.

solid answer

~40 s

A single annotator's verdict is a bare assertion: nothing in the record separates a correct read of a clear-cut post from a careless read of a hard one. Sending each reported post to three independent annotators produces two things one pass cannot. First, an item-level disagreement flag, which routes the contested post into an adjudication path where a senior reviewer applies the written guideline and produces the training label. Second, an operation-level agreement trend, tracked per policy category, which says whether the guideline is teachable at all. Three rather than two is chosen so the common case resolves by majority without escalation and only a genuine split costs adjudicator time. Redundancy is annotation spend, so a category with sustained high agreement can drop to fewer annotators while a new clause keeps three.

go deeper

for a junior

Recall that each reported post gets more than one independent verdict, and that the extra verdicts exist to make disagreement visible rather than to average opinions together.

for a middle

Explain the two consumers of the second and third verdict: the item-level route into adjudication, and the per-category agreement trend that says whether the written guideline is teachable.

for a senior

Show where redundancy is spent. High-agreement categories drop to fewer annotators while edited clauses keep three, capture is blind, and the adjudicated verdict rather than the majority becomes the training label.

for a principal

Argue redundancy as a budget line. Every extra verdict is capacity not spent on fresh volume or on auditing automatic decisions, so the split is justified per policy category by the cost of a wrong decision there.

## The problem with a single verdict A labeling operation exists to **manufacture** the training label for a moderation model: the platform does not find labels lying in its logs, it pays people to produce them against a written policy. A single annotator's verdict on a reported post is an assertion with nothing attached to it. The record says `remove`, and the record alone cannot distinguish a correct read of an obvious case from a rushed read of a genuinely hard one. Nothing in the operation flags that item as worth a second look, and nothing in the aggregate tells you whether the policy document the annotator worked from is even usable. Redundancy is what turns an opinion into a measurement. ## What the second and third verdicts actually produce They produce two separate outputs, consumed by different people: - **An item-level disagreement flag.** A post on which the three annotators split is, by definition, a post the guideline does not resolve cleanly. That flag is the routing signal into the adjudication path. - **An operation-level agreement trend, per policy category.** Rising disagreement inside one category is the earliest evidence that its wording is ambiguous, that it overlaps another category, or that what the queue is sending has shifted. - **A qualification surface.** Overlap between a new annotator and the established pool is how you tell, in their first days, whether they have absorbed the guideline, before their verdicts silently enter the training set. ## The adjudication path 1. The three verdicts are captured **blind**: no annotator sees another's answer. If the third reviewer can see the first two, you have bought an echo rather than a verdict. 2. Unanimous items pass straight through; the shared verdict becomes the training label. 3. A split item leaves the normal queue for an escalation queue, where a senior reviewer decides and **cites the guideline clause** that decided it. 4. The **adjudicated verdict, not the majority**, becomes the training label. Two annotators being outvoted by the reviewer is the path working as designed. 5. When the adjudicated case is the first of its kind, it is written back into the guideline as a worked example. This is how a policy document accumulates its own case law and stops producing the same split next month. ## Why three rather than two | annotators per item | what you learn | relative cost | where it fits | |---|---|---|---| | one | nothing about the verdict's correctness | 1x | settled categories already covered by known-answer items | | two | that they disagree, never which one is right | 2x | categories with a long history of high agreement | | three | a majority resolves the common case, so only a real split escalates | 3x | new clauses, new annotator cohorts, contested categories | Two annotators detect disagreement but resolve nothing: every split goes to an adjudicator, which is the most expensive seat in the operation. A third verdict lets the routine case settle itself and reserves escalation for items that are genuinely on the boundary. ## Agreement is a health metric, not a proof of correctness High agreement means the pool applies the guideline **consistently**. It does not mean the pool applies it **correctly**: a group trained together can settle on a shared misreading of a clause and agree with each other all the way to a wrong training set. Consistency and correctness are different properties, and only an item whose correct verdict is already known can separate them. Two further cautions: - With three or more annotators the measure has to correct for the agreement you would get by chance, which is what statistics such as Fleiss' kappa and Krippendorff's alpha are for; how any of them reads depends on the category's base rate. - The useful artefact is the **trend per category**, not one number for the whole operation. A single operation-wide score averages a healthy category against a broken one and hides both. ## Spending the redundancy where it pays Every extra verdict is annotation capacity that could have bought coverage of fresh volume instead, so redundancy is allocated, not applied uniformly: - A category at sustained high agreement can drop to fewer annotators per item, with a known-answer rate holding the floor, and the freed capacity moves to a contested category. - A newly edited clause or a newly onboarded cohort goes back to three until the agreement trend recovers. - Disagreement is **never** resolved by discarding the item. The contested posts sit exactly on the decision boundary the model most needs to learn; dropping them leaves a training set made only of the easy cases, and a model that looks excellent until it meets a real one.

  • If two annotators say keep and the reviewer says remove, which verdict enters the training set?
    The reviewer's. The majority is a routing device for the routine case, not the authority on the label: a split item is by definition one the guideline did not resolve, and the adjudication path exists precisely to settle it against the written clause. The record keeps all three original verdicts alongside the adjudicated one, because the disagreement itself is evidence about the guideline.
  • Why must the three annotators not see each other's verdicts?
    Visible verdicts collapse three independent readings into one anchored reading. The later annotators drift toward what is already on screen, agreement rises for a reason that has nothing to do with the guideline, and the disagreement flag stops firing on exactly the ambiguous items it exists to catch. You have paid three times for one opinion.
  • When is one annotator per item defensible?
    In a category whose agreement has been high for a long time and where a steady rate of known-answer items keeps scoring each annotator individually, so correctness is still measured even though disagreement is not. It is a spending decision reviewed per category, and it reverts to three the moment the clause is edited or the cohort changes.

Two clocks that disagree at least tell you one of them is wrong; a single clock tells you nothing at all. The third clock is what lets the room act on the common case without stopping to reset everything.

saying these in an interview costs you the question

  • Majority vote alone is enough, so no adjudication path is needed
  • Treats every disagreement as annotator error rather than a possible guideline gap
  • Assumes an experienced annotator needs no second verdict on any item
  • Thinks agreement near 1.0 proves the labels match the guideline
  • Applies three annotators to every category regardless of what it costs
  • Resolves disagreement by dropping the contested post from the set