skip to content

Inter-annotator agreement on a moderation queue collapses the week after a policy edit. What does that drop indicate?

level: seniorimportance: should knowfreq 47%

answer

  1. suspect the wording before the people
  2. is the drop local to edited categories?
  3. an edit changes the target definition
  4. labels now straddle two guideline versions
  5. measure the flip rate on a re-adjudicated sample

basics

~20 s

Suspect the guideline before the people: a freshly edited clause that two careful annotators can read differently shows up immediately as disagreement. The diagnostic is locality, whether the drop sits only in the categories the edit touched.

solid answer

~40 s

Treat the wording as the defect until proved otherwise. A policy edit changes the target definition, and an under-specified clause, or one that overlaps a category already in the guideline, makes two conscientious annotators reach different verdicts the moment it lands. The diagnostic is **locality**. If the drop is confined to the categories the edit touched, the new wording is at fault and the fix is worked examples plus an adjudication precedent, not annotator retraining. If agreement also fell in untouched categories, something else changed that week, such as a new annotator cohort, a change to what the queue is sampling, or an interface change that hid context. The second consequence is that the label set now straddles two definitions, which is why every label record carries the guideline version it was produced under.

code

json · 15 lines
json
{
  "postId": "p-8813204",
  "guidelineVersion": "policy-2026-09-14",
  "verdicts": [
    { "annotatorId": "an-114", "verdict": "remove", "clause": "3.2-harassment" },
    { "annotatorId": "an-207", "verdict": "keep", "clause": null },
    { "annotatorId": "an-318", "verdict": "keep", "clause": null }
  ],
  "unanimous": false,
  "routedTo": "adjudication",
  "adjudicatedVerdict": "remove",
  "adjudicationClause": "3.2-harassment",
  "trainingLabel": "remove",
  "labeledAt": "2026-09-16T11:04:00Z"
}

go deeper

for a junior

Recall that the written guideline defines what a verdict means, so changing it changes what every later label is asserting, even about posts that were never edited.

for a middle

Explain why the wording is the first suspect and what decomposing the agreement trend by category, by cohort and by incoming mix tells you about which explanation holds.

for a senior

Show the operational response: re-adjudicate a stratified pre-edit sample to measure the flip rate, stamp the guideline version on every verdict, and fix the clause with worked examples and a decision order rather than retraining people.

for a principal

Treat a policy edit as a versioned change to a dataset contract. Decide in advance what flip rate forces re-labelling, who pays for it, and whether the edit ships to annotators and to the training set on the same day.

## What a policy edit actually changes An edit to the written guideline changes the **target definition** the operation is manufacturing labels against. Everything downstream inherits that: the verdicts produced after the edit answer a slightly different question from the verdicts produced before it, on posts whose text never changed at all. Two consequences follow immediately, and they are usually confused with each other. 1. **A measurement consequence.** If the new clause can be read two ways, careful annotators diverge and agreement falls in the affected categories from the first shift after the edit. 2. **A dataset consequence.** The label set now contains verdicts made under two different definitions, and without the version recorded on each verdict you cannot tell the two apart later. ## Reading the drop: locality is the diagnostic Before anyone is retrained, decompose the agreement trend: - **By category.** A drop confined to the categories the edit touched points squarely at the wording. This is the common case and the cheapest to fix. - **By annotator cohort.** A drop that is concentrated in one cohort, across categories, points at onboarding rather than at policy - a group that started that week, or a group that attended one particular calibration session. - **By queue composition.** Agreement also falls when the mix of posts arriving gets harder, with no change to policy or people. If the sampling that feeds the queue was changed in the same window, that is a competing explanation. - **By surface or interface.** If reviewers stopped seeing the surrounding thread, the profile history or the media preview, they are deciding on less context and will diverge on exactly the borderline items. | observation | most likely cause | the fix | |---|---|---| | drop only in edited categories | the new clause is ambiguous or overlaps an existing one | rewrite with worked examples and a decision order between overlapping clauses | | drop across all categories, one cohort | onboarding or calibration for that cohort | re-qualify that cohort against known-answer items | | drop across all categories, all cohorts | queue composition or interface change | compare the incoming mix and the reviewer's available context week over week | ## The two definitions now inside your labels Every verdict has to be stamped with the guideline version that produced it, alongside the annotators' original verdicts, the adjudicated verdict where one exists, and the clause cited. Without the version stamp you cannot answer the one question that matters when a label disagrees with today's policy: **did the verdict change, or did the annotator?** With the stamp you can measure the size of the problem directly. Pull a stratified sample of pre-edit posts from the affected categories, send them back through the adjudication path under the new wording, and count how many verdicts flip. A small flip rate means the edit was a clarification, the older verdicts still encode the same target, and they can stay with their version recorded. A large flip rate means the older verdicts are answers to a different question in those categories, and they have to be re-labelled or held out of the affected categories rather than quietly mixed in. ## Fixing the guideline, not the annotators A guideline that produces disagreement is not fixed by adding another sentence of rule. What reliably moves agreement is: - **Worked examples on both sides of the line**, drawn from real adjudicated posts, especially near-miss pairs that differ in exactly one respect. - **A stated decision order** where two clauses can both apply, so an annotator does not have to invent a precedence rule on the spot. - **Adjudication precedents written back** into the document, which is how a guideline accumulates case law and stops generating the same split every week. - **Re-qualification against refreshed known-answer items**, whose recorded answers were re-derived under the new wording, not the old one. ## What not to conclude - Do not conclude the pool got careless. A simultaneous collapse across conscientious annotators is far better explained by the thing that changed that week. - Do not wait for it to settle. Agreement does eventually recover on its own, because annotators converge on *some* shared reading, and there is no guarantee that reading is the one the policy intended. - Do not assume the pre-edit labels are fine because the posts are unchanged. The post is unchanged; the question asked about it is not.

  • Agreement also fell in categories the edit never touched. What does that suggest?
    That something other than the wording changed in the same window. The usual candidates are a new annotator cohort onboarded that week, an interface change that removed context such as the surrounding thread, and a shift in what the queue is sampling toward harder posts. Decompose the trend by cohort, by category and by incoming mix before touching the guideline, because retraining a pool for a queue-composition change fixes nothing.
  • What do you do with verdicts made under the previous guideline version?
    Measure before deciding. Re-adjudicate a stratified sample of pre-edit posts in the affected categories under the new wording and count the flips. A small flip rate means they encode the same target and can stay, with the version on the record. A large flip rate means those older verdicts answer a different question and must be re-labelled or excluded from the affected categories.
  • How do you tell an ambiguous clause from one that simply overlaps an existing category?
    Look at the shape of the disagreement. An ambiguous clause produces splits between applying it and not applying it at all. An overlap produces splits between two different categories, each annotator confident and each citing a different clause. The first is fixed with worked examples and a sharper definition, the second with a stated decision order between the two clauses.

saying these in an interview costs you the question

  • Blames the annotator pool and schedules retraining before rereading the clause
  • Treats the agreement drop as noise that will settle on its own
  • Mixes pre-edit and post-edit verdicts with no guideline version recorded
  • Assumes older labels stay valid because the posts themselves did not change
  • Patches the guideline with another rule and no worked example
  • Scores annotators against known-answer items still keyed to the old policy