Expert annotators agree only 92% of the time - how much headroom does your model really have?
answer
- disagreement is not the same as error
- if they differ, usually only one is wrong
- watch out for skewed label sets
- ambiguity versus inconsistent guidelines
- noisy test labels hide real gains
basics
~20 sPairwise agreement of 92% is not an 8% error floor. If annotators err independently, each is wrong on roughly 4% of items. Adjudicate a sample to split irreducible ambiguity from fixable guideline drift before funding more modelling.
solid answer
~50 sAn 8% disagreement rate is a symptom, not a floor. Two annotators disagreeing means at least one is wrong, so under independent errors each is wrong on about 4% of items - the floor is nearer 4% than 8%, and lower still once you subtract agreement that happens by chance on a small label set. More importantly, disagreement has two very different causes. **Irreducible ambiguity** - the item genuinely admits two readings - really is floor, and no model or budget removes it. **Guideline drift** - annotators applying different rules - is fixable, and fixing it lowers the floor *and* cleans the labels the model learns from. So I adjudicate a stratified sample of disagreements with a panel and classify each one; that decides whether the next quarter buys modelling or labelling. It also caps measurement - a single-annotator test set adds noise, so improvements smaller than that noise are invisible.
go deeper
Recall that human labels are not perfect and that a task where experts disagree has an error floor above zero. Know that reported accuracy is measured against those imperfect labels.
Explain why a disagreement rate overstates per-annotator error, and why agreement on a skewed label set needs chance correction before it means anything.
Show the adjudication workflow: sample the disagreements, classify each as ambiguity, guideline drift or carelessness, and turn that split into a defended floor estimate with the remaining headroom attached.
Own the spending decision that follows. Recommend modelling, labelling-process investment, or stopping, and defend the floor estimate behind it - including buying clean evaluation labels so improvements are observable at all.
## Agreement is not error A reported "92% inter-annotator agreement" is a *pairwise* statistic: on 92 of every 100 items two annotators chose the same label. It is tempting to read the complement as the error floor and announce that no model can beat 8% error. That is wrong twice over. **First, disagreement double-counts.** When two annotators disagree, at least one is wrong - but usually only one. Model each annotator as erring independently with probability e. In a binary task they agree when both are right or both are wrong, so agreement = (1-e)^2 + e^2. Setting that to 0.92 gives e of about 0.04. With more than two labels, agreeing while both wrong is rarer still, and e sits just under half the disagreement rate. So a defensible reading of 92% agreement is **a floor around 4%, not 8%**. **Second, raw agreement flatters you on skewed label sets.** If 90% of items belong to one class, two annotators who label everything with the majority class agree 90% of the time while carrying no information. Chance-corrected measures - Cohen's kappa for two annotators, or a multi-rater equivalent - subtract the agreement expected from the marginal label frequencies. A raw 92% with a kappa near zero says the annotation process is not measuring anything, and no floor estimate built on it is meaningful. ## Two kinds of disagreement, two different bills The decisive question is not *how much* the annotators disagree but *why*. Take a stratified sample of disagreements - say two hundred, spread across classes and across annotator pairs - and have a panel adjudicate each one to a consensus label, recording the reason: - **Genuine ambiguity.** The item supports two readings; competent people who follow the same rules still split. This is irreducible for the current inputs. It sets a real floor and it does not respond to money. - **Guideline drift.** The rules are silent, contradictory, or interpreted differently across annotators or over time. This is *not* floor. Rewriting the guideline, adding worked examples and re-labelling the affected slice removes it. - **Carelessness or fatigue.** Rushed work, throughput incentives, a bad UI. Also not floor; a process problem. - **Missing information.** The annotator could resolve the item with context the schema does not carry - a preceding message, a customer record. This is floor *for the current feature set* and disappears if you add the input. That distinction matters, because it converts a labelling question into a feature question. In practice the split is rarely dominated by genuine ambiguity, which is why "the labels are just noisy, nothing to do" is the weakest possible answer. ## Sizing the headroom Once the sample is adjudicated you can state the number honestly: "of the 8 points of disagreement, roughly 3 are ambiguity, 4 are guideline drift and 1 is carelessness; the achievable floor is near 2%, not 8%." Now compare with where the model sits. A model at 7% error against an estimated 2% floor has five points of headroom and deserves investment. The same model against a genuine 6% floor has one point, and the next quarter should be spent elsewhere - on latency, coverage, monitoring, or the decision the prediction feeds. ## The measurement ceiling Label noise does not only cap the model; it caps your ability to *see* progress. If the test labels are produced by single annotators who are each wrong 4% of the time, then a perfect model is scored as 4% wrong, and two models three tenths of a point apart cannot be distinguished from label noise. Two consequences follow, and stating them is what separates a lead-level answer: - **Buy quality where you measure.** Adjudicate the evaluation set with multiple annotators even if the training set stays single-pass. Clean test labels are cheap relative to the decisions they inform, and a noisy test set silently rejects real improvements. - **Set the decision threshold by noise, not by hope.** Declare in advance the improvement that will count as real, and require that it exceed the label-noise band and the sampling error of the evaluation set together. ## The organisational call The deliverable here is a spending recommendation with a stated floor and a stated uncertainty on it, not a model change. Three plausible outcomes: fund modelling because the headroom is large; fund the labelling process because most disagreement is drift and fixing it raises every future model's ceiling; or stop optimising because the remaining error is ambiguity and the honest move is to change the inputs, split the ambiguous class into two, or let the product abstain on low-confidence items instead of guessing. Deciding which of those three, and defending the floor estimate that justifies it, is the whole question.
- Why is raw pairwise agreement misleading when one class covers most of the data?Because two annotators who both default to the majority class agree at the base rate while conveying nothing. On a 90/10 split, 90% agreement is the floor of meaninglessness, not a sign of quality. Chance-corrected agreement - Cohen's kappa and its multi-rater relatives - subtracts the agreement expected from the label frequencies and exposes that case.
- How does label noise in the test set change how you compare two candidate models?It puts a band around every measured score. If test labels are wrong on a few percent of items, a perfect model scores below 100% and small differences between candidates are indistinguishable from annotation error. I adjudicate the evaluation set with multiple annotators and set, in advance, the improvement that must be exceeded to count as real.
- The panel finds most disagreement is genuine ambiguity. What do you recommend?Stop buying model capacity. The options are to change the inputs so the ambiguity is resolvable, to redefine the label - splitting a contested class or merging two that nobody separates reliably - or to let the system abstain and route low-confidence items to a human. Reporting a floor and the remaining headroom is more valuable than another modelling quarter.
If two experienced doctors reading the same scan reach different conclusions eight times in a hundred, you learn nothing until you find out whether the scans were genuinely borderline or the two were working from different checklists. One is a limit of the medium; the other is a fixable process.
saying these in an interview costs you the question
- Reads 8% disagreement directly as an 8% error floor
- Quotes raw agreement on a heavily skewed label set
- Treats all annotator disagreement as irreducible noise
- Never inspects the disagreeing items themselves
- Ignores label noise when comparing two models on a test set
- Assumes better labels only matter for training, not evaluation