Why report Cohen's kappa instead of raw percent agreement between two annotators?
answer
- raw agreement has no fixed zero
- subtract what chance would give you
- observed minus expected, over the headroom
- skewed classes crush the denominator
- two raters Cohen, many raters Fleiss
basics
~20 sPercent agreement counts the agreements two annotators would hit by chance alone. Cohen's kappa removes that baseline: kappa = (observed agreement - chance agreement) / (1 - chance agreement), so 0 means chance-level and 1 means perfect.
solid answer
~50 sRaw agreement is unanchored. If 95% of comments are not toxic, two annotators who both label almost everything not-toxic will agree about 90% of the time while carrying almost no shared judgement. Kappa fixes the scale by estimating the agreement expected if each annotator applied their own observed label rates independently, then reporting how much of the remaining headroom the pair actually captured. Cohen's kappa is for two annotators on nominal labels; Fleiss' kappa generalises to a fixed number of raters per item; weighted kappa or Krippendorff's alpha handle ordinal scales, and alpha also tolerates missing ratings. One trap is worth naming: on very skewed labels chance agreement is already high, so kappa can look poor even when the pair genuinely agrees -- report the prevalence and the confusion table beside it. The commonly quoted 0.6-0.8 'substantial' band is a convention, not a law.
code
python · 13 lines# Two annotators, 100 comments, label = toxic / not toxic
both_yes, a_yes_b_no, a_no_b_yes, both_no = 15, 10, 5, 70
n = both_yes + a_yes_b_no + a_no_b_yes + both_no
p_o = (both_yes + both_no) / n # observed agreement
a_yes = (both_yes + a_yes_b_no) / n # annotator A's toxic rate
b_yes = (both_yes + a_no_b_yes) / n # annotator B's toxic rate
p_e = a_yes * b_yes + (1 - a_yes) * (1 - b_yes) # chance agreement
kappa = (p_o - p_e) / (1 - p_e)
print("observed", round(p_o, 3)) # 0.85
print("chance ", round(p_e, 3)) # 0.65
print("kappa ", round(kappa, 3)) # 0.571go deeper
Know that agreement between annotators has to be measured, and that a raw agreement percentage flatters you when one class dominates. Being able to state that kappa subtracts chance agreement is enough here.
Explain the formula and what each term means, name which coefficient fits two raters, many raters and ordinal scales, and describe how skewed prevalence deflates kappa without the labels being bad.
Show you would run agreement on a deliberate overlap set early, diagnose low agreement as a guideline defect, and drive the adjudicate-rewrite-remeasure loop on a fresh sample rather than the adjudicated items.
Own how much redundancy the annotation budget buys: when to spend on more items versus more raters per item, what agreement level gates a dataset into production use, and who arbitrates the label definition.
## The problem with percent agreement Suppose three crowd annotators are asked whether a comment is toxic, and you compute pairwise agreement on the items two of them both saw. You get 92%. Is that good? The number alone cannot say, because two annotators who never read the comments and simply guessed at the observed base rate would already agree most of the time when one class dominates. Percent agreement has no fixed zero: its floor moves with the label distribution. ## What kappa does Cohen's kappa rescales agreement so that chance sits at zero: ``` kappa = (p_o - p_e) / (1 - p_e) ``` - `p_o` is **observed agreement** -- the fraction of items the two annotators labelled the same. - `p_e` is **expected agreement** -- what they would reach if each annotator kept their own observed rate of using each label but assigned labels independently of the item. For two classes, `p_e = a_yes*b_yes + (1-a_yes)*(1-b_yes)` where `a_yes` and `b_yes` are the two annotators' marginal rates of the positive label. The numerator is the agreement above chance; the denominator is the agreement above chance that was available. So kappa answers: of the headroom that existed, how much did they use? Kappa is 1 for perfect agreement, 0 when the pair does no better than their own marginals predict, and negative when they systematically disagree more than chance -- which usually means a label was inverted somewhere. ## The prevalence trap This is the single most-asked follow-up. When one class is rare, `p_e` is already close to 1, so `1 - p_e` is tiny and kappa becomes extremely sensitive. A pair can agree on 95% of items and still score kappa near 0.2. That is not a bug -- it is telling you honestly that almost all of the agreement was about the easy majority and very little shared judgement was demonstrated on the cases that matter. The correct response is not to abandon kappa but to report it with the prevalence and the full agreement table, and often to measure agreement on an enriched sample where positives are over-represented. ## Which coefficient for which situation - **Cohen's kappa** -- exactly two annotators, nominal categories, both rating the same items. - **Fleiss' kappa** -- a fixed number of ratings per item, where the raters need not be the same people on every item. This is the crowd-annotation case. - **Weighted kappa** -- ordinal categories, where being one step apart should count as a smaller disagreement than being three steps apart. - **Krippendorff's alpha** -- the most general: any number of raters, missing ratings allowed, and a distance function chosen for nominal, ordinal or interval data. Using Cohen's kappa on three raters, or an unweighted coefficient on an ordinal scale, is a common and visible mistake. ## Using it in practice Agreement is a gate you run *before* you trust a labelled dataset, on a deliberate overlap set: route some fixed share of items -- often 5-10% -- to two or more annotators, and compute agreement only on that overlap. Do it early, on the first batch, while the guidelines can still be changed. Interpretation benchmarks such as 0.41-0.60 'moderate' and 0.61-0.80 'substantial' come from an old convention and are rules of thumb, not thresholds you can defend on their own; what matters is whether the disagreement rate is small relative to the effect you are trying to detect. ## When agreement comes back low The reflex is to blame annotators. Usually the guideline is at fault. Low kappa most often means the label definition is ambiguous at exactly the cases that occur most: the boundary between sarcasm and toxicity, between a complaint and a threat, between 'inactive' and 'churned'. The productive loop is: pull the disagreed items, adjudicate them with a senior annotator, turn each adjudication into a worked example in the guideline, retrain the annotators, and re-measure on a *fresh* overlap set. Re-measuring on the same items you just adjudicated only tells you the adjudication happened. The agreement number is also a planning input. A pair whose kappa is 0.9 barely benefits from a third rater; a pair at 0.4 tells you that single-rater labels are close to worthless and that either the guideline needs work or every item needs multiple ratings with a resolution rule.
- Three crowd annotators rate each item instead of two. What changes?Cohen's kappa is defined for two raters, so you move to Fleiss' kappa, which assumes a fixed number of ratings per item but allows different people to supply them -- exactly the crowd case. If coverage is uneven, some items rated twice and others four times, use Krippendorff's alpha, which tolerates missing ratings and lets you pick a distance function suited to nominal or ordinal labels.
- Two annotators agree on 95% of items but kappa is 0.2. What do you conclude?The label is heavily skewed. When one class covers almost everything, chance agreement is already near 95%, so the correction removes nearly all of the observed agreement and kappa collapses. It is telling you the pair demonstrated little shared judgement on the rare class. Report prevalence and the agreement table alongside the coefficient, and consider measuring agreement on a sample enriched with positives.
- Agreement comes back low. What is your first move?Assume the guideline is ambiguous before assuming the annotators are careless. Pull the disagreed items, have a senior annotator adjudicate them, and convert each ruling into a worked edge-case example in the written definition. Retrain annotators on the updated guideline, then re-measure on a fresh overlap set -- re-scoring the items you just adjudicated proves nothing.
- How does the agreement number inform how many annotators you buy per item?It sets the value of redundancy. If kappa is very high, a second rating buys almost nothing and the budget is better spent on more distinct items. If kappa is mediocre, single-rater labels are unreliable, so you need multiple ratings per item plus a resolution rule -- majority vote, or adjudication by a senior annotator for ties and for any item flagged as hard.
saying these in an interview costs you the question
- Reports 95% agreement with no chance correction
- Reads kappa as the percentage of items agreed on
- Uses Cohen's kappa for three or more annotators
- Treats low kappa as careless annotators, not a vague guideline
- Says kappa of 0 means the annotators never agreed
- Uses an unweighted coefficient on an ordinal rating scale