skip to content

Why report Cohen's kappa instead of raw agreement for a 5-grade severity classifier?

level: seniorimportance: should knowfreq 38%

answer

  1. compared with agreeing by luck
  2. observed minus expected, then rescaled
  3. chance term comes from the marginals
  4. zero means no better than the distribution-matcher
  5. ordered grades want distance-based penalties

basics

~10 s

Cohen's kappa subtracts the agreement expected by chance from the observed agreement and rescales the remainder, so a model that merely mimics the grade distribution scores near zero even when raw agreement looks high.

solid answer

~50 s

Raw agreement — the share of cases where the predicted grade matches the true grade — flatters any model on a skewed label distribution, because guessing the common grade is already right most of the time. Kappa corrects for that: `kappa = (p_o - p_e) / (1 - p_e)`, where `p_o` is observed agreement and `p_e` is the agreement two independent labellers with those same marginal distributions would reach by luck. It is 1 for perfect agreement, 0 for chance-level, and negative for systematically worse than chance. On a 5-grade severity classifier with 82% observed agreement and 74% chance agreement, kappa is (0.82 - 0.74) / 0.26 = 0.31 — modest, not excellent. Two cautions: kappa depends on the marginals, so values are not comparable across datasets with different grade mixes; and for *ordered* grades use a weighted kappa, which penalises calling a grade-5 case a grade-1 far more than calling it a grade-4.

go deeper

for a junior

Recall the shape of the statistic: observed agreement minus chance agreement, divided by one minus chance agreement. Know that zero means chance-level and one means perfect, and that raw agreement flatters skewed label distributions.

for a middle

Compute the chance term from the marginals and explain why a model that copies the label distribution scores near zero. Be ready to work a small numeric example and to state the range including negative values.

for a senior

Show you know kappa's limits in operation: marginal dependence makes it non-comparable across test sets, negative values usually mean a label-mapping bug, and ordinal grades need distance-weighted kappa. Pair it with the per-class table rather than reporting it alone.

for a principal

Decide whether a chance-corrected statistic belongs in the organisation's model-acceptance criteria at all, what weighting scheme encodes the real cost of misgrading, and how to keep an evaluation set stable enough that the number means the same thing quarter to quarter.

## What kappa fixes A 5-grade severity classifier is scored by comparing its grade to the gold grade, case by case. The obvious statistic is the share that match. The problem is that severity grades are rarely uniform: if 70% of cases are grade 2, a model that outputs grade 2 for everything already matches 70% of the time while carrying no information. Any statistic that cannot distinguish that model from a useful one is not measuring skill. Cohen's kappa is the chance-corrected version: ``` kappa = (p_o - p_e) / (1 - p_e) ``` - `p_o` — **observed agreement**, the proportion of cases where prediction and gold grade coincide (the diagonal of the confusion matrix over N). - `p_e` — **expected agreement**, what you would get if the predictions and the gold labels were statistically independent but kept their own marginal frequencies. For each grade g, multiply the share of gold labels equal to g by the share of predictions equal to g; sum over grades: `p_e = sum_g (row_g / N) * (col_g / N)`. The numerator is the agreement above chance; the denominator is the agreement *available* above chance. The ratio therefore answers: of the improvement that was there to be had, how much did the model achieve? ## Reading the scale - **1.0** — every case matches. - **0.0** — exactly chance-level; the model's grade distribution explains its hits entirely. - **Negative** — worse than chance, which in practice signals a systematic mapping error such as swapped or shifted grade codes. A negative kappa is a bug signal, not a bad-model signal. The published verbal bands ("substantial", "moderate") are conventions from the literature, not properties of the statistic; do not present them as thresholds. Whether 0.31 is acceptable depends entirely on what the grade drives. ## A worked case Suppose 1,000 severity cases, mostly grades 2 and 3, and the model matches on 820 of them, so `p_o = 0.82`. Computing `p_e` from the marginals gives 0.74 — the label distribution is skewed enough that a distribution-matching guesser lands nearly three quarters of the time. Then: ``` kappa = (0.82 - 0.74) / (1 - 0.74) = 0.08 / 0.26 = 0.31 ``` The headline "82% agreement" and the honest "kappa 0.31" describe the same model. Only the second survives the question "compared with what?" ## Kappa depends on the marginals This is the caveat that separates people who have used kappa from people who have read about it. `p_e` is computed from the observed marginals of this dataset, so the same model can score very different kappas on two test sets with different grade mixes — and a *more* balanced test set generally yields a *higher* kappa for the same underlying behaviour, because chance agreement is lower. Consequences: - Never compare kappas across datasets, time periods, or sites whose grade distributions differ, and never chart kappa over time if the case mix drifts. - Two models with identical accuracy can have different kappas if their *prediction* marginals differ, since `p_e` uses both marginals. - Kappa is a summary, not a diagnosis. It never tells you which grade is failing; keep the per-class table beside it. ## Weighted kappa for ordered grades Severity grades are **ordinal**: grade 1 to grade 5 is a worse mistake than grade 4 to grade 5, and plain kappa treats every disagreement as equally wrong. Weighted kappa assigns each off-diagonal cell a penalty by grade distance — linear weights penalise proportionally to the number of grades apart, quadratic weights penalise the square of that distance, which is the common choice when large misgrades are disproportionately costly. Reporting unweighted kappa on ordered grades throws away the ordering the label scheme was designed around, and it is a fair thing for an interviewer to catch. ## Matthews correlation, the binary sibling For a two-class problem the analogous chance-aware summary is the Matthews correlation coefficient: ``` MCC = (TP*TN - FP*FN) / sqrt((TP+FP)(TP+FN)(TN+FP)(TN+FN)) ``` It is the correlation between predicted and true labels, running from -1 to +1, and unlike F1 it uses **all four** confusion-matrix cells — including the true negatives that F1 ignores entirely. On a skewed document-type model where one type is 3% of the corpus, a classifier can post a respectable F1 on the rare class while its behaviour on the vast negative pool goes unexamined; MCC is high only when the model does well on both sides. A multiclass generalisation exists and behaves similarly. MCC is likewise marginal-dependent, so the same non-comparability warning applies. ## When to reach for these Use chance-corrected statistics when the label distribution is skewed enough that raw agreement is unpersuasive, when you must defend a model to a reviewer who will ask "how much of that is luck", or when grades are ordinal and distance matters. Do not use them as the only number: report kappa *and* the confusion matrix, because kappa is designed to answer one narrow question and answers nothing else.

  • Why is a kappa of 0.55 on one test set not comparable to 0.55 on another?
    Expected agreement is computed from each dataset's own marginals, so a more balanced set has lower chance agreement and yields a higher kappa for identical model behaviour. Comparing across sets, sites or time periods with drifting class mix compares denominators, not models. Fix the evaluation set, or report the confusion matrix alongside.
  • When would you report Matthews correlation instead of F1 on a binary classifier?
    When true negatives matter and the classes are skewed. F1 is built only from TP, FP and FN, so it never looks at the large negative pool; MCC uses all four cells and is high only when both classes are handled well. On a 3%-positive document-type model that difference is exactly where a flattering F1 hides weak behaviour.
  • What does a negative Cohen's kappa usually indicate in practice?
    Systematic disagreement rather than mere weakness — the model matches less often than its own marginals predict. In real projects that almost always means a wiring bug: grade codes swapped, an off-by-one in the label mapping, or predictions aligned to the wrong rows. Check the mapping before concluding the model is bad.

Scoring a five-option quiz where blind guessing already earns you points: kappa asks how far above the guesser you actually got, not how many you got right.

saying these in an interview costs you the question

  • Reads kappa as a percentage of correct predictions
  • Compares kappa across datasets with different class mixes
  • Believes kappa cannot be negative
  • Uses unweighted kappa on ordered severity grades
  • Claims chance agreement is always one over the number of classes
  • Treats published verbal bands as hard acceptance thresholds

context