What does McNemar's test compare when two classifiers are scored on the same test set?
answer
- four paired cells, only two matter
- agreement cells cancel from the gap
- each disagreement is a fair coin flip
- binomial on b out of b plus c
- chi-square form is (b - c) squared over (b + c)
basics
~20 sIt compares only the discordant examples: those one model gets right while the other gets them wrong, counted each way. Under the null those split like fair coin flips, so a lopsided split is the evidence.
solid answer
~50 sLay the paired outcomes in four cells: both right, both wrong, and the two disagreement counts `b` (A right, B wrong) and `c` (A wrong, B right). McNemar's test looks only at `b` and `c`, because the accuracy gap is `(b - c) / n` - the agreement cells contribute equally to both models and cancel. Under the null of equal accuracy, each discordant example is equally likely to fall either way, so `b` follows a Binomial(`b + c`, 0.5). Test it exactly, or use the chi-square form `(b - c)^2 / (b + c)` with one degree of freedom. With `b = 42` and `c = 18`, that is `576 / 60 = 9.6`, well past the 3.84 cutoff, giving a p-value around 0.002: model A is genuinely ahead. Use the exact binomial version when `b + c` is small.
go deeper
Know that McNemar's test exists for comparing two classifiers on one shared test set, and that it looks at the examples where they disagree. Being able to build the four-cell paired table from a description is enough here.
Derive why the agreement cells cancel, state the binomial null on the discordant pairs, and compute the chi-square form on given counts without hesitating over which value goes where.
Show judgment about when the approximation fails, insist on reporting the effect size next to the p-value, and flag clustered test examples where the fair-coin argument has to move to the cluster level.
Be ready to argue what a significant discordant imbalance should and should not authorise, and to set the bar at which a statistically detectable accuracy gap becomes a decision your team acts on.
## The paired table Two classifiers are scored on the same held-out set. Each example produces a *pair* of outcomes, so the natural summary is four counts: ``` B right B wrong A right a b A wrong c d ``` Here `a` is both right, `d` is both wrong, `b` is A right and B wrong, and `c` is A wrong and B right. The cells `b` and `c` are the **discordant** cells. ## Why the agreement cells are ignored Model A's correct count is `a + b`; model B's is `a + c`. Subtract: ``` accuracy_A - accuracy_B = ((a + b) - (a + c)) / n = (b - c) / n ``` The `a` cell cancels because it adds one correct answer to *both* models, and `d` never appears because it adds to neither. So an example that both models handle the same way, however hard or easy, carries no information at all about which of the two is better. Throwing those examples out is not an approximation; it is the exact structure of the quantity being tested. This is also why McNemar's test can be decisive on a test set where the two models agree 95% of the time - all the signal is concentrated in the remaining 5%. ## The null hypothesis and the reference distribution The null is that the two models have the same accuracy, which by the identity above means `b` and `c` have the same expectation. Condition on the total number of disagreements `b + c`; then the null says each disagreement is a fair coin flip about which model wins it. Therefore ``` b ~ Binomial(b + c, 0.5) ``` and the p-value is a two-sided binomial tail. That is the **exact** McNemar test, and it is always valid. The familiar chi-square form is a normal approximation to that binomial: ``` X^2 = (b - c)^2 / (b + c) ``` compared against a chi-square distribution with one degree of freedom (5% cutoff about 3.84, 1% cutoff about 6.63). A continuity-corrected variant, `(|b - c| - 1)^2 / (b + c)`, tracks the discrete binomial a little better. ## Working the numbers Suppose model A is right and B wrong on `b = 42` examples, and the reverse on `c = 18`. There are 60 disagreements and a net of 24 in A's favour. ``` X^2 = (42 - 18)^2 / (42 + 18) = 576 / 60 = 9.6 ``` Against one degree of freedom that gives a p-value near 0.002 - far past the usual threshold - and the continuity-corrected version, `(24 - 1)^2 / 60 = 8.82`, gives about 0.003. Either way, a 42-to-18 split of 60 coin flips is not something the null produces often, so model A is ahead by more than sampling noise. The effect size is separate: if the test set had 10,000 examples, the accuracy gap is `24 / 10000 = 0.24` percentage points, which is significant and possibly still too small to care about. Always report both. ## When to prefer the exact version The chi-square approximation degrades when the number of discordant pairs is small; a common rule of thumb is to use the exact binomial when `b + c` is below about 25. Since the exact test costs nothing to compute, using it by default is a defensible habit, and saying so signals that you know the approximation is the convenience, not the definition. ## What the test does and does not tell you It tells you whether the *marginal* accuracies differ on this test set. It does not tell you how big the gap is - report `(b - c) / n` and an interval alongside. It does not tell you *where* the models differ; two models can have identical accuracy while disagreeing on 30% of examples, and that disagreement pattern is often the more interesting finding. And it does not extend beyond the population your test set was drawn from. ## Extensions worth naming McNemar's test is defined for a 0/1 outcome and exactly two models. If the metric is continuous per example, use the per-example differences directly rather than binarising them. If the test examples are clustered - many per user or per document - the fair-coin argument applies at the cluster level, not the example level, and treating clustered examples as independent will produce a p-value that is too small.
- When should you use the exact binomial version rather than the chi-square form?When the number of discordant pairs is small - a common rule of thumb is fewer than about 25 - because the chi-square form is a continuous approximation to a discrete binomial and is unreliable in that range. The exact two-sided binomial test on b out of b + c with probability 0.5 is valid at any size and costs nothing to compute.
- Two models disagree on 600 of 10,000 examples, with b = 320 and c = 280. Is that a win?No. The statistic is (320 - 280)^2 / 600 = 1600 / 600 = 2.67, below the 3.84 cutoff for one degree of freedom, so the split is consistent with fair coin flips. The observed gap is 40 / 10000 = 0.4 percentage points, and this test set cannot separate it from noise.
- What does a significant McNemar result fail to tell you?It gives no effect size: a large discordant imbalance on a huge test set can be significant while the accuracy gap is a fraction of a point. It also says nothing about which kinds of examples flipped, and nothing about behaviour outside the population the test set represents. Report the gap and an interval next to the p-value.
Score a match hole by hole: the holes where both golfers made par cannot change who is ahead, so only the holes they played differently count.
saying these in an interview costs you the question
- Includes the both-right and both-wrong counts in the statistic
- Runs an unpaired proportion test on the two accuracy totals
- Reads a small p-value as evidence of a large improvement
- Uses the chi-square form when only a few pairs disagree
- Thinks a high disagreement count alone shows one model is better