Model B beats model A by 0.3 accuracy points on a 1,000-example test set - is that real?
answer
- how many examples is 0.3 points?
- same test set means paired
- only the disagreements carry evidence
- gap equals (b - c) divided by n
- net of three caps the statistic near 3
basics
~20 sNot on that evidence. On 1,000 examples, 0.3 accuracy points is three examples, and a net of three can never reach statistical significance in a paired comparison. The gap sits inside the test set's noise.
solid answer
~50 sAccuracy on 1,000 examples moves in steps of 0.1 points, so a 0.3-point gap is exactly three examples: model B got three more right than model A. Since both were scored on the same examples, the honest analysis is paired. Let `b` be the examples B gets right and A wrong, and `c` the reverse; the accuracy gap is `(b - c) / n`, so `b - c = 3`. McNemar's statistic is `(b - c)^2 / (b + c)`, and because `b + c` is at least 3, the statistic is at most `9 / 3 = 3` - reached only if A wins none of the disagreements - which is below the 3.84 cutoff for one degree of freedom. The exact version is even blunter: three coin flips landing the same way give a two-sided p of 0.25. No split of the disagreements makes this win significant.
go deeper
Be ready to convert a percentage gap into a count of examples on the spot, and to say out loud that a difference of a few examples is not evidence. Knowing that both models were scored on the same data is the key detail.
Explain why the accuracy gap equals the difference of the two disagreement counts divided by n, and show that a net of three examples cannot push McNemar's statistic past the one-degree-of-freedom cutoff.
Show what you would actually do next: quantify the test set's resolution, estimate how many examples the comparison would need, and refuse to declare a winner in the write-up rather than hedging with a footnote.
Own the standard your team applies before any model comparison starts: how big a gap counts, how large the evaluation set must be to see it, and why a culture of shipping on unresolvable differences quietly produces a stream of neutral releases.
## What "0.3 accuracy points" actually means Accuracy on a 1,000-example test set can only move in steps of 0.1 percentage points, because one example is `1/1000`. So a 0.3-point gap is not an abstract quantity - it is **three examples**. Whatever the two models did on the other 997, model B ended up with three more correct answers than model A. Saying it that way already reframes the conversation: nobody ships a model on the strength of three examples without checking whether three is more than the test set's ordinary jitter. ## The comparison is paired, not two separate scores Both models were run over the *same* 1,000 examples. That means the two accuracy numbers are not two independent measurements; they share every quirk of the test set. The information about which model is better lives entirely in the examples where the two models **disagree**. Write the four paired outcome counts for the test set: - `a` - both models right - `b` - model B right, model A wrong - `c` - model A right, model B wrong - `d` - both models wrong Model B's correct count is `a + b` and model A's is `a + c`, so the accuracy gap is ``` accuracy_B - accuracy_A = ((a + b) - (a + c)) / n = (b - c) / n ``` The `a` and `d` cells cancel exactly. A 0.3-point gap on `n = 1000` therefore means `b - c = 3`, no matter whether the models disagreed on 5 examples or 300. ## Why a net of three examples cannot be significant McNemar's test asks whether the disagreements split more lopsidedly than fair coin flips would. Its chi-square form, with one degree of freedom, is ``` X^2 = (b - c)^2 / (b + c) ``` The 5% critical value for one degree of freedom is about **3.84**. With `b - c = 3`, the numerator is fixed at 9, and the denominator `b + c` is at least 3 (you cannot have a net of three from fewer than three disagreements). So ``` X^2 <= 9 / 3 = 3 < 3.84 ``` and the best case - `b = 3`, `c = 0`, meaning model A never once wins an example that B loses - still falls short. The exact form makes the same point more directly: under the null, the discordant examples behave like `b + c` fair coin flips, and three flips all landing the same way has two-sided probability `2 * 0.5^3 = 0.25`. Every other split of a net-three win is weaker still. There is no arrangement of a 0.3-point gap on 1,000 examples that clears a conventional threshold. ## How much data would you need? Suppose the two models keep disagreeing on a fraction `r` of examples and keep the same 0.3-point edge. Then `b - c = 0.003n` and `b + c = rn`, so ``` X^2 = (0.003n)^2 / (rn) = 0.000009 * n / r ``` Setting that at 3.84 gives `n = 426,667 * r`. With 5% disagreement that is roughly 21,000 examples; with 10% disagreement, roughly 43,000. A 0.3-point effect is simply a large-sample question, and a 1,000-example test set was never going to answer it. ## What to report instead Report the paired gap with an uncertainty statement rather than a bare winner: "model B is ahead by 3 examples out of 1,000; the paired interval comfortably covers zero; this test set cannot resolve differences below roughly one point." Naming the resolution of your test set is far more useful to the reader than a p-value, and it turns the next conversation into "do we need a bigger evaluation set?" rather than "which model won?". ## Traps that show up in interviews The most common mistake is comparing the two accuracy numbers as if they were independent proportions - that both ignores the pairing and still would not rescue a three-example gap. The second is treating a positive difference as evidence of any size at all, when the sign of a tiny difference is close to a coin flip. The third is quoting the standard error of a *single* accuracy estimate (about 0.95 points at 90% accuracy on 1,000 examples) as the yardstick for the gap; that number is too conservative for a paired comparison, because the paired standard error of the difference is usually much smaller. Use the paired machinery, and then note that even the paired machinery cannot make three examples significant.
- How large would the test set have to be before a 0.3-point gap could be detected?It depends on how often the models disagree. If they disagree on a fraction r of examples, McNemar's statistic is roughly 0.000009 * n / r, so reaching 3.84 needs about 426,667 * r examples: around 21,000 at 5% disagreement and around 43,000 at 10%. Small effects need large evaluation sets, and no amount of analysis rescues 1,000.
- Why is the standard error of a single accuracy estimate a misleading yardstick here?At 90% accuracy on 1,000 examples, one model's accuracy has a standard error near 0.95 points, which makes any small gap look hopeless. But the two scores come from the same examples and move together, so the paired standard error of the difference is usually much smaller. Judge the gap with the paired analysis, not by eyeballing two single-model error bars.
- How would you report this result honestly rather than claiming a win?State the gap in examples, not just percent: three out of a thousand. Give an interval for the paired difference and say plainly that it includes zero. Then state the resolution of the test set - the smallest gap it could have detected - so the reader knows the comparison was underpowered rather than negative.
Three votes apart in a thousand-vote straw poll is not a mandate, it is a rounding error - and you would ask for a recount, not a coronation.
saying these in an interview costs you the question
- Calls a 0.3-point gap an improvement because the sign is positive
- Compares the two accuracy numbers with no uncertainty at all
- Treats scores from one shared test set as independent samples
- Never converts the percentage gap into a count of examples
- Argues the newer architecture must be better, so the gap is real