Why is comparing two models on one shared held-out set a paired problem, not two proportions?
answer
- one random draw, not two
- hard examples are hard for both
- the covariance term is not zero
- dropping it inflates the standard error
- analyse per-example differences instead
basics
~20 sBoth models are scored on the very same examples, so their accuracy estimates are strongly positively correlated. The two-independent-proportions formula sets that correlation to zero, inflates the standard error of the difference, and wastes power.
solid answer
~50 sThe two scores come from one draw of test examples, not two. Hard examples are hard for both models, so the estimates move together. The variance of the gap is `Var(pA) + Var(pB) - 2*Cov(pA, pB)`, and that covariance is large and positive. The unpaired two-proportion formula silently drops it, so it reports a standard error that is too big and a test that is too conservative. Concretely: two models near 90% accuracy on 10,000 shared examples that disagree on 4% of them have a paired standard error of about 0.2 percentage points on the gap, while the unpaired formula gives about 0.42 - more than double. What pairing throws away is not correctness but sensitivity. The fix is to work with per-example differences: the disagreement counts for accuracy, or per-example loss differences for a continuous metric.
go deeper
Remember the headline: one test set means one random draw, so the two scores are linked. Be able to say that the comparison should be made example by example rather than by subtracting two totals.
Write the variance of the difference with its covariance term and explain which term the unpaired formula drops. Being able to say that the paired variance for accuracy equals the disagreement rate is the answer that lands.
Demonstrate the operational consequence: an unpaired analysis is conservative, so it hides real wins, and teams then ship nothing. Show how you would restructure the evaluation around per-example differences.
Own the evaluation contract for the org: every model comparison runs on one frozen shared set, reports an interval on the gap rather than two separate bars, and states the resolution the set can achieve.
## One draw of data, two numbers When you evaluate two models on the same held-out set, there is exactly one random thing in the experiment: which examples landed in that set. Both accuracy figures are computed from that single draw. The two-independent-proportions machinery assumes there were two separate draws, one per model, and that assumption is simply false here. The consequence is a variance error. For the difference of two estimates, ``` Var(pA - pB) = Var(pA) + Var(pB) - 2 * Cov(pA, pB) ``` The unpaired formula keeps the first two terms and assumes the third is zero. But models trained on similar data make similar mistakes: an ambiguous, mislabelled or genuinely hard example tends to be missed by both. That shared difficulty is exactly a positive covariance, and it is often the dominant term. ## A concrete size for the error Take 10,000 shared test examples, both models at 90% accuracy, disagreeing on 4% of examples. Fill in the paired table: both right on 88%, both wrong on 8%, and 2% each way in the two disagreement cells. Then ``` Cov = 0.88 - 0.9 * 0.9 = 0.07 Var per example, each model = 0.9 * 0.1 = 0.09 Var of the per-example difference = 0.09 + 0.09 - 2 * 0.07 = 0.04 ``` So the paired standard error of the accuracy gap is `sqrt(0.04 / 10000) = 0.002`, or 0.2 percentage points. The unpaired formula gives `sqrt(2 * 0.09 / 10000) = 0.0042`, or 0.42 points. The correlation between the two per-example outcomes is `0.07 / 0.09 = 0.78` - very high, and entirely typical for two models trained on the same corpus. Notice also that the variance of the per-example difference, 0.04, is exactly the disagreement rate. That is the general identity: for accuracy, the paired variance is driven only by how often the models disagree. Two models that agree on 99% of examples can have a tiny standard error on their gap even on a modest test set. ## Conservative is not the same as safe An unpaired test on paired data does not usually inflate your false-positive rate - overstating the standard error makes it harder, not easier, to declare a difference. It is conservative. But it is the wrong kind of caution: it makes real improvements invisible, so teams conclude "no significant difference" from evaluations that would have resolved the gap easily. Reviewers care about this because the cost lands as silently rejected good work rather than as a visible error. ## What to do instead Work with the **per-example difference** as the unit of analysis: - For accuracy or any 0/1 metric, the differences are `+1`, `0`, `-1`, and the whole comparison collapses to the two disagreement counts - which is what McNemar's test consumes. - For a continuous per-example metric such as log-loss or squared error, take `d_i = loss_A(i) - loss_B(i)` and analyse the mean of the `d_i`, either with a resampling procedure or by reporting its interval directly. - For a metric that does not decompose per example, resample the test examples once and score **both** models on each identical resample, so the pairing survives the resampling. In all three, the rule is the same: one shared random draw, then a difference computed inside it. ## Overlapping intervals prove nothing A related trap: plotting a confidence interval for each model's accuracy and concluding "no difference" because the intervals overlap. Overlap of two individual intervals is neither necessary nor sufficient for the difference to be significant, and it is especially misleading when the two estimates are positively correlated - which is exactly the situation here. Build the interval on the **gap**, and check whether that interval excludes zero. ## When pairing does not help Pairing helps because the correlation is positive. If the two models had negatively correlated per-example outcomes - one succeeding precisely where the other fails - the variance of the difference would be *larger* than the unpaired formula suggests, and ignoring the covariance would be anti-conservative. That case is rare between two versions of the same system but is worth naming, because it shows you understand the covariance term rather than reciting a rule. Finally, if the two models were evaluated on *different* test sets, none of this applies: the pairing is gone, and any gap now confounds model quality with test-set difficulty. That is a design flaw to fix, not a statistic to correct.
- Can pairing ever make a difference harder to detect rather than easier?Yes, if the per-example outcomes are negatively correlated - each model succeeding where the other fails. Then the covariance term is negative and the variance of the difference exceeds what the unpaired formula gives. For two versions of the same system trained on the same data this is rare; the correlation is usually strongly positive, so pairing almost always buys sensitivity.
- Two models' individual accuracy intervals overlap. Does that mean the gap is not significant?No. Overlap of two individual intervals is neither necessary nor sufficient for the difference to be significant, and it is especially misleading when the estimates are positively correlated. Build a single interval on the difference itself and check whether it excludes zero.
- What breaks if the two models were scored on different test sets?The pairing disappears entirely, so there is no covariance left to exploit, and worse, the gap now mixes model quality with test-set difficulty. You cannot tell whether one model is better or its test set was easier. Fix the evaluation design - re-score both on one shared set - rather than reaching for a different statistic.
Timing two runners in the same race tells you far more about who is faster than timing them on different days on different courses, where the weather does half the talking.
saying these in an interview costs you the question
- Runs a two-proportion test on scores from one shared test set
- Says the two accuracy estimates are independent because the models are
- Concludes no difference because the two intervals overlap
- Thinks pairing only applies to repeated measurements on people
- Shrugs off the power loss because the unpaired test is conservative