skip to content

How do you put a confidence interval on the AUC difference between two models?

level: seniorimportance: nice to knowfreq 23%

answer

  1. not a mean over examples
  2. resample examples, not predictions
  3. one resample, both models scored on it
  4. percentiles of the recorded differences
  5. DeLong supplies the paired covariance

basics

~20 s

Resample the test examples with replacement, recompute both models' AUC on each identical resample, and take percentiles of the differences. DeLong's method is the analytic alternative, giving a closed-form variance for the difference between two correlated AUCs.

solid answer

~50 s

AUC is not an average of per-example terms - it is built from comparisons between positive and negative examples - so you cannot form one difference per example and work with those. The paired bootstrap does work: draw `n` test examples with replacement, score **both** models on that single resample, record `AUC_A - AUC_B`, repeat a couple of thousand times, and read the 2.5th and 97.5th percentiles of the recorded differences. Resampling once and scoring both models on it is what preserves the pairing; two independent resamples would reintroduce exactly the independence you are trying to avoid. Stratify the resample by class when positives are scarce, so each resample keeps enough of them for AUC to be stable. DeLong's method is the analytic route: it estimates the covariance between two correlated AUCs directly and yields a closed-form interval for their difference.

go deeper

for a junior

Know that any reported metric gap deserves an interval, and that for a ranking metric the practical way to get one is to resample the test examples rather than to compute a formula by hand.

for a middle

Explain why this metric does not decompose into per-example terms, and describe the resampling loop precisely - one resample per iteration, both models scored on it, percentiles of the differences.

for a senior

Show the operational details: stratified resampling under class imbalance, DeLong as the cheap analytic alternative and where its normal approximation strains, and an explicit statement of what the interval excludes.

for a principal

Decide what evidence standard a model swap requires: whether an interval over test-example sampling alone is sufficient, or whether seed variance and evaluation-set refresh policy belong in the bar your team commits to.

## Why this metric needs its own treatment For accuracy or log-loss, each test example contributes one number to each model, so the comparison collapses to per-example differences and the paired machinery is straightforward. AUC does not decompose that way: it is computed from comparisons between positive and negative examples, so a single example participates in many comparisons and there is no well-defined "AUC of example `i`". That rules out the per-example approaches and forces you to resample at the level of examples while recomputing the whole metric. ## The paired bootstrap, step by step 1. Fix the held-out set of `n` examples, with both models' scores already computed for each. 2. Draw `n` examples with replacement to form one resample. 3. Compute `AUC_A` and `AUC_B` **on that same resample** and record `delta = AUC_A - AUC_B`. 4. Repeat for `B` resamples - two thousand is a reasonable default for an interval. 5. Report the 2.5th and 97.5th percentiles of the recorded `delta` values as a 95% interval for the difference. The load-bearing detail is step 3. Because both models are evaluated on the identical resample, every draw carries the same examples for both, and the shared difficulty cancels inside each `delta`. If you instead drew a separate resample per model, the two AUCs would vary independently, the interval on the difference would be much wider, and you would have thrown away the whole benefit of a shared test set. ## Stratifying by class AUC depends on having both positives and negatives present. When positives are rare - a few dozen in a large test set - an unstratified resample can contain very few, making the recomputed AUC unstable or, in the extreme, undefined. Resampling positives and negatives separately, keeping each class count fixed at its observed value, avoids that and is standard practice for imbalanced evaluation sets. Mention this and you signal that you have run the procedure rather than read about it. ## DeLong's method The analytic alternative estimates the variance of each AUC and, crucially, the **covariance** between two AUCs computed on the same data, from which a variance for the difference follows: ``` Var(AUC_A - AUC_B) = Var(AUC_A) + Var(AUC_B) - 2 * Cov(AUC_A, AUC_B) ``` A normal-based interval on the difference then follows immediately. Its appeal is that it is deterministic and cheap - no resampling loop, no Monte Carlo noise - and its covariance term is exactly the paired structure that the naive treatment of two independent AUCs would drop. Its cost is that it leans on a large-sample normal approximation, which is shakier when the test set is small or one class is very thin. The two approaches usually agree closely on reasonably sized sets; the bootstrap is the safer default when they do not. ## Do not compare two individual intervals A common error is to draw an interval around each model's AUC and check whether they overlap. Overlap is neither necessary nor sufficient for the difference to be significant, and it is especially misleading here because the two AUCs are strongly positively correlated - the interval on the *difference* is typically far narrower than the individual intervals suggest. Build the interval on the gap and see whether it excludes zero. ## What the interval covers, and what it does not A bootstrap over test examples quantifies one source of randomness: which examples ended up in the test set. It does not cover training-seed or data-order variance - retrain either model with a different seed and its AUC will move by an amount this interval knows nothing about. It does not cover the effect of having chosen either model by looking at this same test set. And it does not cover any difference between the test distribution and the one the model will meet in the future. When an offline gap is small, seed variance is frequently the larger of the two effects, and a defensible comparison retrains each model a handful of times and reports that spread alongside the interval. ## Reporting Write the interval on the gap, not two numbers with a winner circled: "model A's AUC is 0.008 higher, 95% interval 0.001 to 0.015, from a paired bootstrap over 2,000 resamples of the test examples, stratified by class." That sentence states the effect, its uncertainty, the procedure and the scope, which is exactly what a reviewer needs in order to disagree with you productively.

  • Why not build a separate interval for each AUC and check whether they overlap?
    Overlap of two individual intervals is neither necessary nor sufficient for a significant difference, and it is badly misleading when the two estimates are positively correlated, as two models scored on one test set always are. The interval you need is on the difference itself, computed so the pairing survives; it is typically much narrower than the individual intervals imply.
  • Why stratify the bootstrap resample by class?
    AUC needs both positives and negatives, and when positives are rare an unstratified resample can contain very few, making the recomputed value unstable or undefined. Resampling within each class, keeping the observed class counts fixed, removes that failure mode and reduces the extra variance that random class composition would otherwise inject into the interval.
  • What uncertainty does a bootstrap over test examples fail to capture?
    Only test-example sampling is covered. Training-seed and data-order variance, any selection you did while looking at this same test set, and any shift between the test distribution and future data are all outside it. For a small gap, seed variance is often larger than the interval you just computed, so retraining each model a few times is worth the cost.

saying these in an interview costs you the question

  • Bootstraps each model on its own separate resample of the test set
  • Checks whether the two individual AUC intervals overlap
  • Tries to sign-flip per-example differences for a ranking metric
  • Uses an unstratified resample when positives are very rare
  • Treats the interval as covering retraining and deployment variability

context