skip to content

Why is ROC-AUC unchanged after an evaluation set is downsampled from 3% to 50% positives?

level: seniorimportance: should knowfreq 52%

answer

  1. each rate lives inside one class
  2. denominators never mix the classes
  3. dropping negatives shrinks FP and TN together
  4. pairwise ranking probability has no prior in it
  5. rates hold, alert counts do not

basics

~20 s

True positive rate is computed only among positives and false positive rate only among negatives, so randomly discarding negatives leaves both rates unchanged in expectation. ROC-AUC measures ranking quality and is therefore blind to the class mix.

solid answer

~50 s

Both coordinates of the ROC curve are within-class rates: `TPR = TP / (TP + FN)` uses only actual positives and `FPR = FP / (FP + TN)` uses only actual negatives. Randomly dropping negatives to lift the positive rate from 3% to 50% shrinks FP and TN by the same factor, so their ratio - and the whole curve - is unchanged in expectation. Equivalently, AUC is the probability a random positive outranks a random negative, which depends on the two score distributions, not on how many of each you sampled. That invariance is useful: you can compare AUC across periods whose base rate drifts. It is also a trap, because counts are not invariant. At a fixed threshold false alarms number `FPR * N_neg`, so the same model throws roughly seventeen times more of them in the real 3% population.

go deeper

for a junior

Know the headline: ROC-AUC does not move when you change the ratio of positives to negatives, because its two axes are rates measured inside each class separately.

for a middle

Be able to derive it. Show that dropping negatives at random scales FP and TN by the same factor so FPR is preserved, and that the pairwise-ranking definition contains no class prior at all.

for a senior

Demonstrate the operational consequence: identical rates, very different alert volumes. Also name the conditions under which the invariance breaks - score-correlated subsampling, hard-negative selection, and small-sample noise.

for a principal

Decide what the team reports. Prior-invariance makes AUC the right cross-period ranking tracker and the wrong single health metric, so pair it with count-based operational monitoring rather than letting one number stand in for both.

## The mechanical reason Write down the two rates the ROC curve is made of: ``` TPR = TP / (TP + FN) # denominator = all actual positives FPR = FP / (FP + TN) # denominator = all actual negatives ``` Each one is a **conditional** quantity: the fraction of the positive class above the threshold, and the fraction of the negative class above the threshold. Neither denominator contains a single example from the other class. Now downsample: keep every positive and randomly keep, say, one in seventeen negatives, lifting the positive rate from 3% to 50%. In the retained sample, FP and TN each shrink by roughly the same factor, because the retention was independent of the score. Their ratio is preserved, so FPR at every threshold is unchanged in expectation. TPR is untouched entirely - no positives were removed. Every point of the curve therefore lands where it was, and the area under it is the same. The probabilistic view gives the same answer in one line. `AUC = P(s_pos > s_neg) + 0.5 * P(s_pos = s_neg)`, where the two scores are drawn independently from the positive and negative **score distributions**. That probability is a property of those two distributions. How many draws you take from each - the class prior - never appears in it. ## Where the invariance genuinely helps - **Drifting base rates.** A fraud or churn rate that moves from 3% to 5% across quarters would move threshold-dependent count metrics on its own. AUC tracked over those quarters isolates whether the model's *ranking* got better or worse. - **Cheap evaluation.** On a huge, heavily negative log you can score a random subsample of negatives and get an unbiased AUC estimate far more cheaply than scoring everything. - **Training-set rebalancing.** If you trained on a rebalanced sample and evaluated on the natural distribution, a large AUC gap between the two evaluations points at something other than the class mix - leakage, drift, or a broken split. - **Cross-dataset comparison.** Two teams with differently balanced holdouts can still compare rankers, which threshold-based metrics like accuracy cannot do. ## Where it becomes a trap Invariance to the prior means AUC deliberately throws away the information you need to judge operational load. The distinction to hold onto is **rates versus counts**. At a fixed threshold with false positive rate `f`, the number of false alarms is `f * N_neg`. Take a model operating at `FPR = 0.05`. On a balanced 10,000-example set that is 250 false alarms against 5,000 positives. On the real 3%-positive population of the same size, the same threshold gives about `0.05 * 9,700 = 485` false alarms against only 300 positives - the flagged pile is now overwhelmingly negatives, while TPR, FPR and AUC all read exactly the same. Nothing about the model changed; the population did. The share of flagged cases that are genuinely positive moves with the prior even though the curve does not. So an unchanged AUC after a prevalence shift is **not** evidence that the deployed system still feels the same to whoever works the alerts. It is evidence only that the ordering is as good as it was. ## When the invariance actually breaks The result holds **in expectation, under sampling of one class that is independent of the score**. Three ways it fails: 1. **Non-random subsampling.** Dropping "easy" negatives, hard-negative mining, or filtering by a feature correlated with the score changes the negative score distribution itself, and the curve moves - usually AUC falls, because the surviving negatives are harder. 2. **Sampling noise.** The invariance is an expectation. Downsampling to a few hundred negatives leaves the AUC estimate unbiased but much noisier, so a single measurement can wander. 3. **Time or segment shifts disguised as prior shifts.** If the base rate moved because the population changed - a new traffic source, a new market - the score distributions moved too, and the AUC change you see is real, not an artefact. ## Answering it well A strong answer has three beats. First the mechanism: both coordinates are within-class rates, so the class mix cancels. Second the equivalent framing: AUC is a pairwise ranking probability over two score distributions, and the prior is not in that expression. Third the caveat, unprompted: prior-invariance is a property, not a virtue - it is exactly why AUC alone cannot tell you what an operating point will cost you when the population is 3% positive. Candidates who only give the first beat sound like they memorised it; the third beat is what marks production experience.

  • Under what kind of subsampling does AUC actually change?
    When the sampling is not independent of the score. Dropping easy negatives, hard-negative mining, or filtering on a feature correlated with the score reshapes the negative score distribution, so the curve genuinely moves - usually downward, because the survivors are harder to separate. Small samples also add variance: the estimate stays unbiased but a single measurement can wander well away from the population value.
  • If AUC is prior-insensitive, why does the same model feel much worse when the real base rate is 3%?
    Because rates are invariant but counts are not. At a fixed false positive rate `f`, false alarms number `f * N_neg`, and there are seventeen times more negatives at a 3% prior than at 50%. The ranking is identical, but the flagged pile fills up with negatives, so the share of flagged cases that are truly positive collapses even though the ROC curve did not move.
  • Does an unchanged AUC across two quarters prove the model has not degraded?
    No. It proves the ranking quality is stable in aggregate. Score distributions may have shifted so the same threshold now sits at a different operating point, a subpopulation may have degraded while another improved, and the volume of negatives may have grown. Prior-invariance means AUC deliberately hides exactly those effects, so it should be tracked alongside count-based operational monitoring.

saying these in an interview costs you the question

  • Says AUC falls automatically as data becomes more imbalanced
  • Claims downsampling negatives is what raised the AUC
  • Assumes a high AUC guarantees a tolerable false-alarm volume
  • Thinks AUC accounts for the cost of a false positive
  • Expects accuracy and AUC to move together when the base rate shifts

context