skip to content

What does an ROC-AUC of 0.78 mean in probability terms for a display-ad click model?

level: middleimportance: must knowfreq 76%

answer

  1. pick one of each class
  2. compare the two scores
  3. probability the positive ranks higher
  4. ties are worth half a win
  5. same as the normalised rank-sum statistic

basics

~20 s

Draw one clicker and one non-clicker at random: the model scores the clicker higher 78% of the time, counting ties as half a win. ROC-AUC is a statement about ranking order, not about the numeric size of the scores.

solid answer

~50 s

The area under the ROC curve has a clean probabilistic reading: `AUC = P(score of a random positive > score of a random negative) + 0.5 * P(tie)`. So 0.78 for a click model means that if you sample one user who clicked and one who did not, the clicker gets the higher score in 78% of such pairs. That is a pure statement about ordering. It is the same number as the normalised Mann-Whitney U statistic: rank all scores from lowest to highest, sum the ranks of the positives to get `R1`, then `AUC = (R1 - n1*(n1+1)/2) / (n1*n0)` where `n1` and `n0` are the positive and negative counts. Because it depends only on ranks, any strictly increasing transform of the scores leaves AUC untouched. 0.5 is the coin-flip diagonal, 1.0 is perfect separation, and below 0.5 means the ranking runs backwards.

code

python · 21 lines
python
scores = [0.9, 0.8, 0.7, 0.6, 0.55, 0.4, 0.3, 0.2]
labels = [1,   0,   1,   1,   0,    0,   1,   0]

pos = [s for s, y in zip(scores, labels) if y == 1]
neg = [s for s, y in zip(scores, labels) if y == 0]

# 1) pairwise wins: every positive plays every negative, a tie is half a win
wins = sum((p > n) + 0.5 * (p == n) for p in pos for n in neg)
auc_pairs = wins / (len(pos) * len(neg))

# 2) Mann-Whitney form: rank ascending, sum the positives' ranks
order = sorted(range(len(scores)), key=lambda i: scores[i])
rank = [0] * len(scores)
for r, i in enumerate(order, start=1):
    rank[i] = r

n1, n0 = len(pos), len(neg)
rank_sum = sum(rank[i] for i in range(len(scores)) if labels[i] == 1)
auc_rank = (rank_sum - n1 * (n1 + 1) / 2) / (n1 * n0)

print(auc_pairs, auc_rank)

go deeper

for a junior

Memorise and be able to say the pairwise sentence: a random positive outranks a random negative this often. Also know the anchors - 0.5 is random, 1.0 is perfect - and that AUC is not accuracy.

for a middle

Explain why the area equals that probability and connect it to the rank-sum statistic. Be ready to show that only score order matters, so any monotone transform leaves the number unchanged.

for a senior

Use the interpretation as a debugging tool. A sub-0.5 AUC points at an inverted label or score path; a near-1.0 AUC on a hard problem points at leakage; a flat 0.5 points at a constant or broken feature pipeline.

for a principal

Decide when a single ranking scalar is the right thing for a team to report at all. Push for the pairwise sentence in reviews rather than the bare number, so stakeholders stop hearing AUC as accuracy.

## From area to probability The ROC curve plots true positive rate against false positive rate over all thresholds. The **area under it** collapses that whole curve to one number between 0 and 1. The reason that number is worth memorising is its exact probabilistic meaning: ``` AUC = P(s_pos > s_neg) + 0.5 * P(s_pos = s_neg) ``` where `s_pos` is the score of a uniformly random positive example and `s_neg` the score of a uniformly random negative one, drawn independently. In interview register for a display-ad click model: **"AUC 0.78 means that if I pick a random clicker and a random non-clicker, my model gives the clicker the higher score 78 times out of 100."** That sentence, said out loud, is what the question is fishing for. The intuition behind the identity: sweeping the threshold down past a negative example moves the operating point right by `1/n0`, and the height of the curve at that moment is the fraction of positives already above the threshold - that is, the fraction of positives that beat this negative. Summing height times width over all negatives averages "fraction of positives that outrank this negative" over the negatives, which is precisely the pairwise win rate. ## The Mann-Whitney connection Count, over all `n1 * n0` positive-negative pairs, how many the positive wins (ties counting half). That count is the **Mann-Whitney U statistic**, the same statistic behind the Wilcoxon rank-sum test. AUC is simply U normalised by the number of pairs: ``` U = R1 - n1*(n1 + 1)/2 # R1 = sum of the positives' ranks, ranked ascending AUC = U / (n1 * n0) ``` Three consequences follow immediately. 1. **AUC is a rank statistic.** Only the order of the scores enters. Multiply every score by 10, take its logarithm, or squash it through any strictly increasing function: the AUC does not move by a thousandth. This is why AUC says nothing about whether a score of 0.30 corresponds to a 30% event rate - that is a separate property of the scores entirely. 2. **It is non-parametric.** No assumption about the shape of either class's score distribution is needed. 3. **Ties are handled explicitly.** A model that outputs the same constant for everyone has every pair tied and scores exactly 0.5 - the same as random guessing, which is the right answer. ## Reading the number - **1.0** - perfect separation: every positive outranks every negative. In practice this usually means a leaked feature rather than a great model. - **0.5** - the diagonal. The score orders positives and negatives no better than a coin flip. - **Between 0.5 and 1.0** - the normal range. Roughly, 0.6 is weak-but-real signal, 0.7-0.8 is a workable production ranker for a noisy behavioural target like ad clicks, and above 0.95 on a hard problem deserves a leakage audit before celebration. - **Below 0.5** - the ranking is systematically *backwards*. A product-returns model reporting **0.43** is not "useless": it is informative, just inverted, and flipping the sign of the score would give 0.57. Almost always the cause is mechanical - the positive class was defined one way when training and the other way when scoring, the negative-class score was compared against positive labels, or a sign was flipped somewhere in the scoring path. Only on a tiny evaluation set is sampling noise alone a plausible explanation for a dip meaningfully below 0.5. ## What the number does not say AUC is a **ranking** summary and nothing else. - It does not name a threshold. Two teams shipping the same model at different cut-offs share an AUC and have completely different error profiles. - It does not report error counts. It is built from within-class rates, so it does not tell you how many alerts an operating point generates. - It does not say the scores are meaningful as probabilities. A perfectly ordered set of scores that are all between 0.01 and 0.02 still has AUC 1.0. - It weights the whole curve, including false-positive-rate regions your product may never operate in. ## Saying it well The strongest version of the answer in an interview does four things in about twenty seconds: gives the pairwise-probability sentence, names the Mann-Whitney equivalence, notes the invariance to monotone transforms, and states what 0.5 and sub-0.5 mean. Candidates who instead say "AUC 0.78 means the model is 78% accurate" have made the single most common error on this topic - accuracy depends on a threshold and on the class balance, and AUC depends on neither.

  • How exactly does AUC relate to the Mann-Whitney U statistic?
    U counts the positive-negative pairs the positive wins, ties counting half, and AUC is U divided by the total number of pairs, `n1 * n0`. Computationally you rank all scores ascending, sum the positives' ranks to get `R1`, then `U = R1 - n1*(n1+1)/2`. So AUC is a normalised rank-sum statistic, which is why it is non-parametric and depends only on ordering.
  • A product-returns model reports an AUC of 0.43 on a large evaluation set. What do you check first?
    Below 0.5 means positives systematically rank *below* negatives, so the signal is inverted rather than absent - flipping the score would give 0.57. Check the mechanics before the modelling: which class was treated as positive at scoring time versus training time, whether the negative-class score was scored against positive labels, and whether a sign was flipped. Genuine anti-correlation with no bug is rare on a large set.
  • If you apply a strictly increasing transform to every predicted score, what happens to AUC?
    Nothing. AUC depends only on the ordering of the scores, and a strictly increasing transform preserves order, so every pairwise comparison and every rank is unchanged. Rescaling, log-transforming, or passing scores through a monotone calibration map all leave AUC exactly where it was - which also means a rising AUC can never be credited to rescaling.

It is a round-robin tournament between the two classes. Every positive plays every negative, and the higher score wins the match. AUC is simply the positives' overall win rate, with draws worth half a point.

saying these in an interview costs you the question

  • Says AUC of 0.78 means the model is 78% accurate
  • Calls AUC the probability that a prediction is correct
  • Insists AUC below 0.5 is impossible or a computation bug only
  • Claims rescaling the scores raises AUC
  • Treats a high AUC as proof the scores are meaningful probabilities

context