A fraud model shows ROC-AUC 0.97 but average precision 0.15. Why the gap?
answer
- same recall axis, different second axis
- count the negatives in the denominator
- false positives barely dent a huge FPR denominator
- chance level here is not 0.5
- AP is measured against the prevalence line
basics
~20 sBoth are right; they use different denominators. With 0.2% fraud, tens of thousands of false positives barely move the false positive rate, but they dominate the flagged set and crush precision. Read average precision against a 0.002 baseline, not 0.5.
solid answer
~50 sAssume one million transactions with 2,000 frauds. At a threshold that catches 1,600 of them, suppose 20,000 legitimate transactions are also flagged. The ROC point looks superb: recall 0.80 and false positive rate 20,000/998,000 = 0.02, because the FPR denominator is the enormous pool of negatives. The precision-recall point is grim: 1,600/21,600 = 0.074, because precision's denominator is only what you flagged. ROC-AUC never touches the flagged set's composition, so it stays flattering under heavy imbalance; average precision does nothing else. Neither is broken — but read AP against its own baseline: a random scorer has precision 0.002 at every recall, so AP of 0.15 is about seventy-five times chance, whereas ROC's chance level is 0.5. If the workflow only ever acts on flagged cases, AP is the number that describes the experience of the people acting on them.
go deeper
Recall that both numbers can be correct at once, and that the fraction of positives in the data has to be quoted before any average precision value means anything.
Write out the two formulas and point at the denominators: all negatives for false positive rate, only the flagged rows for precision. Then work a small numeric example out loud.
Demonstrate that you would pick the reported metric from the workflow — alert-queue products live and die on the flagged set — and that you know average precision shifts when the base rate shifts.
Own the reporting standard: decide which number the organisation tracks for rare-event detectors, insist prevalence is published beside it, and prevent teams from claiming wins driven by a moving base rate.
## Two metrics, two denominators The whole discrepancy is arithmetic about what sits under the fraction bar. An ROC curve plots the true positive rate against the false positive rate: - `TPR = TP / (TP + FN)` — the positives you caught, over all positives. Same thing as recall. - `FPR = FP / (FP + TN)` — the negatives you wrongly flagged, over **all negatives**. A precision-recall curve plots precision against recall: - `recall = TP / (TP + FN)` — identical to TPR. - `precision = TP / (TP + FP)` — true positives over **everything you flagged**. The two curves share one axis. They differ entirely in the second one, and the difference is the denominator of the false-positive term: all negatives versus the flagged set. ## The fraud arithmetic Take a month of one million card transactions with a 0.2% fraud rate: 2,000 frauds, 998,000 legitimate. Pick a threshold that catches 1,600 frauds and also flags 20,000 legitimate transactions. - ROC reading: TPR = 1,600/2,000 = 0.80, FPR = 20,000/998,000 = 0.020. The point (0.02, 0.80) sits far into the top-left corner. Sweep the threshold and you can easily accumulate an area of 0.97. - PR reading: recall = 0.80, precision = 1,600/21,600 = 0.074. Of every hundred alerts, seven are fraud. Same model, same threshold, same confusion matrix. Twenty thousand false positives is a rounding error against 998,000 negatives and a catastrophe against 21,600 alerts. Because the negatives outnumber the positives roughly 500 to 1, the false-positive term is divided by a number 500 times larger on the ROC side — so ROC-AUC is structurally optimistic-looking whenever positives are rare, while average precision is not. ## The baselines are not the same The second half of the answer is that 0.97 and 0.15 are scored on different scales. A classifier that assigns scores at random flags a random subset, so its precision equals the positive rate at every recall. Its precision-recall curve is a **flat line at the prevalence** — here 0.002 — and its average precision is 0.002. The same random classifier has an ROC curve on the diagonal and an area of exactly 0.5, whatever the prevalence. So AP = 0.15 is roughly 75x chance, while AUC = 0.97 is roughly 1.94x chance. Stated as lift over the no-skill line, the two metrics are far less contradictory than the raw numbers suggest. A candidate who calls 0.15 "terrible" without asking the base rate has misread the scale; a candidate who calls 0.97 "production ready" without asking about alert volume has misread the workflow. ## Which one to trust It depends on what is done with the scores. - If a downstream process acts only on the flagged cases — an alert queue, a hold, an automated block — then the composition of the flagged set *is* the product experience, and average precision is the honest headline. Nobody in that workflow ever meets the correctly-ignored negatives that prop up FPR. - If the model's job is to rank, and the whole ranked population is used, ROC-AUC is a legitimate summary of ordering quality. The common practice is to report both, always beside the evaluation set's positive rate. Reporting AP without prevalence is like reporting a temperature without units. ## Prevalence dependence and comparability One more consequence, and it is the practical one. ROC-AUC is invariant to the class prior: it is built from TPR and FPR, each computed within one class, so re-weighting how many negatives you sample leaves it roughly unchanged. Average precision is **not** invariant — its baseline moves with prevalence, and so does the metric. That has a sharp operational edge. If the fraud rate drops from 0.2% to 0.1% because an upstream rule started blocking a fraud pattern, average precision will fall on identical model behaviour, while ROC-AUC sits still. The reverse also happens: a fraud wave inflates AP without the model improving. So AP is excellent for judging one model on one population, and dangerous for comparing across periods or segments with different base rates unless you either quote the lift over each period's baseline or re-score both periods at a common prevalence. ## What not to say Do not conclude that ROC-AUC is useless on imbalanced data — it is a valid, prevalence-stable ranking summary. Do not conclude that the model is broken because AP is 0.15 — check the baseline first. And do not conclude the two metrics disagree about the ranking: they are computed from the same ranking and are asking different questions of it.
- Average precision fell from 0.30 to 0.22 across two quarters while ROC-AUC held at 0.96. Is the model worse?Not necessarily. Check the positive rate in each period first. Average precision moves with prevalence, and ROC-AUC does not, so a falling base rate produces exactly this pattern with no change in model behaviour. Compare each AP against its own period's baseline, or re-score both periods at a common prevalence, before calling it degradation.
- Does this mean ROC-AUC should never be reported on imbalanced data?No. It remains a correct summary of how well the model orders positives above negatives, and its prevalence-invariance is genuinely useful for tracking a ranker across periods where the base rate moves. The mistake is presenting it alone as evidence that a rare-event detector is deployable, when nobody downstream ever sees the true negatives that make it look good.
- Why is the no-skill precision-recall baseline the positive rate rather than 0.5?A random scorer flags a random subset of rows, so the fraction of that subset which is truly positive equals the overall positive rate, at every recall. The curve is therefore a flat line at prevalence. The ROC diagonal gives area 0.5 regardless of prevalence, which is why the two metrics' chance levels differ.
You pull 21,600 straws out of a haystack of a million and 1,600 are needles. Measured against the haystack you have barely disturbed it; measured against your handful, most of what you are holding is straw.
saying these in an interview costs you the question
- Declares the model broken because average precision is 0.15
- Declares the model production-ready because ROC-AUC is 0.97
- Compares average precision against 0.5 as chance
- Thinks precision and false positive rate share a denominator
- Assumes the two metrics must agree or one is miscomputed
- Says average precision is comparable across periods with different base rates