skip to content

A fraud model shows 0.97 ROC-AUC but poor live precision — which sklearn.metrics calls do you reach for?

level: seniorimportance: should knowfreq 40%

answer

  1. ranking summary versus operating point
  2. the rare-class denominator problem
  3. one curve summary is prevalence-sensitive
  4. three arrays, one is shorter
  5. re-measure at your own cut, not 0.5

basics

~20 s

Score the ranking with average_precision_score, then call precision_recall_curve to choose an operating threshold: it returns precision and recall arrays one element longer than thresholds. Rebuild confusion_matrix at that threshold instead of trusting predict()'s default 0.5.

solid answer

~40 s

Two different things are being measured. `roc_auc_score` summarises ranking across every threshold and is insensitive to how rare the positive class is; live precision is one *operating point* on a prevalence-heavy stream. So I switch to the prevalence-sensitive summary, `average_precision_score(y_true, y_score)`, and then to `precision_recall_curve(y_true, y_score)`, which returns `precision, recall, thresholds` — with `thresholds` one element **shorter**, because a trailing (precision=1, recall=0) point has no threshold. Aligning them with `precision[:-1]` is the detail people get wrong. From that curve I pick the threshold meeting the precision floor the reviewers can absorb, then call `confusion_matrix(y_true, scores >= t)` and `classification_report` at that threshold rather than at `predict()`'s implicit 0.5. One caution: summarise the PR curve with `average_precision_score`, not `auc(recall, precision)` — the trapezoidal integration is optimistic on a curve that moves in steps.

code

python · 15 lines
python
import numpy as np
from sklearn.metrics import (average_precision_score, confusion_matrix,
                             precision_recall_curve)

def pick_threshold(y_true, scores, min_precision=0.5):
    precision, recall, thresholds = precision_recall_curve(y_true, scores)
    ok = np.where(precision[:-1] >= min_precision)[0]
    return thresholds[ok[np.argmax(recall[:-1][ok])]]

y_true = np.array([0, 0, 1, 0, 1, 0, 0, 1])
scores = np.array([0.10, 0.20, 0.90, 0.35, 0.60, 0.40, 0.05, 0.55])

t = pick_threshold(y_true, scores)
print(average_precision_score(y_true, scores), t)
print(confusion_matrix(y_true, (scores >= t).astype(int)))

go deeper

for a junior

Know that ROC-AUC is threshold-free while precision describes one chosen cut, and that precision_recall_curve plus average_precision_score are the functions to reach for on rare-positive problems.

for a middle

Explain the array-length mismatch precision_recall_curve returns and how to align it, and show that predict() applies its own fixed threshold so metrics must be recomputed at the cut you selected.

for a senior

Drive the whole diagnosis: separate the ranking question from the operating-point question, defend average_precision_score over trapezoidal auc, and convert the confusion matrix at the chosen threshold into review load and missed-case cost.

for a principal

Own the operating-point policy — who sets the precision floor, how the threshold is re-tuned as prevalence drifts, and which numbers appear in a launch review — so that model quality is argued in terms of downstream capacity rather than a single curve summary.

## Why the two numbers disagree The report is not contradictory. `roc_auc_score` integrates true-positive rate against false-positive rate over all thresholds. The false-positive rate has the negative count in its denominator, so when negatives outnumber positives a thousand to one, thousands of false alarms barely move it. Precision has the *predicted-positive* count in its denominator, so those same false alarms dominate it. A model can rank almost perfectly and still hand a review queue mostly noise, and no operating threshold is even implied by an AUC figure. In `sklearn.metrics` terms: you asked a threshold-free ranking question and got a threshold-free answer. ## The summary to switch to `average_precision_score(y_true, y_score)` summarises the precision-recall curve instead. It takes the same continuous score array — `predict_proba(X)[:, 1]` or `decision_function(X)` — and returns the precision-weighted mean over the recall steps. Its baseline is the positive-class prevalence rather than the fixed 0.5 of ROC-AUC, so it moves when the class balance moves, which is exactly the sensitivity you wanted. A specific caution: do not compute this as `auc(recall, precision)`. The generic `auc` helper applies trapezoidal integration, which linearly interpolates between adjacent curve points. The PR curve advances in discrete steps as samples cross the threshold, and interpolating across those steps is optimistic — scikit-learn's own documentation flags this. `average_precision_score` uses the step-wise summation and is the one to quote. ## Reading precision_recall_curve correctly `precision_recall_curve(y_true, y_score)` returns three arrays: `precision`, `recall`, `thresholds`. The first two have length `n + 1`; `thresholds` has length `n`. The extra point is the degenerate endpoint at recall 0 and precision 1, which corresponds to predicting nothing positive and therefore has no threshold behind it. Any code that zips the three arrays together, or indexes `thresholds` with an index taken from `precision`, is off by one at the tail. The correct alignment is `precision[:-1]` and `recall[:-1]` against `thresholds`. This off-by-one is one of the most reliable senior-level details on this whole topic, precisely because it produces a plausible wrong threshold rather than an exception. `roc_curve` returns `fpr, tpr, thresholds`, all the same length, with a synthetic first threshold above the maximum score — a different convention, so do not carry the habit across. ## Choosing and then measuring at the operating point Once you have the curve, threshold selection is a constrained pick: the largest recall whose precision clears the floor your downstream process can absorb, or the highest precision that still catches enough cases, depending on which side is scarce. Then — and this is the step most often skipped — you must *re-measure at that threshold*. `clf.predict(X)` does not use your threshold; for probabilistic classifiers it takes the argmax of `predict_proba`, which in binary is the fixed 0.5 cut, and for margin classifiers the sign of `decision_function`, which is 0. So every label metric computed from `predict()` describes a threshold you did not choose. The re-measurement is `confusion_matrix(y_true, (scores >= t).astype(int))` and `classification_report(y_true, (scores >= t).astype(int))`. The confusion matrix at the chosen threshold is what actually translates into queue volume: false positives are reviewer hours, false negatives are losses. `precision_score` and `recall_score` on those hard labels then match what operations will see. ## Reporting honestly A good evaluation of this system therefore carries three layers: a ranking summary (`roc_auc_score`) that is comparable across datasets with different prevalence, a prevalence-sensitive summary (`average_precision_score`) that reflects this stream, and the confusion matrix at the deployed threshold that converts to headcount and money. Quoting only the first is what produced the original surprise. `PrecisionRecallDisplay.from_predictions` is a convenient way to put the curve in front of stakeholders alongside those numbers. ## The evaluation-data caveat One more thing to check before blaming the metric: the threshold must be chosen on data whose class balance matches production. A threshold tuned on a downsampled or class-balanced validation set will not deliver the same precision on a live stream where positives are a hundred times rarer, because precision depends on prevalence while the score distribution does not shift with it. Tune the threshold on a held-out sample with realistic prevalence, and re-check it as prevalence drifts.

  • Why does precision_recall_curve return a thresholds array one element shorter than precision and recall?
    The curve includes a terminal point at recall 0 with precision 1, representing the degenerate classifier that predicts nothing positive. No finite threshold produces it, so it has no entry in `thresholds`. Align with `precision[:-1]` and `recall[:-1]`. Note `roc_curve` uses a different convention — all three arrays are equal length, with a synthetic leading threshold above the maximum score.
  • What is wrong with computing PR-AUC as auc(recall, precision)?
    `auc` integrates trapezoidally, linearly interpolating between adjacent points. The precision-recall curve advances in steps as individual samples cross the threshold, and interpolating across those steps overstates the area. `average_precision_score` sums the step-wise contributions instead and is the summary scikit-learn recommends, which is also what the built-in 'average_precision' scorer string uses.
  • You picked a threshold on a class-balanced validation set. Why might live precision still disappoint?
    Precision depends on prevalence, unlike recall or the score distribution. Downsampling negatives raises apparent precision at every threshold, so a cut that gives 0.6 precision on a balanced set can give far less on a stream where positives are a hundred times rarer. Tune the threshold on data with production prevalence, or reweight, and monitor precision as prevalence drifts.
  • After choosing a threshold, why can't you report metrics computed from clf.predict(X)?
    `predict()` ignores your threshold. For probabilistic classifiers it takes the argmax of `predict_proba`, which in the binary case is a fixed 0.5 cut, and for margin-based ones the sign of `decision_function`. Every label metric derived from it therefore describes a different operating point. Recompute with `(scores >= t).astype(int)` before quoting precision, recall or the confusion matrix.

saying these in an interview costs you the question

  • Treating high ROC-AUC as proof the operating point is usable
  • Zipping precision, recall and thresholds as equal-length arrays
  • Summarising a PR curve with auc(recall, precision)
  • Reporting precision from predict() after choosing a custom threshold
  • Tuning the threshold on class-balanced data and shipping it unchanged

context