skip to content

Why does scikit-learn's roc_auc_score need scores rather than predict() labels?

level: middleimportance: must knowfreq 64%

answer

  1. area under a curve of thresholds
  2. ranking, not a single decision
  3. hard labels give one interior point
  4. probability column or decision margin
  5. invariant to monotone transforms

basics

~20 s

ROC-AUC measures how well a continuous score ranks positives above negatives, so it needs predict_proba(X)[:, 1] or decision_function(X). Hard 0/1 labels from predict() collapse the curve to one threshold and silently return a lower, meaningless number instead of raising.

solid answer

~50 s

`roc_auc_score(y_true, y_score)` integrates the ROC curve, which is traced by sweeping a threshold across a continuous score. So `y_score` must be a ranking signal: `clf.predict_proba(X)[:, 1]` for probabilistic estimators, or `clf.decision_function(X)` for margin-based ones like `SVC`. Any strictly increasing transform of the score gives the same AUC, because only the ordering matters. The trap is that passing `clf.predict(X)` does not raise — hard labels are a valid numeric array, so scikit-learn computes the area under a curve with a single interior point, which is exactly balanced accuracy at that one threshold. Reported AUCs of 0.7 that should be 0.9 usually come from this. For multiclass targets you must pass the full `predict_proba` matrix, whose columns follow the estimator's `classes_` order, and set `multi_class='ovr'` or `'ovo'`; the default `multi_class='raise'` refuses. `average_precision_score` and `precision_recall_curve` take the same kind of score array.

code

python · 10 lines
python
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score

X, y = make_classification(n_samples=500, weights=[0.9], random_state=0)
clf = LogisticRegression(max_iter=1000).fit(X, y)

print(roc_auc_score(y, clf.predict(X)))
print(roc_auc_score(y, clf.predict_proba(X)[:, 1]))
print(roc_auc_score(y, clf.decision_function(X)))

go deeper

for a junior

Remember that ROC-AUC needs probabilities or decision-function values, not the output of predict(). Knowing to write predict_proba(X)[:, 1] for a binary classifier covers most of what is asked here.

for a middle

Explain why the function does not raise on hard labels and what the resulting number actually equals, and describe the multiclass requirements: full probability matrix plus multi_class='ovr' or 'ovo'.

for a senior

Diagnose a suspicious AUC from its symptoms — a value near 0.5 from label input, a value below 0.5 from an inverted positive class — and argue when AUC is the wrong headline metric because ranking quality is not the operational concern.

for a principal

Set the convention for how model quality is reported across teams, including whether ranking metrics or calibration metrics are the contract, so that AUC numbers from different projects mean the same thing and are not quietly computed from labels.

## What the function is integrating A ROC curve is not a property of a set of predicted labels; it is a property of a *ranking*. You take a continuous score per sample, sweep a decision threshold from above the maximum score down to below the minimum, and at each threshold record the true-positive rate and false-positive rate produced by calling everything above the threshold positive. The curve is the trace of those points, and `roc_auc_score` is its area. Equivalently, the AUC is the probability that a randomly chosen positive sample receives a higher score than a randomly chosen negative one. That definition is why the input must be a score. If you have only hard labels, there is no ranking to sweep — every sample is already at 0 or 1. ## Where the score comes from in scikit-learn The estimator API gives you two response methods. Probabilistic classifiers (`LogisticRegression`, `RandomForestClassifier`, `GaussianNB`, gradient boosting) expose `predict_proba(X)`, returning shape `(n_samples, n_classes)`. For a binary problem you pass column 1 — but 'column 1' means the column corresponding to `clf.classes_[1]`, not the column you assume is positive. If your labels are strings, check `clf.classes_` before slicing, because indexing the wrong column gives you `1 - AUC`, and an AUC of 0.13 is usually a positive-class mix-up, not a broken model. Margin-based classifiers expose `decision_function(X)`, a signed distance from the boundary with no probability calibration. That is fine: AUC is invariant to any strictly monotonic transform, so raw margins, log-odds and calibrated probabilities all give the identical AUC. This also means ROC-AUC tells you nothing about calibration; a model can have AUC 0.95 and probabilities that are systematically far from the truth. If calibration is the question, `log_loss` or `brier_score_loss` is the metric, not AUC. ## The silent failure Passing `clf.predict(X)` is the classic mistake because it fails quietly. The array is numeric and one-dimensional, so scikit-learn happily computes an ROC curve with exactly three points: (0,0), the single (FPR, TPR) pair produced by the model's own threshold, and (1,1). The area under that trapezoid works out to (TPR + TNR) / 2 — precisely balanced accuracy. So the number is not garbage, it is a *different metric wearing AUC's name*, and it is almost always noticeably worse than the real AUC. Two symptoms give it away: an AUC suspiciously close to a balanced-accuracy figure you already computed, and an AUC that does not move when you change the decision threshold — which is impossible for a genuine AUC, since AUC is threshold-independent by construction. ## Multiclass The signature includes `multi_class`, defaulting to `'raise'`. For more than two classes you must pass `y_score` of shape `(n_samples, n_classes)` and choose `'ovr'` (one class against the rest) or `'ovo'` (all class pairs), combined with `average='macro'` or `'weighted'`. Column order follows `classes_`, and for `'ovr'` scikit-learn expects the rows to sum to one, so pass raw `predict_proba` output rather than a hand-assembled matrix. Passing a two-column `predict_proba` matrix for a *binary* target is also rejected — binary wants the single positive-class column. ## The neighbouring functions `roc_curve(y_true, y_score)` returns `fpr, tpr, thresholds` if you want the curve itself, and `auc(fpr, tpr)` integrates it — `roc_auc_score` is the convenience wrapper over both. `RocCurveDisplay.from_estimator` handles the plotting and picks the right response method for you, which is a good reason to prefer it in notebooks. `average_precision_score` and `precision_recall_curve` take the same score array and are the imbalanced-data counterparts. ## Inside cross-validation When you pass `scoring='roc_auc'` to a search or `cross_val_score`, the built-in scorer selects `predict_proba` or `decision_function` for you, so this whole class of bug disappears — one of the better arguments for using the scorer strings instead of computing metrics by hand on a held-out fold.

  • You get an AUC of 0.12 from a model that clearly learned something. What happened?
    Almost certainly the positive class is inverted. Either the wrong `predict_proba` column was sliced, or `pos_label` does not match what the score array ranks high. An AUC of 0.12 means the ranking is 0.88 correct in the opposite direction, so check `clf.classes_` and slice the column matching the intended positive label rather than assuming index 1.
  • Why does the AUC stay identical when you apply a sigmoid or a log to the score array?
    AUC depends only on the relative ordering of scores, since it is the probability that a random positive outranks a random negative. Any strictly increasing transform preserves that ordering, so the curve's shape in FPR/TPR space is unchanged. The practical consequence is that AUC cannot detect miscalibration, which is why log_loss or brier_score_loss is used when probability quality matters.
  • What does roc_auc_score require differently for a five-class problem?
    Pass the full `predict_proba` matrix of shape (n_samples, 5) and set `multi_class='ovr'` or `'ovo'`; the default `'raise'` rejects multiclass input outright. Columns must follow the estimator's `classes_` order, and `average` then controls whether the per-class or per-pair AUCs are combined unweighted ('macro') or by support ('weighted').

saying these in an interview costs you the question

  • Passing clf.predict(X) into roc_auc_score and trusting the result
  • Assuming the function raises when given hard labels
  • Slicing predict_proba column 1 without checking classes_
  • Claiming AUC proves the probabilities are well calibrated
  • Expecting binary roc_auc_score to accept a two-column probability matrix

context