skip to content

Metrics & Evaluation

Picking the right metric is often the real interview question: accuracy collapses on imbalanced data, precision and recall trade off, and ROC-AUC and F1 answer different things. Covers classification, regression, and clustering metrics plus custom scorers.

on this pageshow

questions

6

In scikit-learn, how do you read the output of confusion_matrix?

level: juniorimportance: must knowfreq 58%

answer

  1. rows and columns mean different things
  2. true on one axis, predicted on the other
  3. diagonal is what went right
  4. ordering comes from sorting, not appearance
  5. tn, fp, fn, tp flattening for 0/1

basics

~20 s

confusion_matrix returns a square array where rows are true classes and columns are predicted classes, so C[i, j] counts samples of true class i predicted as class j. Classes appear in sorted order unless you pass labels= to fix the order.

solid answer

~50 s

`confusion_matrix(y_true, y_pred)` returns an `n_classes x n_classes` integer array with **rows = true class, columns = predicted class**: `C[i, j]` is how many samples of true class `i` the model called class `j`. The diagonal is correct predictions. Class order is the sorted union of labels found in both arrays, not order of appearance and not the estimator's `classes_` — which matters for string labels, where '10' sorts before '9'. Pass `labels=` to pin the order explicitly. Because sorting puts 0 before 1, the binary idiom `tn, fp, fn, tp = confusion_matrix(y_true, y_pred).ravel()` works — but only for 0/1 labels in that orientation, so I pass `labels=[0, 1]` to make it explicit. `normalize='true'` divides each row by its support, turning the diagonal into per-class recall; `'pred'` normalises columns and `'all'` by the grand total. `ConfusionMatrixDisplay.from_predictions` plots it. Note that some textbooks and other tools use the transposed convention.

code

python · 9 lines
python
from sklearn.metrics import confusion_matrix

y_true = [0, 1, 0, 1, 1, 0]
y_pred = [0, 1, 1, 1, 0, 0]

tn, fp, fn, tp = confusion_matrix(y_true, y_pred, labels=[0, 1]).ravel()
print(tn, fp, fn, tp)
print(confusion_matrix(y_true, y_pred, labels=[1, 0]))
print(confusion_matrix(y_true, y_pred, normalize="true"))

go deeper

for a junior

Be able to point at a 2x2 matrix and name each cell, and state that rows are the true classes. Knowing the tn, fp, fn, tp unpacking line for 0/1 labels is a common screening expectation.

for a middle

Explain that label order is the sorted union of both arrays, why that bites with string labels, and what normalize='true' versus 'pred' turns the diagonal into.

for a senior

Use the off-diagonal structure diagnostically — naming which class pairs are confused and in which direction — and insist on labels= so matrices stay comparable across folds and across model versions.

for a principal

Decide what an error-analysis artefact must contain before a model ships, so that reviews argue about specific confusions and their business cost rather than a single accuracy figure that hides an unusable rare class.

## The layout `confusion_matrix(y_true, y_pred, *, labels=None, sample_weight=None, normalize=None)` returns a square array of counts. The orientation is fixed and worth memorising because half of all confusion-matrix confusion is an orientation mix-up: **row index is the true class, column index is the predicted class**. So `C[i, j]` is the number of samples whose true label is class `i` and whose prediction was class `j`. Reading along a row tells you what happened to a real class; reading down a column tells you what ended up in a predicted bucket. The diagonal holds correct predictions and everything off-diagonal is an error. This convention is not universal. Several statistics textbooks, and some other libraries, put predictions on the rows. If you copy a formula from elsewhere and it gives nonsense, the transpose is the first thing to check. ## Label ordering When `labels` is not given, scikit-learn builds the class set from the union of the values present in `y_true` and `y_pred`, and sorts it. Three consequences follow. First, ordering is *lexicographic for strings*. Labels `'1'`, `'10'`, `'2'` come out in that order, which quietly scrambles a plot's axis. Second, a class that appears in neither array is absent from the matrix entirely, so two folds of a cross-validation can produce matrices of different sizes — pass `labels=` if you intend to sum them. Third, the order is *not* taken from the fitted estimator, so if you want the matrix aligned with a model's own `classes_`, pass `labels=clf.classes_`. ## The binary unpacking idiom For 0/1 labels, sorting places 0 first, so the matrix is - `C[0, 0]` true 0 predicted 0 = true negatives - `C[0, 1]` true 0 predicted 1 = false positives - `C[1, 0]` true 1 predicted 0 = false negatives - `C[1, 1]` true 1 predicted 1 = true positives Flattened row-major, that is `tn, fp, fn, tp`, which is why `tn, fp, fn, tp = confusion_matrix(y_true, y_pred).ravel()` is such a common line. It is safe *only* when the labels sort with the negative class first. With labels `'no'`/`'yes'` it still happens to work; with `'positive'`/`'negative'` it silently inverts, because `'negative'` sorts first but is your negative class — fine — whereas with `'benign'`/`'malignant'` you get benign first, which may or may not be what you meant. The habit that removes all doubt is `confusion_matrix(y_true, y_pred, labels=[neg, pos]).ravel()`. ## Normalisation `normalize` accepts `'true'`, `'pred'`, `'all'` or `None`. With `'true'`, each row is divided by its total, so entries are the fraction of each real class routed to each prediction and the diagonal becomes per-class recall — this is the version to show when class sizes differ wildly, because raw counts make a rare class invisible. With `'pred'`, columns sum to one and the diagonal becomes per-class precision. With `'all'`, everything is divided by the number of samples. Normalised output is float, so the `ravel()` unpacking no longer gives integer counts. ## Related helpers `ConfusionMatrixDisplay.from_predictions(y_true, y_pred)` and `ConfusionMatrixDisplay.from_estimator(clf, X, y)` render the matrix with labelled axes and handle the class ordering for you. `classification_report` gives the derived per-class precision, recall, F1 and support in text form, or as a dict with `output_dict=True`. For multilabel data, `multilabel_confusion_matrix` returns one 2x2 matrix per label instead of a single square matrix, since a sample can belong to several labels at once. ## What to actually look at The value of the matrix over a single scalar is that it names the *specific* confusions. A model with 90% accuracy across ten classes might be perfect on nine and useless on one, or might systematically confuse two visually similar classes in both directions. Only the off-diagonal structure shows that, and it is usually the fastest route from 'the metric is bad' to a hypothesis about why.

  • Two cross-validation folds produce confusion matrices of different shapes. Why, and how do you fix it?
    Without `labels=`, the class set is inferred from the values present in that fold's `y_true` and `y_pred`. A rare class missing from one fold — or never predicted in it — simply drops out, shrinking the matrix. Passing `labels=clf.classes_` or the full sorted class list forces every matrix to the same shape so the folds can be summed.
  • What does the diagonal become when you pass normalize='true'?
    Per-class recall. Row normalisation divides each row by the number of true samples in that class, so `C[i, i]` becomes the fraction of real class `i` that was correctly identified. Using `normalize='pred'` instead normalises columns, which makes the diagonal per-class precision. Both turn the array into floats, so integer count unpacking no longer applies.
  • When is the tn, fp, fn, tp unpacking unsafe?
    Whenever the sorted label order does not put the negative class first, or the target is not binary. Sorting is lexicographic for strings, so a positive label that sorts before the negative one silently swaps every term, and a third class makes the flattened array nine elements long and the unpack raises. Passing `labels=[neg, pos]` explicitly removes both risks.

saying these in an interview costs you the question

  • Reading columns as true classes and rows as predictions
  • Assuming class order follows first appearance in y_true
  • Unpacking ravel() as tp, fp, fn, tn
  • Expecting the matrix to use the estimator's classes_ order automatically
  • Summing per-fold matrices that were built without labels=

context

open as a page

In scikit-learn, what does precision_score's average parameter control?

level: middleimportance: must knowfreq 68%

basics

~20 s

The average parameter decides how per-class precision values are reduced to one number. It defaults to 'binary', which scores only one class and rejects multiclass targets; multiclass data needs 'macro', 'micro', 'weighted', or None for the per-class array.

open as a page

Why does scikit-learn's roc_auc_score need scores rather than predict() labels?

level: middleimportance: must knowfreq 64%

basics

~20 s

ROC-AUC measures how well a continuous score ranks positives above negatives, so it needs predict_proba(X)[:, 1] or decision_function(X). Hard 0/1 labels from predict() collapse the curve to one threshold and silently return a lower, meaningless number instead of raising.

open as a page

In scikit-learn, when do you use silhouette_score versus adjusted_rand_score?

level: middleimportance: should knowfreq 32%

basics

~20 s

silhouette_score(X, labels) is internal: it scores cluster shape from the feature matrix alone, so it works without ground truth. adjusted_rand_score(labels_true, labels_pred) is external: it compares two labelings and requires true labels you usually do not have.

open as a page

In scikit-learn, how does make_scorer turn a metric into a scorer?

level: middleimportance: should knowfreq 46%

basics

~10 s

make_scorer wraps a metric function into a callable of the form scorer(estimator, X, y). greater_is_better=False negates the value so larger always means better, and response_method decides whether the scorer calls predict, predict_proba, or decision_function.

open as a page

A fraud model shows 0.97 ROC-AUC but poor live precision — which sklearn.metrics calls do you reach for?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Score the ranking with average_precision_score, then call precision_recall_curve to choose an operating threshold: it returns precision and recall arrays one element longer than thresholds. Rebuild confusion_matrix at that threshold instead of trusting predict()'s default 0.5.

open as a page