skip to content

In scikit-learn, what does precision_score's average parameter control?

level: middleimportance: must knowfreq 68%

answer

  1. one number out of many classes
  2. the default assumes two classes
  3. macro treats every class equally
  4. weighted follows the majority class
  5. micro pooling collapses to accuracy

basics

~20 s

The average parameter decides how per-class precision values are reduced to one number. It defaults to 'binary', which scores only one class and rejects multiclass targets; multiclass data needs 'macro', 'micro', 'weighted', or None for the per-class array.

solid answer

~50 s

`precision_score` always computes precision per class first, then reduces it according to `average`. The default `average='binary'` reports only the class named by `pos_label` (1 by default), so handing it multiclass labels raises a `ValueError` telling you to pick an averaging mode. `average='macro'` is the unweighted mean over classes, so a 20-sample class counts as much as a 20,000-sample one. `average='weighted'` weights each class by its support, which is why it tracks the majority class and can hide terrible rare-class performance. `average='micro'` pools the true-positive and false-positive counts globally; in single-label multiclass that makes micro precision, micro recall and micro F1 all equal accuracy. `average=None` returns the per-class array, which is what I look at before choosing any summary. `average='samples'` is for multilabel indicator data only. The same parameter appears on `recall_score`, `f1_score` and `fbeta_score`, and `classification_report` prints the macro and weighted rows for you.

code

python · 10 lines
python
import numpy as np
from sklearn.metrics import precision_score

y_true = np.array([0, 1, 2, 2, 1, 0])
y_pred = np.array([0, 2, 2, 2, 1, 0])

print(precision_score(y_true, y_pred, average=None))
print(precision_score(y_true, y_pred, average="macro"))
print(precision_score(y_true, y_pred, average="weighted"))
print(precision_score(y_true, y_pred, average="micro"))

go deeper

for a junior

Know that precision_score defaults to binary scoring and needs an explicit average for three or more classes. Being able to say 'macro means unweighted mean over classes' is enough at this level.

for a middle

Explain all four reduction modes and what each hides, including that micro averaging equals accuracy on single-label multiclass. Mention pos_label, labels= and zero_division as the knobs that change the number.

for a senior

Show judgment about which average goes into a report and which goes into a model-selection scorer, and be ready to explain a large macro-versus-weighted gap on an imbalanced dataset as a diagnosis rather than a formatting detail.

for a principal

Own the reporting convention for a team: which averaged metric is the headline number, whether rare classes are weighted equally, and how zero_division is set, so that results from different projects stay comparable rather than each author picking the flattering mode.

## The shape of the problem Precision is defined for one class at a time: of the samples the model assigned to class k, what fraction really were class k. With two classes you conventionally care about one of them (the positive class), so a single number is natural. With five classes there are five precisions, and any single number you report is a *choice about how to combine them*. In scikit-learn that choice is the `average` keyword on `precision_score`, `recall_score`, `f1_score` and `fbeta_score`. ## Why the default rejects your multiclass data The signature is roughly `precision_score(y_true, y_pred, *, labels=None, pos_label=1, average='binary', sample_weight=None, zero_division='warn')`. Because `average='binary'` is the default, the function scores exactly the class in `pos_label` and ignores the rest. If `y_true` contains three or more distinct labels, scikit-learn raises a `ValueError` saying the target is multiclass but average is 'binary', and lists the alternatives. This is the single most common first contact people have with the parameter: their binary notebook worked, they added a third class, and the metric call broke. Nothing is wrong with the model; the reduction is simply undefined. Note also that `pos_label=1` is a literal label value, not "the second class". If your labels are the strings 'spam' and 'ham', the binary default fails until you pass `pos_label='spam'`. ## The reduction modes **`average=None`** returns an array of per-class precisions, ordered by sorted class labels, or by whatever you pass in `labels=`. This is the honest starting point: it shows you which class is dragging the model down before any averaging hides it. **`average='macro'`** takes the arithmetic mean of those per-class values with equal weight per class. It treats every class as equally important regardless of how many samples it has, so it is the mode that punishes a model for failing on rare classes. On heavily imbalanced data macro scores are usually much lower than weighted ones, and that gap is informative. **`average='weighted'`** takes the same per-class values but weights each by its support (the number of true samples in that class). Because support is dominated by the majority class on imbalanced data, the weighted score largely reports majority-class performance. It is the number people quote when they want a flattering figure, and interviewers know it. **`average='micro'`** does not average the per-class scores at all. It sums true positives, false positives and false negatives across all classes and computes one global ratio. In single-label multiclass, every misclassification is simultaneously a false positive for one class and a false negative for another, so micro precision equals micro recall equals micro F1 equals accuracy. If you report "micro-F1" as though it were a robustness metric on single-label data, you are reporting accuracy under a fancier name. Micro averaging becomes genuinely distinct in the multilabel setting, where those counts no longer balance. **`average='samples'`** applies only to multilabel indicator targets. It computes the metric per sample across its label set and averages over samples. Passing it ordinary multiclass labels raises. ## Restricting the class set The `labels` argument does double duty. With `average=None` it fixes the order of the returned array. With macro or weighted averaging it restricts *which* classes are included, which is how you exclude a garbage or 'other' class from the summary, or how you include a class that never appears in `y_pred` so that it contributes a zero instead of vanishing. ## zero_division When a class is never predicted, its precision is 0/0. By default scikit-learn warns and substitutes 0.0. Passing `zero_division=0`, `1`, or `np.nan` chooses the substitute explicitly and silences the warning. Using `np.nan` makes macro averaging skip the undefined class rather than pulling the mean down with a zero, which changes the reported number — say which you used. ## In practice A defensible reporting habit is: look at `average=None` or `classification_report` first, quote macro when classes matter equally, quote weighted only alongside macro so the gap is visible, and never quote micro on single-label multiclass without admitting it is accuracy. When you plug a metric into a search via `scoring=`, the string names encode the same choice: 'f1_macro', 'f1_micro', 'f1_weighted', 'precision_macro' and so on.

  • On single-label multiclass data, why do micro precision, micro recall and micro F1 all come out identical?
    Micro averaging pools counts globally. In single-label multiclass every wrong prediction is a false positive for the predicted class and a false negative for the true class, so the global FP and FN totals are equal. Precision and recall therefore share a denominator and coincide, and their harmonic mean coincides with them. The common value is exactly accuracy.
  • How do you keep a class that the model never predicts from silently disappearing from a macro average?
    Pass it explicitly in `labels=`. Scikit-learn otherwise derives the class set from the labels actually present in `y_true` and `y_pred`. Including it forces the class into the computation with an undefined precision, which `zero_division` then resolves to 0.0, 1.0 or `np.nan`. Choosing `np.nan` skips it in the mean; choosing 0 penalises the model for it.
  • Which averaging mode would you report to a stakeholder for a 95/5 imbalanced problem, and why?
    I would show macro alongside the per-class array, and give weighted only for context. Weighted averaging is dominated by the 95% class, so it can read above 0.9 while the minority class is near zero. Macro exposes that, and the per-class array names which class is failing so the conversation is about the actual failure, not the summary.

Per-class precision is a scorecard with one row per class; average is the rule for turning the scorecard into a single grade, and different rules can rank the same two models differently.

saying these in an interview costs you the question

  • Calling precision_score on multiclass labels and expecting a default average
  • Assuming macro and weighted averaging give the same number
  • Quoting weighted average as evidence of rare-class performance
  • Believing micro-F1 differs from accuracy on single-label multiclass
  • Thinking average='samples' works for ordinary multiclass labels

context