skip to content

How do macro, micro and weighted averaging of multiclass F1 differ?

level: middleimportance: must knowfreq 78%

answer

  1. one vote per class versus per sample
  2. average the scores or pool the counts
  3. support weights favour the head classes
  4. in single-label multiclass one of them is accuracy
  5. the gap between two of them is the diagnosis

basics

~20 s

Macro averages per-class F1 scores with equal weight, so a rare class counts as much as a common one. Micro pools every class's true positives and errors first, so frequent classes dominate. Weighted averages per-class scores by class support.

solid answer

~50 s

All three start from per-class one-vs-rest counts and differ only in how they aggregate. **Macro** computes F1 for each class and takes the plain mean — one vote per class, so the long tail can sink the score. **Micro** sums TP, FP and FN across all classes and computes a single F1 from those pooled counts — one vote per sample, so the head classes dominate; in single-label multiclass where every sample gets exactly one predicted class, micro-F1 equals plain accuracy. **Weighted** averages the per-class F1 scores by class support, which tracks micro closely and is easily mistaken for tail coverage. A 40-category marketplace taxonomy routinely shows macro-F1 near 0.46 against micro-F1 near 0.76: the big categories work, the small ones do not. Pick the average that matches the cost — macro when every category must work, micro or weighted when errors cost the same per item.

code

python · 19 lines
python
# (tp, fp, fn) per class; 127 samples, every one given exactly one label
counts = {"A": (90, 8, 10), "B": (5, 12, 15), "C": (2, 10, 5)}

def f1(tp, fp, fn):
    p = tp / (tp + fp) if tp + fp else 0.0
    r = tp / (tp + fn) if tp + fn else 0.0
    return 2 * p * r / (p + r) if p + r else 0.0

# macro: average the per-class scores, one vote per class
macro = sum(f1(*c) for c in counts.values()) / len(counts)

# micro: pool the counts first, one vote per sample
tp = sum(c[0] for c in counts.values())
fp = sum(c[1] for c in counts.values())
fn = sum(c[2] for c in counts.values())
micro = f1(tp, fp, fn)

print("macro", round(macro, 3))  # 0.463 - dragged down by the two small classes
print("micro", round(micro, 3))  # 0.764 - equals accuracy, 97 correct of 127

go deeper

for a junior

Know the three names and the one-line difference: macro treats classes equally, micro treats samples equally, weighted scales by class size. Be ready to say which one a rare class can influence.

for a middle

Derive them from per-class TP, FP and FN on the whiteboard, and explain why micro-F1 collapses to accuracy when each sample gets exactly one label. Expect a small numeric example where macro and micro disagree sharply.

for a senior

Justify a choice against real error costs, not taste. Show that you report macro and micro together with the per-class table, and that you read the gap between them as evidence about how the tail behaves rather than picking whichever looks better.

for a principal

Own the reporting standard. Decide which average is the organisation's headline for many-class systems, make the definition explicit so numbers compare across teams and quarters, and resist the pull toward whichever aggregation makes a launch look ready.

## The common starting point For K classes you first build K one-vs-rest tables: for class c, `TP_c` is the count correctly predicted c, `FP_c` the other classes wrongly called c, `FN_c` the true c's called something else. Precision, recall and F1 for each class follow from those three counts. The three averages differ only in the aggregation step afterwards. ## Macro ``` macro-F1 = (1/K) * sum over c of F1_c ``` Every class contributes equally regardless of size. A category with 12 test samples moves the score exactly as much as one with 12,000. Macro is the right choice when each class carries independent business meaning — every product category must be shelvable, every ticket queue must be routable — and it is the metric that exposes a long tail. Its weakness is volatility: tiny classes have noisy F1 scores, and with 40 classes a handful of ten-sample categories can swing macro-F1 by several points between runs. ## Micro ``` micro-precision = sum(TP_c) / sum(TP_c + FP_c) micro-recall = sum(TP_c) / sum(TP_c + FN_c) micro-F1 = harmonic mean of the two ``` Counts are pooled *before* any ratio is taken, so every sample contributes one unit and large classes dominate. In **single-label** multiclass, where every sample is assigned exactly one class, each error simultaneously creates one FP for the predicted class and one FN for the true class. The two sums are therefore equal, micro-precision equals micro-recall, and both equal the fraction of samples classified correctly — micro-F1 *is* accuracy. That identity breaks in multilabel settings and whenever the model may abstain. ## Weighted ``` weighted-F1 = sum over c of (support_c / N) * F1_c ``` Per-class F1 scores averaged by how many true samples each class has. It sits between macro and micro in spirit but is closer to micro in behaviour, because support weights are exactly what make big classes dominate. Its real hazard is rhetorical: teams report weighted-F1 believing they have accounted for imbalance, when weighting by support does the opposite of protecting the tail. Note also that weighted-F1 is not bounded between the weighted precision and weighted recall, so it can behave in ways a single F1 cannot. ## A worked contrast With three classes and per-class (TP, FP, FN) of A = (90, 8, 10), B = (5, 12, 15), C = (2, 10, 5): per-class F1s are 0.909, 0.270 and 0.211. Macro-F1 is 0.463. Pooling gives TP = 97, FP = 30, FN = 30 over 127 samples, so micro-F1 = 97/127 = 0.764 — identical to accuracy, as expected. Weighted-F1, using supports of 100, 20 and 7, comes out at 0.770. One model, three defensible numbers spanning 0.30 of range. Reporting "F1 = 0.77" without saying which average is not a metric, it is a decision hidden in a word. ## Choosing Ask what one error costs. In support-ticket routing across nine queues where one queue holds 60% of tickets, micro and weighted are both essentially "how often is a ticket routed correctly", which is fine if every misroute costs the same handling delay. But if the eight small queues include a legal-escalation queue, macro is the honest headline, or better, you report the small queues' per-class recall directly. Conversely, when classes are near-balanced the three numbers converge and arguing about them is wasted time — check the supports before you argue. A useful discipline: report macro **and** micro side by side plus the per-class table. The gap between the two is itself the diagnostic — a large gap says head and tail behave differently, and no single number will summarise that away. ## Two traps worth naming **Macro-F1 is not the F1 of the macro-averaged precision and recall.** Averaging the F1 scores and computing an F1 from averaged precision and recall are different quantities, because the harmonic mean is not linear. Both are called "macro F1" by different tools, and they disagree most exactly when classes are uneven. Say which one you computed. **Micro is not always above macro.** It usually is, because the biggest classes are usually the easiest — but if a rare class is trivially separable while the dominant class is hard, macro can exceed micro. Treat the ordering as an observation, never an assumption. ## Beyond F1 The same three aggregations apply to precision and recall separately, and to threshold-free measures. Macro one-vs-rest AUC binarises the problem once per class, computes each class's AUC against all the rest, and averages them — with the caveat that each binarisation has a different positive rate, so the per-class AUCs are not measuring equally hard problems. A support-weighted variant exists and shifts the emphasis back to the head classes exactly as weighted-F1 does.

  • When exactly does micro-averaged F1 equal plain accuracy?
    In single-label multiclass where every sample receives exactly one predicted class. Each mistake then adds one false positive to the predicted class and one false negative to the true class, so the pooled FP and FN totals are equal, micro-precision equals micro-recall, and both equal the fraction correct. It fails in multilabel settings and when the model can abstain.
  • A 40-category product taxonomy shows macro-F1 0.46 and micro-F1 0.76 — what do you conclude?
    The head categories carrying most of the volume are handled well and the long tail is not; macro is pulled down by many small categories scoring poorly. Whether that is acceptable depends on cost: if every category must be shelvable, optimise and report macro plus the worst per-class scores rather than the flattering pooled number.
  • How do you extend AUC to a 10-class problem?
    Macro one-vs-rest AUC: binarise once per class, compute that class's AUC against all the rest, and average the ten values. A support-weighted average is the alternative. The caveat is that each binarisation has its own positive rate, so the ten AUCs are not measuring equally hard sub-problems and should be shown individually too.
  • Why can weighted-F1 mislead a team that chose it to handle imbalance?
    Weighting by support gives the biggest classes the biggest say, which is the opposite of protecting rare ones. It tracks micro closely and hides a failing tail behind a healthy head. If the concern is imbalance, macro or the per-class table is the honest report; weighted is appropriate only when error cost really is proportional to volume.

Macro averaging is one vote per class; micro averaging is one vote per sample. It is the same election counted two ways, and small constituencies only matter under the first.

saying these in an interview costs you the question

  • Reports an F1 number without saying which average
  • Claims weighted averaging protects rare classes
  • Asserts micro-F1 is always higher than macro-F1
  • Thinks macro-F1 is the F1 of the pooled counts
  • Says micro-F1 equals accuracy in every setting
  • Confuses averaging the F1 scores with F1 of averaged precision and recall

context