skip to content

Why is F1 the harmonic mean of precision and recall rather than the arithmetic mean?

level: middleimportance: should knowfreq 62%

answer

  1. an average that punishes imbalance
  2. the smaller number dominates
  3. reciprocals, not sums
  4. zero on either side means zero overall
  5. 0.90 and 0.10 land at 0.18

basics

~20 s

The harmonic mean is dominated by the smaller of the two values, so F1 stays low unless precision and recall are both decent. Precision 0.90 with recall 0.10 gives F1 0.18, while the arithmetic mean would flatter it at 0.50.

solid answer

~40 s

F1 is `2 * P * R / (P + R)`, the harmonic mean of precision and recall. The point of choosing it over the plain average is that the harmonic mean is pulled toward the smaller value, so a model cannot buy a good score by maxing one metric and abandoning the other. Precision 0.90 with recall 0.10 scores 0.18, not 0.50 — and a model that predicts one confident positive out of a million, or one that labels everything positive, both score near zero instead of near half. F1 is therefore a *both-or-nothing* summary. It also inherits the property that it ignores true negatives entirely, which is why it suits rare-positive problems. Its assumption is that the two errors cost the same, which is usually a convenience rather than a fact.

code

python · 12 lines
python
def harmonic(p, r):
    return 2 * p * r / (p + r)

def arithmetic(p, r):
    return (p + r) / 2

for p, r in [(0.90, 0.10), (0.60, 0.40), (0.50, 0.50)]:
    print(p, r, round(arithmetic(p, r), 3), round(harmonic(p, r), 3))

# 0.9 0.1 0.5 0.18
# 0.6 0.4 0.5 0.48
# 0.5 0.5 0.5 0.5

go deeper

for a junior

Know the formula F1 = 2 * P * R / (P + R) and that it is a harmonic mean, not a plain average. Remember that a very low precision or recall drags the score down hard.

for a middle

Explain why the harmonic mean is pulled toward the smaller value, work an example such as 0.90 and 0.10 giving 0.18, and note that true negatives never enter the score.

for a senior

Show you never ship F1 alone: quote the precision and recall split and the raw error counts, and say out loud that F1 assumes the two errors cost the same.

for a principal

Decide when a single scalar is the right contract for a team at all, and argue for a cost-aware target where the errors are genuinely asymmetric rather than defaulting to the leaderboard metric.

## The formula and what it is `F1 = 2 * P * R / (P + R)`, where P is precision and R is recall. That is exactly the **harmonic mean** of the two — the reciprocal of the average of the reciprocals, `1 / F1 = (1/P + 1/R) / 2`. The point of a single number is to let you rank models and set one target; the point of *this* single number is what it does to lopsided models. ## Why the arithmetic mean fails The arithmetic mean rewards a total. Any two numbers that add to 1.0 average 0.50, so a classifier with precision 0.90 and recall 0.10 scores exactly as well as one with 0.50 and 0.50, and exactly as well as one with 0.99 and 0.01. That is precisely the ranking you do not want. A model with recall 0.10 finds one positive in ten and is useless no matter how clean the few it finds are; a model with precision 0.10 fires nine false alarms for every real hit. The harmonic mean separates these. Run the same three pairs: - P 0.90, R 0.10 → `2 * 0.09 / 1.00` = **0.18** - P 0.60, R 0.40 → `2 * 0.24 / 1.00` = **0.48** - P 0.50, R 0.50 → `2 * 0.25 / 1.00` = **0.50** Same arithmetic mean, very different F1. The general property: the harmonic mean of two positive numbers is always at most their arithmetic mean, with equality only when they are equal, and it collapses toward the smaller value as the gap widens. In the limit, if either metric is 0 then F1 is 0, whatever the other one is. That is the behaviour you want from a summary of a pair you refuse to let a model trade away. ## The intuition to say out loud The harmonic mean is the right average when the quantities are *rates* and you care about the bottleneck. Averaging speeds over a fixed distance works the same way: an hour at 90 km/h and an hour's worth of distance crawled at 10 km/h does not average to 50 for the trip. F1 treats a weak metric as a bottleneck, not as something a strong metric can compensate for. ## What F1 does not tell you **It hides the split.** Two models both at F1 0.70 may be at (0.70, 0.70) and (0.90, 0.57). Those are very different products — one fires often and is usually right, the other rarely fires but is nearly always right. Always report P and R next to F1. **It ignores true negatives.** Neither precision nor recall touches TN, so neither does F1. That makes it robust to a huge easy-negative population — an advantage on rare-positive problems, and a blind spot if performance on negatives is itself the product. **It assumes symmetric cost.** The 2 in the formula is an equal weighting of the two metrics, which is a statement that a false positive and a false negative are worth about the same. When they are not — a missed disease case against a repeat visit, a quarantined invoice against one junk email — F1 is answering a question you did not ask, and a weighted variant is the honest metric. **It is not comparable across datasets.** F1 has no fixed baseline the way accuracy has the majority class; a random model's F1 depends on the positive rate, so an F1 of 0.40 can be excellent on one problem and dismal on another. ## When it is the right tool F1 earns its place when you genuinely need one number — a leaderboard, a regression check in a pipeline, a model-selection sweep — the positive class is the one that matters, and you have no credible cost ratio to justify weighting one error over the other. Under those three conditions it is a good default. Outside them it is a habit.

  • Two models both report F1 of 0.70. What does that number fail to tell you?
    The split. One could be at precision 0.70 and recall 0.70, the other at 0.90 and 0.57 — very different behaviour in production, one firing often and usually right, the other rarely firing but almost always right. F1 is a ranking device, not a description; always print precision and recall beside it.
  • Does F1 use the true-negative count anywhere?
    No. Precision and recall are both defined only on the positive class, so true negatives never enter F1. That is why it stays informative when negatives vastly outnumber positives, and why it says nothing at all about how the model treats the negative class.
  • Is a high F1 always the goal?
    No — F1 encodes the assumption that a false positive and a false negative are worth the same. When one error is clearly costlier, optimising F1 deliberately mis-weights the thing you care about, and a weighted variant or a direct cost measure is the honest target.

saying these in an interview costs you the question

  • Calls F1 the average of precision and recall
  • Reads F1 near 0.5 as both metrics near 0.5
  • Thinks F1 accounts for true negatives
  • Treats maximising F1 as always correct
  • Compares F1 across datasets as if it had a fixed baseline

context