skip to content

Precision Versus Recall

Precision asks how many flagged items were right, recall how many real positives you caught, and F-beta trades them off. Nearly every screen asks which one a described product should favour.

on this pageshow

questions

4

What is the difference between precision and recall, and why do they trade off?

level: juniorimportance: must knowfreq 88%

answer

  1. two rates, two different denominators
  2. one counts what the model claimed
  3. the other counts what was true
  4. loosen the rule, catch more, be wrong more
  5. everything-positive scores recall 1.0

basics

~20 s

Precision is the share of predicted positives that are truly positive; recall is the share of actual positives the model catches. Loosening the decision rule labels more examples positive, which raises recall and usually lowers precision.

solid answer

~50 s

Both are built from the same three counts but divide by different things. Precision is `TP / (TP + FP)` — of everything the model called positive, how much really was. Recall, also called sensitivity or the true-positive rate, is `TP / (TP + FN)` — of everything that really was positive, how much the model found. They pull against each other because a single model has one decision rule: relax it and you claim more positives, so you catch more real ones (recall up) while sweeping in more wrong ones (precision usually down). At the extreme, calling everything positive gives recall 1.0 and precision equal to the positive rate in the data. The trade-off is along one model's decision rule — a genuinely better model, with better features or a better ranking of examples, can lift both at once.

go deeper

for a junior

Be able to state both formulas from the four confusion-matrix cells without hesitating, and say which denominator belongs to which. Know that recall and sensitivity are the same thing.

for a middle

Explain the trade-off mechanically: the decision rule grows the predicted-positive set, recall cannot fall, precision usually does. Distinguish moving the rule from actually improving the model.

for a senior

Show you pick the metric from the cost of each error in the specific product, not by habit, and that you report both numbers with absolute error counts so a stakeholder can see the real volume.

for a principal

Own the question of which error the organisation is willing to make. Argue for a single agreed metric per system, and for the review capacity or fallback path that makes the losing side of the trade-off survivable.

## The two counts they do not share Every binary prediction lands in one of four cells: a **true positive (TP)** the model called positive and was right about, a **false positive (FP)** it called positive and was wrong about, a **false negative (FN)** it called negative but was actually positive, and a **true negative (TN)** it called negative and was right about. - **Precision** = `TP / (TP + FP)` — the denominator is everything the model *predicted* positive. - **Recall** = `TP / (TP + FN)` — the denominator is everything that *actually is* positive. That is the whole difference, and it is the thing candidates most often blur. Precision is read down the model's positive column; recall is read across the true positive row. Recall goes by three names depending on the field — recall, **sensitivity**, and the **true-positive rate** — and they are the same quantity. Precision is also called **positive predictive value**. In words: precision answers *"when this thing fires, how often should I believe it?"* Recall answers *"of everything I was supposed to catch, how much did I actually catch?"* Note that neither metric uses TN anywhere. Both are entirely about the positive class, which is exactly why they are the pair of choice when positives are the rare and interesting thing. ## Why they trade off A classifier does not emit a label directly; it emits a score, and a decision rule turns that score into a label. Slide that rule so more examples get labelled positive and two things happen at once. The set of predicted positives grows, so it can only gain true positives — recall never goes down. But the added examples are the ones the model was less sure about, so proportionally more of them are wrong, and precision typically falls. The two extreme cases make it concrete. Label *everything* positive: you catch every real positive, so recall is 1.0, but precision collapses to the fraction of the data that is genuinely positive — on a rare-positive problem, a terrible number. Label only the single example you are most confident about: precision may be a perfect 1.0 while recall is almost zero. Neither model is useful, and both look excellent on one metric alone. This is why the two are always quoted as a pair. The crucial caveat: this trade-off describes moving the decision rule of *one fixed model*. It is not a law that precision and recall cannot both improve. A model with better features, more data, or a better ordering of examples by risk dominates the old one — at the same recall it makes fewer false positives. That is what "the model got better" means, as opposed to "the model got more or less trigger-happy". ## The degenerate corners Precision is undefined when the model predicts no positives at all — the denominator is zero. Tooling usually reports 0 in that case, which quietly hides a model that has stopped firing entirely. Recall is undefined if the evaluation set contains no actual positives, which happens more often than people expect on small held-out slices of a rare-event problem. ## Mapping the pair onto cost The choice between them is not statistical, it is an argument about which mistake hurts. Take an email spam filter. A **false positive** is a legitimate message — a customer's invoice — quarantined into a folder nobody reads; the recipient may never learn it existed. A **false negative** is one piece of junk mail sitting in the inbox, which the user deletes in a second. The costs are wildly asymmetric in favour of precision, so a spam filter is tuned to almost never quarantine real mail even though that means letting some junk through. Flip the asymmetry and the answer flips. When the positive class is a disease, a fraud case, or a safety fault, a miss is expensive and a false alarm merely triggers a cheap second look — there, recall is the metric you protect and false positives are the price you pay. ## What to report Quote both numbers, and quote the raw cell counts behind them. "Precision 0.82, recall 0.61" is informative; "precision 0.82" alone is a sales pitch. Counts matter too, because the same rates on a stream of ten million items mean a very different absolute number of angry users than they do on a batch of two hundred.

  • A model labels every example positive. What are its precision and recall?
    Recall is 1.0 — no actual positive is missed. Precision falls to the positive rate in the data, so on a set that is 3% positive, precision is 0.03. It is the standard sanity check that a high recall number alone proves nothing.
  • Which of the two can you fail to compute at all, and when?
    Precision, when the model predicts no positives: the denominator TP + FP is zero, so the ratio is undefined. Tools often print 0.0, which disguises a model that has stopped firing. Recall is undefined only if the evaluation set contains no actual positives.
  • Can precision and recall both improve at the same time?
    Yes — but not by moving the decision rule of one model, which only trades one for the other. Both rise when the model itself gets better: better features, more data, or a better ordering of examples, so that at any given recall it makes fewer false positives.
  • Why does neither precision nor recall involve true negatives?
    Both are defined purely on the positive class — what was claimed positive and what was truly positive. That makes them insensitive to a flood of easy negatives, which is precisely why they are preferred to accuracy when positives are rare.

A fishing net. Precision asks what fraction of what you hauled in is actually fish; recall asks what fraction of the fish in the lake ended up in the net. A wider net catches more fish and more rubbish.

saying these in an interview costs you the question

  • Uses precision and recall as interchangeable words
  • Puts actual positives in precision's denominator
  • Quotes recall alone as proof the model is good
  • Claims both metrics can never improve together
  • Thinks precision counts true negatives

context

open as a page

Why is F1 the harmonic mean of precision and recall rather than the arithmetic mean?

level: middleimportance: should knowfreq 62%

basics

~20 s

The harmonic mean is dominated by the smaller of the two values, so F1 stays low unless precision and recall are both decent. Precision 0.90 with recall 0.10 gives F1 0.18, while the arithmetic mean would flatter it at 0.50.

open as a page

Why can a sepsis alert with 95% specificity still be wrong most times it fires?

level: middleimportance: should knowfreq 48%

basics

~20 s

Specificity is measured only among patients who do not have sepsis, so it ignores how rare sepsis is. At a 2% rate, 5% of the huge non-sepsis group yields far more false alarms than true cases, leaving precision near 25%.

open as a page

How do you choose beta in F-beta when a missed case costs far more than a false alarm?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Beta sets how many times more recall matters than precision: F2 weights recall twice as heavily, F0.5 half as heavily, F1 equally. Derive beta from the cost ratio of a missed case to a false alarm.

open as a page