skip to content

Why can a sepsis alert with 95% specificity still be wrong most times it fires?

level: middleimportance: should knowfreq 48%

answer

  1. different denominators, that is the whole trick
  2. specificity conditions on the true label
  3. the negatives vastly outnumber the positives
  4. five percent of 9,800 is 490
  5. precision moves with prevalence, specificity does not

basics

~20 s

Specificity is measured only among patients who do not have sepsis, so it ignores how rare sepsis is. At a 2% rate, 5% of the huge non-sepsis group yields far more false alarms than true cases, leaving precision near 25%.

solid answer

~40 s

Specificity is `TN / (TN + FP)` — conditioned on the patients who truly do not have sepsis. Precision is `TP / (TP + FP)` — conditioned on the patients the alert fires on. Different denominators, so a great specificity says nothing directly about how trustworthy an alert is. Work it through for 10,000 monitored patients at 2% prevalence with sensitivity 0.80 and specificity 0.95: 200 truly septic, of whom 160 are caught; 9,800 not septic, of whom 5% — 490 patients — trip the alert anyway. The alert fires 650 times and is right 160 of them, so precision is about 25%. Sensitivity and specificity are properties of the model on each true class and do not move with prevalence; precision does. That gap is what drives alarm fatigue on rare-event alerts.

go deeper

for a junior

Know that specificity is true negatives over all actual negatives, and that it is a different measure from precision. Remember the false-positive rate is simply one minus specificity.

for a middle

Be able to run the counts out loud from a prevalence, a sensitivity and a specificity, and explain why precision depends on the class mix while the other two do not.

for a senior

Show you convert rates into absolute alert volumes for the real population before shipping, and that you recompute precision locally instead of trusting a figure measured somewhere else.

for a principal

Frame it as an operations problem: alarm fatigue is a system failure, so argue for a false-alarm budget per shift and for who owns the alert when the base rate drifts.

## Four rates, two denominators From the confusion matrix: - **Recall / sensitivity / true-positive rate** = `TP / (TP + FN)` — among the truly positive. - **Specificity / true-negative rate** = `TN / (TN + FP)` — among the truly negative. - **False-positive rate** = `FP / (FP + TN)` = `1 - specificity` — same denominator as specificity, the other slice of it. - **Precision / positive predictive value** = `TP / (TP + FP)` — among those the model *called* positive. The first three are conditioned on the **true** class. The last is conditioned on the **prediction**. That single structural difference explains everything below, and mixing up precision and specificity is one of the most common confusion-matrix errors in interviews — they share the FP term but sit on opposite denominators. ## The arithmetic on a hospital sepsis alert Take a monitoring alert shipped at its default cut, with sensitivity 0.80 and specificity 0.95, running over 10,000 monitored patients where 2% genuinely develop sepsis. - Truly septic: 200. Caught at 80% → **TP = 160**, missed **FN = 40**. - Not septic: 9,800. Correctly left alone at 95% → **TN = 9,310**; the other 5% → **FP = 490**. The alert fires `160 + 490 = 650` times. It is right 160 of those: **precision ≈ 0.25**. Three of every four alerts are false, on a model whose specificity is a genuinely respectable 95%. Nothing here is a flaw in the model or a bug in the metric. It is arithmetic: 5% of a very large group of negatives beats 80% of a very small group of positives. The FP count scales with the number of negatives, and the negatives are 49 times more numerous than the positives. ## Why prevalence moves precision and not the other two Sensitivity and specificity are computed *within* a true class, so changing the mix between classes leaves them alone — they are (to a first approximation) properties of the model and the cut, transportable between populations. Precision divides across the classes: its numerator comes from the positive group and part of its denominator comes from the negative group, so it depends on the ratio between them. Deploy the same alert, unchanged, on a high-acuity ward where 20% of monitored patients develop sepsis. Per 10,000: 2,000 septic → TP 1,600; 8,000 not → FP 400. Precision jumps to `1600 / 2000 = 0.80`. Same model, same cut, same sensitivity and specificity — precision went from 25% to 80% because the population changed. The operational consequence is blunt: **a precision figure is not transferable**. A vendor's quoted PPV, a benchmark number from a curated evaluation set, or last year's figure from a different patient mix tells you very little about what the alert will do on your population. Sensitivity and specificity travel far better; recompute precision locally. ## Why 95% specificity is a weaker claim than it sounds Humans read "95% specific" as "only 5% wrong", and on a rare-event problem the 5% is applied to the overwhelming majority of the data. A one-point change in specificity, from 95% to 96%, removes 98 false alarms per 10,000 patients here — more than half the true-positive count. On rare-event alerting, specificity has to be quoted with the number of negatives it applies to, or converted into an expected false-alarm count per shift, per day, per thousand patients. That absolute count is what determines whether clinicians keep responding to the alert or start ignoring it. ## Which metric to quote to whom For a modeller comparing cuts on a fixed dataset, sensitivity and specificity are the stable pair. For anyone who has to act on the output — a clinician, a reviewer, an on-call engineer — precision is the number that describes their experience, because they only ever see the cases where it fired. Report both, plus the raw counts, and state the prevalence they were measured at.

  • How does the false-positive rate relate to specificity?
    It is the complement: FPR = FP / (FP + TN) = 1 - specificity. Both are read over the truly negative group, so 95% specificity is a 5% false-positive rate. Neither is affected by how many positives exist in the population.
  • The same alert is moved to a ward where 20% of patients develop sepsis. What changes?
    Sensitivity and specificity are unchanged — they are conditioned on the true class. Precision jumps: per 10,000 patients there are now 1,600 true positives against 400 false ones, so precision rises from about 25% to about 80%. The model did not improve; the base rate did.
  • Which two of recall, specificity and precision are safe to quote across populations?
    Recall and specificity, because each is computed within a single true class and does not depend on the class mix. Precision must be recomputed on the target population; a quoted precision from a different prevalence is close to meaningless.
  • Why is specificity a misleading headline number on a rare-event alert?
    Because the 5% it concedes is applied to the overwhelming majority of the data, so a small specificity change swings the absolute false-alarm count enormously. Convert it into expected false alarms per day or per thousand cases before anyone judges whether the alert is usable.

A metal detector at a beach that correctly ignores 95% of ordinary sand. There is so much sand that it still beeps far more often for grit than for coins.

saying these in an interview costs you the question

  • Treats specificity and precision as the same measure
  • Reads 95% specificity as 95% of alerts being correct
  • Carries a quoted precision across populations unchanged
  • Defines specificity as true negatives over all examples
  • Thinks prevalence changes a model's sensitivity

context