A test with 99% sensitivity and 95% specificity flags a disease with 1% prevalence: how likely is a positive to be real?
answer
- the condition is rare
- count people, not percentages
- 9,900 healthy, 5% flagged anyway
- compare 99 true against 495 false
- posterior odds = prior odds x LR
basics
~10 sAbout 17%. In 10,000 people, 100 have the disease and 99 of them test positive, while 495 of the 9,900 healthy people also test positive. Rare conditions make positives mostly false positives.
solid answer
~50 sWork it in natural frequencies. Take 10,000 people: 1% prevalence means 100 have the disease and 9,900 do not. Sensitivity 99% means 99 of the 100 diseased test positive. Specificity 95% means 5% of the 9,900 healthy test positive anyway, which is 495 false positives. Total positives = 99 + 495 = 594, so the positive predictive value is 99/594 = 1/6, about 17%. The odds form gives the same thing: prior odds of disease are 1:99, the positive likelihood ratio is sensitivity/(1 - specificity) = 0.99/0.05 = 19.8, and 1:99 times 19.8 gives posterior odds of about 1:5, i.e. 1/6. The lesson is that the base rate dominates: with a rare condition, even a very accurate test produces far more false alarms than true hits, so a single positive is a reason to confirm, not to conclude.
go deeper
Be ready to state that a positive result on a rare condition is usually a false positive, and to name prevalence as the missing ingredient people forget.
You are expected to produce the 17% yourself, either with a 10,000-person frequency table or with prior odds times the likelihood ratio, and to define sensitivity, specificity and PPV without mixing them up.
Show the operating consequence: pick the confirmatory step, argue that specificity is the binding constraint on a rare condition, and flag that repeating the same assay is not an independent second opinion.
Own the screening policy question. Decide who is inside the screened population at all, since choosing the prior is the biggest lever you have, and be able to defend the false-alarm volume you are about to create downstream.
## What is being asked The question mixes two different conditional probabilities, and confusing them is the single most common error in the whole topic. - **Sensitivity** is `P(test positive | disease)` -- among people who have the disease, what fraction does the test catch. Here 99%. - **Specificity** is `P(test negative | no disease)` -- among healthy people, what fraction the test correctly clears. Here 95%, so 5% of healthy people are falsely flagged. - **Prevalence** (the base rate) is `P(disease)` in the population being tested. Here 1%. - **Positive predictive value (PPV)** is `P(disease | test positive)` -- the thing the person holding the result actually cares about. Sensitivity and specificity are properties of the test. PPV is not: it depends on who you point the test at. ## The natural-frequency solution Counting people is faster and less error-prone than manipulating percentages. Imagine 10,000 people drawn from this population. | | Test positive | Test negative | Total | |---|---|---|---| | Disease | 99 | 1 | 100 | | No disease | 495 | 9,405 | 9,900 | | Total | 594 | 9,406 | 10,000 | - Diseased: 1% of 10,000 = 100. The test catches 99% of them, so 99 true positives and 1 false negative. - Healthy: 9,900. The test wrongly flags 5% of them, so 495 false positives and 9,405 true negatives. Then `PPV = 99 / (99 + 495) = 99/594 = 1/6 = 16.7%`. Five out of six positives are healthy people. The complementary number is reassuring: `NPV = P(no disease | negative) = 9,405/9,406`, about 99.99%. A negative from this test is close to conclusive; a positive is not. ## The Bayes and odds versions Bayes' theorem written out: ``` P(D|+) = P(+|D) P(D) / [ P(+|D) P(D) + P(+|not D) P(not D) ] = (0.99 x 0.01) / (0.99 x 0.01 + 0.05 x 0.99) = 0.0099 / (0.0099 + 0.0495) = 0.167 ``` The odds form is quicker and worth having memorised: **posterior odds = prior odds x likelihood ratio**. Prior odds are `0.01/0.99 = 1:99`. The positive likelihood ratio is `sensitivity / (1 - specificity) = 0.99/0.05 = 19.8`. Posterior odds are `19.8/99 = 0.2 = 1:5`, and odds of 1:5 convert to probability `1/(1+5) = 1/6`. Same answer, no table. ## Why the base rate dominates The two error streams are drawn from pools of wildly different size. True positives come from a pool of 100; false positives come from a pool of 9,900. Even a 5% error rate on the big pool swamps a 99% hit rate on the small one. This is the **base-rate fallacy**: people quote the test's accuracy and ignore how rare the condition is, then read a positive as if it meant 99% likelihood of disease. A useful rule of thumb: a positive result multiplies your odds by the likelihood ratio. It does not set them. Starting odds of 1:99 multiplied by 20 are still only about 1:5. ## What actually moves PPV Three levers, in order of leverage for a rare condition: 1. **Raise specificity.** Going from 95% to 99.5% cuts false positives from 495 to about 50, lifting PPV from 17% to about 67%. On a rare condition, the false-positive rate is the binding constraint. 2. **Raise the prior.** Screen a group with 20% prevalence instead of 1% and the same test yields `0.198/(0.198+0.04) = 83%` PPV. This is why symptoms, exposure history or a referral change the meaning of the identical result -- and why mass screening of a low-risk population is a different proposition from testing someone who already looks sick. 3. **Retest.** A second, mechanistically different confirmatory test applies another likelihood ratio to already-updated odds. Two positives from 1:5 odds and an LR of 20 give 4:1, about 80%. The caveat is that repeating the *same* test on the same person often re-triggers the same idiosyncratic cause of the first false positive, so the results are not conditionally independent and multiplying the likelihood ratios overstates the evidence. ## How to answer in the room Say the number, then show the frequency table, then name the principle. Interviewers are checking three things: that you did not report sensitivity when asked for PPV, that you can actually do the arithmetic under mild pressure, and that you draw the operational conclusion -- with a low base rate, a positive is a trigger for confirmation, not a diagnosis.
- What happens to the answer if you run the same test on a high-risk group with 20% prevalence?PPV rises to about 83%. With 20% prevalence, 0.2 x 0.99 = 0.198 of the group are true positives and 0.8 x 0.05 = 0.04 are false positives, so 0.198/0.238 = 0.83. The test did not change; the prior did. That is exactly why the same result means something different for a screened volunteer and for a symptomatic patient.
- What is the negative predictive value of this test at 1% prevalence?About 99.99%. Of 10,000 people, 9,406 test negative and only 1 of them has the disease, so P(no disease | negative) = 9,405/9,406. A negative is far more informative than a positive here, because a 99% sensitivity leaves almost no diseased people on the negative side, and the negative side is huge.
- If you could improve only one of sensitivity or specificity, which would you pick here and why?Specificity. False positives are generated from the 9,900 healthy people, so shaving that 5% error rate has enormous leverage: raising specificity to 99.5% cuts false positives from 495 to about 50 and lifts PPV from 17% to roughly 67%. Improving sensitivity from 99% to 99.9% recovers at most one extra true positive out of 100.
Fishing with a net that catches almost every trout but also snags 5 in every 100 sticks. In a river full of sticks and nearly no trout, most of what is in the net is still sticks.
saying these in an interview costs you the question
- Says a positive means a 99% chance of having the disease
- Reports sensitivity when asked for positive predictive value
- Ignores prevalence entirely when interpreting the result
- Assumes 95% specificity means 5% of positives are wrong
- Answers P(positive | disease) when asked P(disease | positive)
- Treats predictive value as a fixed property of the test