Why is naive Bayes often accurate even though its posterior probabilities are far off?
answer
- only the ordering decides the label
- log-ratios summed, duplicates enter twice
- amplification pushes an existing direction
- 0.99999 is not a probability here
- fine for labels, unsafe for thresholds
basics
~20 sClassification needs only which class scores highest, not the size of the score. Double-counted correlated evidence inflates the leading class, so a naive Bayes posterior of 0.99999 is meaningless, but the ordering of the classes — and therefore the predicted label — usually survives.
solid answer
~50 sThe decision rule is an argmax over class scores, so it depends only on the *ranking* of the classes, not on the values. Violating conditional independence distorts the values badly: every duplicated feature multiplies its likelihood ratio in again, so the score of whichever class that evidence favours is pushed further ahead than the data warrants. But the push is in the direction that feature already pointed, so when the redundant evidence agrees with the class that would have won anyway, the argmax is unchanged and only the confidence is exaggerated. That is why naive Bayes routinely posts respectable accuracy while reporting posteriors pinned at 0 or 1. The practical consequence: use the label, treat the score as an uncalibrated ranking number, and never plug a raw naive Bayes posterior into a cost-based threshold or blend it with another model's probabilities without a separate calibration step fitted on held-out data.
go deeper
Remember the headline: the predicted label comes from whichever class scores highest, so a distorted score can still give the right label. Know that naive Bayes probabilities are not trustworthy numbers.
Explain the mechanism in log space — one log-likelihood-ratio term per feature, duplicates adding their term twice — and connect that directly to posteriors pinned near 0 and 1.
Show you know which downstream uses this breaks: cost-based thresholds, cross-example confidence comparison, ensembling with other models' probabilities. Say what you would do instead before shipping.
Be ready to rule on when an uncalibrated but accurate classifier is acceptable in a product decision loop, and when the team must pay for calibrated probabilities or a different model class.
## Two different jobs a classifier's output can do A scoring model produces a number per class; what you do with the number decides how much its accuracy matters. - **Deciding a label** uses only the ordering: predict the class with the largest score. Any transformation that preserves the ordering leaves the prediction untouched. - **Deciding an action with a cost** uses the value: block the email if the probability of spam exceeds 0.95, escalate the case if the probability of default exceeds 0.3. Here the number is load-bearing. Naive Bayes is far better at the first job than the second, and the reason is exactly the broken independence assumption. ## Why the numbers go wrong Work in logs, which is how the model is evaluated in practice anyway. For two classes the decision reduces to the sign of ``` log(P(c1)/P(c2)) + sum over features of log(P(x_i|c1)/P(x_i|c2)) ``` Each feature contributes one log-likelihood-ratio term. Now suppose two features are near-duplicates — the tokens `free` and `free!!!` in a spam filter, or `fever` and `temperature above 38C` in a symptom classifier. In truth they are one piece of evidence, but the sum receives two nearly identical terms. The evidence enters twice, so the total moves roughly twice as far from zero as it should. Convert that back out of log space and the posterior collapses onto 0.99999 or 0.00001. Add ten correlated features from the same underlying signal and the exponent gets ten times the weight it deserves. This is why naive Bayes posteriors are famously extreme, and it is a property of the model's structure, not a symptom of too little training data — more data makes each individual likelihood estimate sharper without removing the double counting. ## Why the label survives anyway The distortion is not random noise; it is amplification along a particular direction. A duplicated feature multiplies the pull it was already exerting. If that pull was towards the correct class, the correct class simply wins by a larger margin than it deserves, and the argmax is identical. Predictions only change when the amplified evidence pulls towards a *different* class than the one the honest computation would have chosen and the amplification is large enough to overcome the honest margin. On many real problems the redundant features are redundant precisely because they are all reflecting the true class — several ways of saying "this looks like spam" — so their double counting reinforces the right answer. This robustness of the zero-one-loss decision under violated independence is a known and studied result in the machine-learning literature, not folklore; it is the standard reason naive Bayes remains a credible text-classification baseline decades after its assumption was known to be false. ## What this does and does not license It licenses reporting accuracy, using the predicted label, and shipping the model as a fast baseline. It does not license: - **Thresholding on the score.** A cutoff of 0.95 is meaningless when nearly every prediction reads above 0.99. In practice you end up picking a threshold empirically on a validation set, which works, but be honest that you are choosing an operating point on an uncalibrated score, not a probability. - **Expected-cost decisions.** If a false negative costs eight times a false positive, the arithmetic needs a real probability. A naive Bayes posterior is not one. - **Comparing confidence across examples.** The inflation is not a constant factor. It depends on how many redundant features happened to fire for that particular example, so one prediction at 0.9999 and another at 0.999 are not reliably ordered by true confidence. Rank-based use across examples is therefore less safe than the within-example argmax, even though both look like "just using the ordering". - **Feeding the score into a downstream model or an ensemble average** as if it were a probability, where the extremeness will dominate every other input. If you genuinely need probabilities, keep the naive Bayes score as a ranking signal and fit a separate mapping from score to probability on held-out data — that recalibration step is a distinct piece of work with its own methods. ## How to answer it Name the argmax property first, then explain the multiplicative amplification that wrecks the value, then be explicit about which downstream uses that breaks. Candidates who stop at "it works well in practice" without explaining the mechanism sound like they have read the folklore rather than the model.
- A product manager wants to act only on emails your naive Bayes filter scores above 0.95 spam. What do you tell them?That the 0.95 is not what they think it is: correlated tokens push almost every confident-looking email past 0.99, so the cutoff will not select the top 5% of risk. Either pick the operating point empirically on a validation set and describe it as a rank cutoff, or fit a calibration mapping on held-out data first and threshold the calibrated output.
- Is the overconfidence fixed by training on more data?No. More data sharpens each individual per-feature likelihood estimate, but the double counting is structural — the same evidence is still entered once per redundant feature, so the log-ratio sum is still inflated. The fixes are on the feature side (merge or drop redundant features) or downstream (recalibrate the score).
- Does the same robustness hold with twenty classes as with two?Less reliably. With many classes there are more competitors that inflated evidence can push past a modest leader, so the honest margin is more often overturned. The argmax is most robust when one class holds a wide lead over a small number of alternatives.
saying these in an interview costs you the question
- Treats naive Bayes posteriors as calibrated probabilities
- Says high accuracy proves the independence assumption holds
- Blames the overconfidence on too little training data
- Thinks averaging the per-feature probabilities would fix it
- Cannot say that the decision rule uses only the ordering