skip to content

Why does naive Bayes output posteriors near 0 or 1 even when its accuracy is only moderate?

level: seniorimportance: should knowfreq 48%

answer

  1. thousands of terms, one sum
  2. the sigmoid saturates by log-odds ten
  3. bias compounds, it does not cancel
  4. ranking survives, magnitude does not
  5. accuracy and calibration are different properties

basics

~20 s

The score adds thousands of log-likelihood terms, so small per-feature errors compound and the log-odds land far from zero; exponentiating then saturates the posterior at 0 or 1. Treat the number as a ranking score, not a probability.

solid answer

~50 s

The posterior comes from summing one log-likelihood term per feature. With four thousand tokens, the class log-odds is a sum of four thousand small quantities, and any systematic bias in them — most of it from treating overlapping features as independent evidence — accumulates instead of cancelling. A log-odds of plus or minus 40 is routine, and `1 / (1 + exp(-40))` is 1 to sixteen decimal places. So the ordering of documents by score is usually informative while the magnitude is not: 0.99999 versus 0.99 tells you which ranks higher, not that one is a hundred times surer. Practical consequences: never feed these numbers into an expected-value calculation, do not compare scores across documents of very different length, and if you need real probabilities, fit a separate calibration step on held-out data rather than trusting the raw output.

go deeper

for a junior

Recall that naive Bayes scores are useful for deciding and ranking but should not be quoted as real probabilities. Knowing that much protects you from promising a stakeholder a confidence number that does not mean what it says.

for a middle

Explain the mechanism: the posterior is a sigmoid of a sum of thousands of log-ratio terms, and the sigmoid is already at 0.99995 by a log-odds of ten. Be able to say why the terms compound instead of cancelling.

for a senior

Show the operational consequences you have actually hit — expected-value logic breaking, a nominal 0.5 threshold being meaningless, scores incomparable across item lengths — and say that the fix is a separate calibration step, not more data.

for a principal

Own the decision of whether the system needs probabilities at all. If it does, argue for a model or a calibration stage that supplies them; if ranking suffices, defend keeping the cheap model and forbidding downstream consumers from reading its scores as risk.

## The arithmetic behind the saturation For two classes, the decision is driven by the log-odds: ``` log_odds = log P(pos)/P(neg) + sum over features j of log [ P(x_j | pos) / P(x_j | neg) ] ``` and the reported posterior is `1 / (1 + exp(-log_odds))`. The prior contributes one term. The features contribute one term each — four thousand of them for a long document. The sigmoid saturates fast. A log-odds of 5 already gives 0.993; 10 gives 0.99995; 40 gives 1.0 to the limit of double precision. To produce a moderate, honest-looking 0.7 the whole sum would have to land near 0.85, which a sum of thousands of terms essentially never does. So even a model that is right 85% of the time reports near-certainty on almost every item. ## Why the terms do not cancel If every per-feature log-ratio were an unbiased, independent estimate of the true evidence, errors would partly cancel and the sum would be roughly right. Two things stop that. First, the model counts every feature as a fresh, independent piece of evidence. Real feature sets overlap, so evidence that should be counted once is counted several times, and the log-odds is pushed outward rather than being pulled back. This is a bias, not noise, so it grows with the number of features rather than averaging away. Second, each log-ratio is estimated from finite counts, and the estimates that are most extreme are the ones estimated from the fewest observations. Smoothing bounds how extreme any single term can be, which helps a little, but with thousands of terms the sum still runs away. The crucial point for an interview: this is a defect of the reported *probability*, not necessarily of the *decision*. More training data does not fix it, because it is bias in how evidence is combined rather than variance in the estimates. ## What the number can and cannot be used for **Usable — ordering.** Ranking items by score is generally meaningful: the higher-scored item is the one the model considers more likely positive, and rank-based evaluation still tells you something real about the model. **Not usable — the value itself.** A reported 0.9999 does not mean one error in ten thousand. Anything that consumes the number as a probability breaks: expected-value or cost-weighted decisions, blending the score with another model's probability, abstain-if-unsure thresholds, or reporting a confidence to a user. **Not usable — cross-item comparison of magnitude.** A long document accumulates more log-likelihood terms than a short one, so its score is pushed further out purely by length. Comparing the *sizes* of two scores across items of very different length compares document lengths as much as evidence. **Fragile — the 0.5 threshold.** Since the scores pile up at the ends, the fraction of items between 0.4 and 0.6 is tiny and moving the threshold inside that range changes almost nothing, while the operating point you actually want may sit at 0.999999. Choose the operating point from the ranking on held-out data, not from the nominal 0.5. ## Things that do not fix it - **More data.** The problem is bias in how evidence is combined, so more rows sharpen the same wrong combination. - **Stronger smoothing.** Raising the smoothing constant shrinks every individual term toward neutrality and does mildly temper the extremes, but it cannot undo the compounding across thousands of terms, and pushed far enough it just destroys the signal. - **Switching likelihood family.** Gaussian, multinomial and Bernoulli variants all saturate; the effect comes from summing many terms, not from which family supplies them. What does help is reducing the number of features that enter the sum — a model over 30 features saturates far less than one over 4,000 — and, when genuine probabilities are required, fitting a separate calibration mapping from raw score to probability on held-out data. That calibration step is a topic in its own right; the thing to demonstrate here is knowing that the raw naive Bayes posterior needs one. ## How to talk about it A strong answer separates three claims that candidates routinely merge: the model may be *accurate*, its scores may be well *ordered*, and its probabilities are nonetheless badly *calibrated*. Accuracy and calibration are different properties, and naive Bayes is the standard example of a model that can have the first without the second.

  • Does adding more features make the saturation better or worse?
    Worse. Every extra feature adds another log-ratio term, so the log-odds drifts further from zero and the posterior gets more extreme without getting more correct. A naive Bayes model over 30 features reports far more moderate scores than the same model over 4,000, which is one reason feature-count reduction can improve the usability of the scores even when accuracy barely moves.
  • Two items score 0.99999 and 0.99. What can you legitimately conclude?
    Only the ordering: the model ranks the first ahead of the second. The gap between the two numbers is not a meaningful ratio of likelihoods, and if the two items differ a lot in feature count — a long document versus a short one — part of that gap is just the longer item accumulating more terms.
  • Would stronger smoothing tame the overconfidence?
    Only slightly. A larger smoothing constant pulls each individual likelihood toward uniform, which caps how extreme any single term can be, but the posterior is a sum over thousands of terms and their compounding survives. Push the constant high enough to flatten the scores and you have destroyed the discriminative signal along with them.

saying these in an interview costs you the question

  • Reads a 0.9999 score as a one-in-ten-thousand error rate
  • Says more training data will fix the calibration
  • Calls it overfitting rather than miscalibration
  • Treats accuracy and calibration as the same property
  • Compares score magnitudes across documents of very different length

context