skip to content

Likelihood Variants and Smoothing

Gaussian, multinomial and Bernoulli forms each assume a different feature distribution, and additive smoothing stops one unseen word from zeroing a whole class. Text is the stock example.

on this pageshow

questions

4

How does Laplace smoothing stop one unseen word from zeroing a class score in naive Bayes?

level: juniorimportance: must knowfreq 76%

answer

  1. one missing count outvotes everything else
  2. a product dies on a single zero factor
  3. never-seen is not impossible
  4. add a pseudo-count before dividing
  5. denominator gains alpha times vocabulary size

basics

~20 s

Laplace smoothing adds one pseudo-count to every word-class pair before dividing, so no estimated likelihood is exactly zero. Without it, a word never seen in a class drives that class's whole product to zero whatever the other words say.

solid answer

~50 s

Multinomial naive Bayes estimates `P(word | class)` as that word's count in the class divided by the class's total token count. If a word never appeared in that class during training, the estimate is exactly 0, and because the score is a product over every token, one zero factor vetoes the class outright even when the other four thousand tokens point straight at it. That is an accident of a finite sample, not evidence the word is impossible. Laplace smoothing fixes it by adding a pseudo-count: `P(word | class) = (count + alpha) / (class_total + alpha * V)`, where `V` is the vocabulary size and `alpha = 1` is add-one. The `alpha * V` in the denominator is not optional; without it the per-class distribution no longer sums to one. `alpha` is a hyperparameter worth tuning, not a constant.

code

python · 19 lines
python
import math

# token counts per class from a tiny newswire training set
counts = {"sports":  {"goal": 12, "match": 9, "quarter": 3},
          "finance": {"quarter": 14, "bond": 7, "match": 1}}
vocab = ["goal", "match", "quarter", "bond"]
alpha = 1.0            # add-one; set to 0.0 to see the zero veto

def log_likelihood(doc, cls):
    total = sum(counts[cls].get(w, 0) for w in vocab)
    denom = total + alpha * len(vocab)          # denominator grows too
    return sum(math.log((counts[cls].get(w, 0) + alpha) / denom)
               for w in doc)

doc = ["match", "bond"]  # 'bond' was never seen in the sports class
for cls in counts:
    print(cls, round(log_likelihood(doc, cls), 3))
# sports -4.362 / finance -3.744  -> finance wins, but sports is still
# in the running; with alpha = 0.0 sports would be an outright zero.

go deeper

for a junior

Be ready to say why a probability of exactly zero is fatal when the score is a product, and to write the add-one formula with the corrected denominator. Knowing the phrase Laplace smoothing without the formula will not carry the answer.

for a middle

Expect to explain what alpha does at both extremes: near zero you recover raw counts and the zeros, very large and every likelihood flattens toward uniform so only the prior speaks. Also separate smoothing from log-space arithmetic.

for a senior

Show that you tune the smoothing constant rather than inheriting add-one, and that you have a policy for out-of-vocabulary tokens at scoring time. Interviewers like hearing how sparse training data changes the choice.

for a principal

Own the tradeoff as a data-collection question: heavy smoothing is what you pay for a thin corpus, so argue when more labelled text beats tuning a constant, and how the choice interacts with vocabulary size across the whole pipeline.

## What is being multiplied Multinomial naive Bayes scores a document by combining a class prior with one likelihood term per token: ``` score(class) = P(class) * product over tokens t of P(t | class) ``` and predicts the class with the largest score. Every factor is a probability between 0 and 1, and the factors are multiplied, which is what makes a single zero catastrophic: `anything * 0 = 0`. ## Where the zero comes from The unsmoothed (maximum-likelihood) estimate of a word likelihood is ``` P(t | class) = count(t, class) / sum over vocabulary of count(t', class) ``` Suppose you train a topic classifier on newswire stories and the word `bond` happens never to occur in a story labelled `sports`. Then `count(bond, sports) = 0`, the estimate is exactly 0, and any test document containing `bond` gets `score(sports) = 0` — no matter how many times it says `goal`, `match` or `referee`. One word that the training sample simply failed to show you gets a veto over every other piece of evidence. This is a sampling artifact. Language is heavy-tailed: most words are rare, so in any finite corpus a large fraction of legitimate word-class pairs have a count of zero. Treating "I never saw it" as "it is impossible" is the mistake. ## Additive (Laplace / Lidstone) smoothing The fix is to pretend you saw every word a little bit in every class: ``` P(t | class) = (count(t, class) + alpha) / (N_class + alpha * V) ``` where `N_class` is the total token count in that class and `V` is the vocabulary size. With `alpha = 1` this is Laplace or add-one smoothing; a general `alpha` is often called Lidstone smoothing. Two details candidates get wrong: - **The denominator must grow too.** You added `alpha` to each of `V` numerators, so the normaliser gains `alpha * V`. Adding `alpha` to the numerator alone leaves a "distribution" summing to more than one and silently biases classes with large vocabularies. - **Smoothing is per class.** Each class has its own `N_class`, so the same unseen word gets a slightly different smoothed probability in each class. Intuitively, `alpha` is a number of imagined prior sightings of every word in every class. It pulls each estimate away from the raw count ratio and toward the uniform distribution `1/V`. ## Choosing alpha `alpha` is a hyperparameter, and the two extremes bracket the tradeoff: - `alpha -> 0` returns to raw counts. Rare words get extreme estimates, and zeros come back. - `alpha` very large flattens every likelihood toward `1/V`. The words stop discriminating, and the prediction collapses toward whatever the class prior says. That is underfitting. Large corpora with plenty of counts per class typically want a small `alpha`; small or very sparse training sets want more. Tune it by cross-validation like any other hyperparameter rather than leaving add-one as an unexamined default. ## Zeros and underflow are two different problems A related but distinct numerical issue: multiplying four thousand token probabilities, each around `1e-4`, gives a number near `1e-16000`, far below what double-precision floating point can represent (roughly `1e-308`). It silently becomes 0 and every class ties. The cure is to work with a sum of logs: ``` log_score(class) = log P(class) + sum over tokens of log P(t | class) ``` Logs are monotonic, so the argmax is unchanged, and adding a few thousand numbers of moderate size is numerically safe. Logs do **not** rescue you from a zero count: `log 0` is negative infinity, which vetoes the class exactly as `0` did. Smoothing cures unseen counts; logs cure underflow. A real implementation needs both. ## Words unseen in every class A token that appears in the test document but in no training document at all is out-of-vocabulary. Its smoothed probability is `alpha / (N_class + alpha * V)` — not identical across classes, because `N_class` differs, so it mildly favours whichever class had fewer training tokens. Because the token carries no discriminative information, the usual practice is to score only tokens inside the training vocabulary and drop the rest, or to fold them all into a single reserved bucket that is itself smoothed. ## Other likelihood families The same idea appears wherever a probability can be estimated as exactly zero. A Bernoulli likelihood over present/absent flags smooths as `(docs_with_feature + alpha) / (docs_in_class + 2 * alpha)` — the `2 * alpha` because the feature has two outcomes. A Gaussian likelihood cannot produce a zero probability, but it can produce a zero variance when a feature is constant within a class, which blows the density up; the analogous guard is to add a small floor to every variance.

  • You already take logs to avoid underflow — does that also solve the zero-count problem?
    No. Logs solve a floating-point problem: four thousand small factors multiplied together underflow to zero, while their logs sum safely. But `log 0` is negative infinity, so an unsmoothed zero count still vetoes the class. Smoothing and log-space arithmetic fix two different failures and you need both.
  • How would you choose the smoothing strength rather than defaulting to add-one?
    Treat it as a hyperparameter and cross-validate it. Small values keep the raw counts and their sharpness, which suits large corpora with dense counts; larger values suit small or sparse training sets. Watch the two failure ends: near zero the zeros return, and very large values flatten every likelihood toward uniform so only the prior decides.
  • What happens to a token in the test document that appeared in no training class at all?
    It gets the pure pseudo-count probability `alpha / (N_class + alpha * V)`, which is not quite equal across classes because each class total differs, so it slightly favours the class with fewer training tokens. Since it carries no evidence either way, the usual practice is to restrict scoring to the training vocabulary or route all unknown tokens into one smoothed bucket.

Unsmoothed counts let one juror who has never met the defendant cast a veto. Smoothing takes the veto away and turns it back into a vote.

saying these in an interview costs you the question

  • Says a zero probability just lowers the score a little
  • Adds alpha to the numerator but not the denominator
  • Thinks taking logs removes the zero-probability problem
  • Claims smoothing is only needed for the class prior
  • Treats add-one as a fixed rule rather than a tunable constant

context

open as a page

When would you use Gaussian, multinomial or Bernoulli likelihoods in a naive Bayes classifier?

level: middleimportance: must knowfreq 66%

basics

~20 s

Match the likelihood to the feature type. Gaussian suits continuous readings assumed normal within each class, multinomial suits non-negative counts such as term frequencies, and Bernoulli suits binary present/absent flags where an absence is itself evidence.

open as a page

Why does naive Bayes output posteriors near 0 or 1 even when its accuracy is only moderate?

level: seniorimportance: should knowfreq 48%

basics

~20 s

The score adds thousands of log-likelihood terms, so small per-feature errors compound and the log-odds land far from zero; exponentiating then saturates the posterior at 0 or 1. Treat the number as a ranking score, not a probability.

open as a page

Your naive Bayes prior comes from a 90/10 label count but deployment runs 50/50 — what do you change?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Swap the prior, do not retrain. The class prior enters the score as one additive log term, so replacing the training log-prior with the deployment class mix corrects the model, provided the per-class feature distributions themselves have not changed.

open as a page