skip to content

Naive Bayes Classifiers

A generative classifier that multiplies per-feature likelihoods by a class prior and takes the argmax. Interviewers like it because it stays accurate even when its core assumption is plainly false.

on this pageshow

explore

questions

7

What does naive Bayes' conditional independence assumption claim about the features?

level: juniorimportance: must knowfreq 80%

answer

  1. the class label is already known
  2. a joint likelihood becomes a product
  3. one small table per feature per class
  4. conditional, not overall, independence
  5. 'free' and 'free!!!' counted twice

basics

~20 s

Naive Bayes assumes that once the class label is known, the features carry no further information about each other. That licenses replacing the whole likelihood P(features | class) with a product of one-feature terms P(x_i | class).

solid answer

~50 s

Naive Bayes scores a class as `prior(class) * P(x_1 | class) * P(x_2 | class) * ... * P(x_d | class)` and predicts the argmax. That product is only valid if the features are conditionally independent given the class — knowing the label, learning the value of one feature tells you nothing about another. The point of the assumption is tractability: modelling the full joint likelihood over `d` binary features needs on the order of `2^d` numbers per class, while the factorised version needs one small table or one Gaussian per feature per class, each estimated from simple counts or means. The assumption is called naive because it is almost always false. A spam filter with separate features for `free` and `free!!!` treats one signal as two independent witnesses and multiplies the same evidence twice.

go deeper

for a junior

Be ready to write the scoring rule — prior times a product of one-feature likelihoods — and to say the word conditional out loud. Have one duplicated-feature example ready, such as two spam tokens that mean the same thing.

for a middle

Explain why the factorisation is needed at all: give the parameter-count argument, and distinguish correlation that the class explains from residual correlation inside a class, which is the part the model ignores.

for a senior

An interviewer expects you to move straight from the assumption to its operational consequence — inflated posteriors from double-counted evidence — and to say which features in a proposed feature set you would suspect of restating each other.

for a principal

Own the framing that this is a deliberate bias-for-estimability trade. Be able to say when that trade is the right call for a team: tiny labelled sets, very wide sparse inputs, or a baseline that has to ship this week.

## The classifier the assumption serves Naive Bayes is a generative classifier: it models how each class produces feature values, then inverts that with Bayes' rule to score classes for a new example. For a class `c` and a feature vector `x = (x_1, ..., x_d)` it computes a score proportional to ``` score(c) = P(c) * P(x | c) ``` where `P(c)` is the **prior** (how common the class is, usually the share of training rows with that label) and `P(x | c)` is the **likelihood** (how typical this particular feature vector is for that class). The predicted label is the class with the largest score — the **argmax**. The denominator of Bayes' rule, `P(x)`, is the same for every class, so it can be ignored for prediction and only matters if you want the posterior to be a real probability. ## Why the likelihood has to be factorised The hard term is `P(x | c)` — the probability of a whole *combination* of feature values. Estimate it directly and you must fill in one number for every distinct combination. With 30 binary features that is roughly a billion cells per class, and your training set has, say, 50,000 rows. Almost every cell would be estimated from zero or one example. The model would be unusable not because the maths is wrong but because there is no data to fit it. The conditional independence assumption cuts that down. It states ``` P(x_1, x_2, ..., x_d | c) = P(x_1 | c) * P(x_2 | c) * ... * P(x_d | c) ``` Each factor is a one-dimensional quantity: for a binary feature, the fraction of class-`c` rows in which it fires; for a continuous feature, a mean and variance per class. Thirty binary features now need sixty numbers per class instead of a billion, and each is estimated from the whole class, not from a sliver of it. That is the entire trade the model makes — a false structural assumption bought in exchange for parameters you can actually estimate. ## Conditional, not marginal The most common misreading is that naive Bayes assumes the features are independent overall. It does not, and that distinction matters. Height and weight are strongly correlated across a mixed population of adults and children, but much of that correlation exists *because* both depend on which group a person is in. Naive Bayes is perfectly comfortable with that: the class explains it. What the model assumes away is the *leftover* correlation **inside** each class — that among adults only, height still tells you nothing about weight. That residual dependence is what the factorisation ignores, and it is what causes trouble. ## Where it breaks in practice The assumption fails hardest when two features are near-restatements of each other. A spam model with one feature for the token `free` and another for `free!!!` has, in effect, one piece of evidence entered twice; the product treats them as two independent witnesses who happened to agree, and the score moves twice as far in log terms as the evidence warrants. A symptom classifier that records both `fever` and `temperature above 38C` has the same defect — the second feature is nearly a function of the first, so its likelihood is counted as fresh information when it is not. Two credit-bureau fields such as the number of open cards and the total credit limit move together for the obvious reason, and a naive Bayes model treats their agreement as independent corroboration. The visible symptom of this double counting is a posterior that is far too extreme: scores like 0.99999 on examples that are genuinely borderline, because each duplicated factor pushes the winning class further ahead. It does not follow that the prediction is wrong — the ordering of the classes often survives distortion that ruins the number — but the number itself should not be read as a probability. ## What to say in an interview State the factorisation, name the word *conditional* explicitly, give the parameter-count argument for why the assumption is made at all, and offer one concrete duplicated-feature example. Then volunteer that the assumption is essentially never true and that the model is used anyway — that is the opening the interviewer is usually fishing for.

  • Why does the assumption make naive Bayes trainable on very little data?
    It collapses the parameter count. The unfactorised likelihood over `d` binary features needs on the order of `2^d` cells per class, each starved of data. The factorised form needs `d` one-feature estimates per class, and every one of them is computed from all rows of that class. That is why naive Bayes is a credible baseline on a few hundred labelled examples.
  • Does naive Bayes assume the features are independent of each other overall?
    No — only within a class. Two features can be strongly correlated across the full dataset because both depend on the label, and the model handles that correctly, since the class already explains the shared movement. What it assumes away is the correlation that remains after you condition on the class, and that is the part real features rarely honour.
  • Is the assumption ever actually true on real data?
    Essentially never, and nobody claims otherwise. It can be approximately true for a small, deliberately deduplicated feature set drawn from genuinely different sources — say one behavioural signal, one demographic field, one device attribute. Treat it as a modelling convenience that buys estimability, not as a claim you are asserting about the data-generating process.

Twelve jurors who all read the same newspaper article and reach the same conclusion. Counted as twelve independent opinions, the case looks overwhelming; really there is one piece of evidence repeated twelve times.

saying these in an interview costs you the question

  • Says naive Bayes assumes the features are independent overall
  • Claims the assumption must hold or the model is worthless
  • Confuses it with assuming the classes are independent
  • Says it assumes every feature is normally distributed
  • Cannot state the factorisation of the likelihood

context

open as a page

How does Laplace smoothing stop one unseen word from zeroing a class score in naive Bayes?

level: juniorimportance: must knowfreq 76%

basics

~20 s

Laplace smoothing adds one pseudo-count to every word-class pair before dividing, so no estimated likelihood is exactly zero. Without it, a word never seen in a class drives that class's whole product to zero whatever the other words say.

open as a page

When would you use Gaussian, multinomial or Bernoulli likelihoods in a naive Bayes classifier?

level: middleimportance: must knowfreq 66%

basics

~20 s

Match the likelihood to the feature type. Gaussian suits continuous readings assumed normal within each class, multinomial suits non-negative counts such as term frequencies, and Bernoulli suits binary present/absent flags where an absence is itself evidence.

open as a page

Why is naive Bayes often accurate even though its posterior probabilities are far off?

level: middleimportance: should knowfreq 58%

basics

~20 s

Classification needs only which class scores highest, not the size of the score. Double-counted correlated evidence inflates the leading class, so a naive Bayes posterior of 0.99999 is meaningless, but the ordering of the classes — and therefore the predicted label — usually survives.

open as a page

Why does naive Bayes output posteriors near 0 or 1 even when its accuracy is only moderate?

level: seniorimportance: should knowfreq 48%

basics

~20 s

The score adds thousands of log-likelihood terms, so small per-feature errors compound and the log-odds land far from zero; exponentiating then saturates the posterior at 0 or 1. Treat the number as a ranking score, not a probability.

open as a page

In naive Bayes, which correlated features actually flip the predicted class?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Only redundant evidence that favours a class other than the honest winner, and whose repeated counting is large enough to overcome the honest margin. Correlated features that reinforce the class that would have won anyway just inflate confidence without changing the label.

open as a page

Your naive Bayes prior comes from a 90/10 label count but deployment runs 50/50 — what do you change?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Swap the prior, do not retrain. The class prior enters the score as one additive log term, so replacing the training log-prior with the deployment class mix corrects the model, provided the per-class feature distributions themselves have not changed.

open as a page