When would you use Gaussian, multinomial or Bernoulli likelihoods in a naive Bayes classifier?
answer
- the skeleton is fixed, one factor varies
- ask what kind of number the feature is
- counts versus flags versus measurements
- one variant scores absent features too
- mean and variance per feature per class
basics
~20 sMatch the likelihood to the feature type. Gaussian suits continuous readings assumed normal within each class, multinomial suits non-negative counts such as term frequencies, and Bernoulli suits binary present/absent flags where an absence is itself evidence.
solid answer
~50 sNaive Bayes is a template, not one model: the class prior and the argmax stay fixed while the per-feature likelihood changes with the data type. **Gaussian** estimates a mean and variance per feature per class — the right choice for three continuous machine-sensor readings, provided each looks roughly unimodal within a class. **Multinomial** models non-negative counts, weighting each term's log-likelihood by how many times it occurs, which is why it fits newswire topic labelling over token counts. **Bernoulli** models each feature as a present/absent coin flip and, crucially, multiplies in a `(1 - p)` term for every feature that is *absent*; for 200 app-permission flags in malware triage, a missing permission is real evidence, and multinomial would ignore it. Because the model factorises per feature, you can also mix families across columns and just add their log terms.
go deeper
Be able to name the three families and the feature type each expects: continuous measurements, non-negative counts, binary flags. An interviewer will accept that mapping without the estimation formulas at this level.
Expect to state what is estimated per class in each case — mean and variance per feature, count ratios over a vocabulary, per-feature presence rates — and to explain why only Bernoulli scores the absent features.
Demonstrate diagnosis: recognise a within-class bimodal feature killing a Gaussian term, know the variance floor that stops a constant feature exploding, and be comfortable mixing families across columns of one model.
Frame it as where modelling effort belongs. Argue when a feature deserves a better likelihood family versus when the whole naive Bayes assumption has run out and a discriminative model is the cheaper path for the team.
## The shared skeleton Every naive Bayes variant computes the same thing: ``` log_score(class) = log P(class) + sum over features j of log P(x_j | class) ``` and predicts the argmax. The class prior comes from label counts, and the per-feature independence is assumed throughout. What varies between the named variants is only the family used for `P(x_j | class)` — so the choice is a modelling decision about what kind of quantity each feature is. ## Gaussian likelihood — continuous features For a continuous feature, assume that within each class it follows a normal distribution, and estimate a mean and a variance for that feature in that class from the training rows of that class alone. For three machine-sensor readings and two classes, that is 3 x 2 means and 3 x 2 variances, plus the prior. There is no covariance matrix: independence means each feature gets its own one-dimensional density, which amounts to a diagonal-covariance Gaussian per class. Two practical notes. First, a feature that is constant inside a class gives a variance of zero and an infinite density, so implementations floor every variance with a small positive constant — the Gaussian counterpart of smoothing a zero count. Second, the normality is assumed *within a class*, not overall; a feature that is bimodal across the whole dataset is fine if each class sits on one mode. What breaks it is a feature that is bimodal *inside* one class — say a sensor that idles at 20 and spikes at 80 in the same failure mode. The fitted mean lands near 50 with a huge variance, putting its density peak where the class never actually is. Remedies: transform the feature, split the class, or discretise the reading into bins and switch that column to a categorical likelihood. ## Multinomial likelihood — counts Here each feature is a non-negative count and each class has a probability distribution over the vocabulary. A document contributes ``` sum over distinct terms t of count(t) * log P(t | class) ``` so a term occurring five times contributes five times its log-likelihood: repetition is evidence. Estimation is the count ratio with additive smoothing so no term is exactly zero. The consequence worth stating in an interview is that multinomial scoring only touches terms that are **present**. A term absent from the document contributes nothing at all. For a long newswire story that is fine, because presence carries plenty of signal. ## Bernoulli likelihood — binary flags Bernoulli treats each feature in a fixed feature set as an independent yes/no event, with `p_j = P(feature j present | class)` estimated as the fraction of that class's training rows in which it occurs. Scoring runs over **every** feature, not just the present ones: ``` sum over j of [ x_j * log p_j + (1 - x_j) * log (1 - p_j) ] ``` The second term is the whole point: an absent feature contributes `log(1 - p_j)`, which is a strong penalty for a class where that feature is usually present. For mobile malware triage over 200 present/absent app-permission flags, the fact that an app does **not** request network access is genuinely informative, and Bernoulli encodes it while multinomial would let it pass silently. Smoothing applies here too, as `(rows_with_feature + alpha) / (rows_in_class + 2 * alpha)`; the `2 * alpha` reflects two possible outcomes rather than a vocabulary of size V. Note also that Bernoulli discards repetition entirely — five occurrences and one occurrence look identical. ## Choosing between them - Continuous measurements, roughly unimodal per class: Gaussian. - Counts where how many times matters, and documents long enough that presence carries the signal: multinomial. - Binary indicators, or short items over a modest feature set where absence is informative: Bernoulli. Binarising counts and feeding them to multinomial is not the same as Bernoulli: you lose the repetition information either way, but only Bernoulli adds the absence terms back. On very short texts, where there is barely any repetition to exploit, the explicit modelling of zeros is why Bernoulli can be competitive; on long documents multinomial generally has the advantage. ## Mixing families Because the log-score is a sum of independent per-feature terms, nothing forces one family across the whole feature set. A model over two continuous sensor readings and five binary flags can use Gaussian terms for the first two and Bernoulli terms for the rest and simply add them, keeping one shared class prior. That flexibility is a good sign that the candidate understands the skeleton rather than three memorised recipes.
- One sensor reading is bimodal within a single class — what does that do to Gaussian naive Bayes?It puts the fitted density in the wrong place. The estimated mean lands between the two modes with an inflated variance, so the likelihood peaks at values that class never actually produces and is too low at the values it does. Fix it by transforming the feature, splitting the class if the modes are really two regimes, or discretising the reading into bins and using a categorical likelihood for that column.
- Why can Bernoulli naive Bayes hold its own against multinomial on very short texts?Short items give the multinomial almost no repetition to exploit, so its main advantage — count weighting — is idle. Bernoulli meanwhile adds an explicit term for every feature that is absent, and with a small feature set those absence terms carry real signal. On long documents the balance tips back toward multinomial.
- Can one naive Bayes model use different likelihood families for different features?Yes. The log-score is a sum of per-feature terms under the independence assumption, so each column can use whichever family matches its type — Gaussian for continuous readings, Bernoulli for flags, multinomial for counts — sharing a single class prior. You just add the log terms together.
saying these in an interview costs you the question
- Feeds raw word counts into a Gaussian likelihood
- Says binarised counts under multinomial equal Bernoulli
- Claims Gaussian naive Bayes models feature correlations
- Believes absent features never contribute to any variant
- Ignores that Gaussian normality is assumed within each class