skip to content

Why can naive Bayes beat logistic regression on a small training set but lose on a large one?

level: middleimportance: should knowfreq 42%

answer

  1. bias against variance, indexed by sample size
  2. one converges fast to a worse ceiling
  3. simple per-class summaries versus a fitted boundary
  4. sample complexity: log d against d

basics

~20 s

Naive Bayes trades bias for variance: its generative assumptions are usually wrong, putting a floor under its error, but its per-class estimates stabilise after very few labels. Logistic regression has a lower floor and needs more data to reach it.

solid answer

~50 s

The two sit at different points on the bias-variance tradeoff. Naive Bayes fits a class prior plus a small set of per-class feature statistics; each one is estimated from many examples at once, so the model stabilises after very few labels - but its assumptions about how features are distributed inside a class are usually wrong, which puts a floor under its error. Logistic regression optimises the labelling criterion directly, so its floor is lower, yet with few labels its coefficients are high-variance and it can be beaten. Ng and Jordan's 2001 comparison made this precise: the generative model approaches its higher asymptotic error after a number of examples that grows only logarithmically in the number of features, while the discriminative model needs roughly linearly many. With 200 labelled memos I would reach for naive Bayes; at 200,000 I would expect logistic regression to have overtaken it.

go deeper

for a junior

Recall the headline: fewer labels favour the model with stronger assumptions, more labels favour the one that optimises the label criterion directly. Naming the two families correctly matters more than quoting rates.

for a middle

Explain the mechanism in bias-variance terms - low-variance per-class summaries with a bias floor, against a low-bias fit with noisy coefficients - and state that the generative model converges faster to a worse ceiling.

for a senior

Show how you would decide empirically: learning curves on a fixed held-out set, reading both the current ordering and the remaining slope, and revisiting the family choice as labelled data accumulates.

for a principal

Frame this as an investment question. Whether to buy more labels, add regularisation, or accept a biased model depends on where the crossover sits for your feature width and labelling cost, and that call should be revisited on evidence.

## The claim, stated carefully Compare two classifiers on the same features: a generative one that fits a class prior and a per-class feature model, and a discriminative one that fits `p(y|x)` directly by optimising the labelling criterion. The empirical and theoretical result - the Ng and Jordan comparison of naive Bayes against logistic regression, 2001 - has two halves, and candidates usually remember only one: 1. The generative classifier typically has a **higher asymptotic error**: with unlimited data, the discriminative one wins. 2. The generative classifier **converges to its own asymptotic error much faster**: with limited data, it can be well ahead. Put those together and you get a crossover. On a learning curve - held-out error against training-set size - the generative curve starts lower and flattens early; the discriminative curve starts higher, keeps falling, and eventually crosses underneath. ## Why the generative model stabilises so fast Its parameters are summaries. A class prior is a count of labels. A per-class feature statistic is an average over the examples of that class. Every parameter is estimated from a large slice of the data independently of the others, so each has low variance and none of them has to be searched for - there is no optimisation coupling them together. Ten examples per class already give usable estimates, which is why these models are the reflex choice when labelling is expensive. The price is **bias**. The model has committed to a particular story about how feature vectors are distributed inside a class, and real data almost never follows that story. No amount of extra data repairs a wrong assumption; it only makes the wrong parameters more precisely estimated. That is the floor under the error. ## Why the discriminative model needs more data but ends up lower Logistic regression makes no claim about how features are distributed. It fits coefficients by optimising a criterion tied directly to predicting the label, so it is free to bend whatever way the data says the log-odds bend. That is a weaker set of assumptions, hence lower asymptotic error - but a weaker set of assumptions means more to learn from data, so the coefficients are noisy when examples are scarce. With more features than informative examples, an unregularised fit can chase noise and even separate the training data perfectly while generalising badly. ## The sample-complexity statement The sharper version of "faster" is in terms of the number of features `d`: the generative fit needs a number of labelled examples growing on the order of `log d` to come close to its asymptotic error, while the discriminative fit needs on the order of `d`. Read that as a scaling statement, not a formula to plug numbers into. Its practical content is: **as you add features, the discriminative model's appetite for labels grows far faster than the generative model's.** A wide, label-poor problem is exactly where the generative side shines. ## Where the crossover actually sits There is no universal number of examples. The crossover moves with: - **How wrong the generative assumptions are.** Assumptions close to the truth raise the generative asymptote hardly at all, and the crossover moves far to the right - occasionally past any sample size you will ever have. If the generative model were *correctly* specified, it would not be worse asymptotically at all; the whole ordering depends on misspecification, which is simply the normal case. - **Feature count.** More features push the discriminative model's requirement up faster than the generative model's. - **Regularisation.** Penalising the discriminative fit's coefficients cuts its variance and pulls the crossover left, which is the standard way to make it competitive earlier. Regularisation is itself a way of injecting bias on purpose, which is the same trade under another name. ## How to find it on your own problem Do not argue about it - measure it. Fit both families at several training-set sizes (say 10%, 25%, 50%, 100% of what you have), evaluate each on the same held-out data, and plot two learning curves. Three readings matter: which curve is lower now, whether the discriminative curve is still falling at your current size, and how much slack the generative curve has left. If the discriminative curve is still dropping steeply at full data, more labels are worth buying and the family will likely change once you have them. ## The interview answer in one move When asked which model for a document classifier with 200 labelled internal memos, say: the generative one now, because at that size variance dominates and its bias is the cheaper error to pay. Then add that if the same task grows to 200,000 labelled memos, you expect the discriminative model to be ahead and you would re-run the learning curves rather than assume the original choice still holds. Naming both halves of the tradeoff, and naming the measurement, is what separates the answer from a memorised slogan.

  • Does the crossover always exist?
    No. The ordering assumes the generative model's assumptions about the features are wrong, which is usual but not guaranteed. If they are close to true, its asymptotic error is not higher and the discriminative model may never overtake it. Equally, if you never accumulate enough labels, the crossover exists in theory and is irrelevant in practice - you live on the left half of the curve.
  • How would you locate the crossover on your own dataset?
    Plot learning curves. Fit both families at several training-set sizes, score each on the same held-out data, and plot held-out error against size. The crossing point is where the discriminative curve passes under the generative one. The slope at your current size is the more actionable reading: if the discriminative curve is still falling, more labels will change the answer.
  • You have 200 labelled memos now and expect 200,000 within a year. What do you ship?
    Ship the generative model now - at 200 examples it is the lower-error choice and it is cheap to fit. Ship the evaluation harness with it: a fixed held-out set and the learning curves. Then revisit as labels accumulate rather than on a calendar. The decision is a checkpoint on the curve, not a permanent architectural commitment.

A sketch map drawn from memory is usable after one walk through a town and never gets much better. A proper survey needs many measurements before it is trustworthy, but it eventually beats the sketch on every street.

saying these in an interview costs you the question

  • Says naive Bayes is simply the worse model everywhere
  • Claims more data always favours the model with more assumptions
  • Credits the small-sample win to faster training rather than lower variance
  • Treats the crossover as a fixed universal number of examples
  • Forgets that the generative model's advantage is a bias it never sheds

context