skip to content

What is the difference between a generative and a discriminative classifier?

level: juniorimportance: must knowfreq 62%

answer

  1. which distribution actually gets fitted
  2. the boundary, or the classes themselves
  3. only one of them can produce new x
  4. p(y|x) direct versus p(y) times p(x|y)

basics

~20 s

A discriminative classifier models p(y|x) - the label given the features - directly. A generative classifier models the joint p(x,y), normally as a class prior p(y) times a per-class feature distribution p(x|y), and inverts it to score labels.

solid answer

~50 s

A discriminative classifier estimates `p(y|x)` directly - or, like a support vector machine, just a boundary - and spends all of its capacity on separating the labels. A generative classifier estimates the joint `p(x,y)`, normally as a class prior `p(y)` times a class-conditional `p(x|y)` fitted from that class's examples alone, then recovers the label score by inverting it: `p(y|x)` is proportional to `p(y)p(x|y)`. Naive Bayes and linear discriminant analysis are generative; logistic regression, support vector machines, random forests and gradient boosting are discriminative. The practical consequence is that the generative model also holds a distribution over the features themselves, so it can produce new feature vectors and reason about inputs it has never seen, while the discriminative model can only answer the labelling question - but it answers it by optimising exactly the quantity it is graded on.

go deeper

for a junior

Be ready to define both in one sentence each and name two examples per side. Knowing that naive Bayes is generative and logistic regression is discriminative gets you most of the way.

for a middle

Explain the factorisation p(x,y) = p(y)p(x|y) and how the inversion recovers a label score, and say what each family spends its capacity on. Expect to be asked which side a named model sits on.

for a senior

Show that the choice has operational consequences: what you can do at scoring time with a distribution over features, how each family behaves as labels accumulate, and how a class-balance change is absorbed.

for a principal

Own the framing that a generative model buys capabilities beyond the label at the cost of extra assumptions about data you do not need to describe. Be able to say when that trade is worth paying for across a portfolio of models.

## Two routes to the same decision A classifier is given a feature vector `x` and must choose a label `y`. There are two fundamentally different things you can fit in order to do that, and the split between them is the oldest taxonomy in supervised learning. **Discriminative** methods model `p(y|x)` directly: the probability of each label *given* the features. Some do not even go that far - a support vector machine or a perceptron learns only a decision boundary, a function that says which side of a surface a point falls on, with no probability attached. What all of them share is that training optimises a criterion tied to labelling: get the label right, or get the predicted label probability close to the observed one. Logistic regression, support vector machines, decision trees, random forests, gradient-boosted trees and neural networks are all discriminative. **Generative** methods model the joint distribution `p(x, y)` - how labelled examples as a whole are produced. In practice nobody fits a joint density in one piece; it is factored as ``` p(x, y) = p(y) * p(x|y) ``` so training splits into two easy halves: estimate `p(y)`, how common each class is, and estimate `p(x|y)`, a model of how feature vectors are distributed *inside* each class, fitted from that class's examples only. Prediction then inverts the factorisation: ``` p(y|x) is proportional to p(y) * p(x|y) ``` and you take the label with the largest value. Naive Bayes, linear and quadratic discriminant analysis, and hidden Markov models for sequences are generative. ## Where the name comes from `p(x|y)` is literally a recipe for producing data: draw a class from `p(y)`, then draw a feature vector from that class's distribution, and you have manufactured a new labelled example. That capability - not the ability to write text or paint pictures - is what "generative" means here. Modern generative media models are the same idea scaled up, but in classical ML the word is a statement about *which distribution was estimated*, nothing more. ## What each one spends its effort on The discriminative model spends everything it has on the region where the classes meet. Points deep inside a class contribute little, because moving the boundary does not change how they are labelled. The generative model, by contrast, must describe each class's entire cloud of points - including the parts far from any boundary, which have no influence on the decision at all. That is effort spent on structure the decision does not need, and every assumption it makes about that structure is an assumption that can be wrong. The payoff is that the generative model ends up holding something richer than a decision rule: a description of the data itself. It can produce new feature vectors, evaluate how typical an input is, and re-use `p(x|y)` when the class balance changes. The discriminative model has none of that; it knows only how the label depends on the features. ## Consequences worth having ready - **Probabilities.** Both *can* be probabilistic. Outputting a number between 0 and 1 does not make a model generative - logistic regression does that and is discriminative. - **Labels.** Both need labels to be trained as classifiers. A generative classifier fits one `p(x|y)` per class, which requires knowing which examples belong to which class. - **Sample size.** Generative fits stabilise quickly because their parameters are simple per-class summaries; discriminative fits usually have lower error once enough data arrives. - **New classes.** Adding a class to a generative classifier means fitting one more class-conditional and renormalising the priors; a discriminative model is typically refitted end to end. ## Confusions that cost candidates points *Linear discriminant analysis is discriminative.* It is not. Despite the name, it assumes a normal density per class with a shared covariance matrix, fits those densities plus class priors, and applies the inversion above. The resulting boundary happens to be linear - that is what "linear" in the name refers to - but the object that was fitted is a class-conditional density, which makes the method generative. *Generative means unsupervised.* No. Density estimation without labels is unsupervised; a generative *classifier* is supervised and uses labels to split the data before fitting. *Discriminative means no probabilities.* No. It means the model targets the label given the features rather than the joint distribution. ## Where the line blurs The two routes can arrive at the same *shape* of decision rule. Normal class-conditionals with a shared covariance matrix imply that the log-odds between two classes is a linear function of `x` - exactly the functional form logistic regression fits. The families are still different, because they optimise different objectives from the same data and therefore land on different coefficients. The taxonomy is about what is estimated, not about how curved the boundary is.

  • Is a decision tree generative or discriminative?
    Discriminative. A tree recursively partitions the feature space to separate the labels and stores a class distribution in each leaf, which is an estimate of `p(y|x)` for that region. It never models how feature vectors are distributed within a class, so it cannot produce a new feature vector or say how typical an input is.
  • Why is linear discriminant analysis generative when the name suggests otherwise?
    Because of what it fits. It estimates a class prior and a normal density per class with a shared covariance matrix, then applies the inversion `p(y|x)` proportional to `p(y)p(x|y)`. The shared covariance makes the resulting boundary linear, which is where the name comes from, but the fitted object is a class-conditional density - the defining mark of a generative model.
  • Does a generative classifier need labelled data?
    Yes, to be a classifier. Each class-conditional `p(x|y)` is fitted from the examples carrying that label, so labels are required to split the data. What the generative form adds is that unlabelled examples are not useless: they carry information about the feature distribution and can be folded into the fit, which is the usual entry point for semi-supervised methods.

A discriminative model learns the few giveaway sounds that separate two languages. A generative model learns to speak each language, then asks which speaker was more likely to have produced the sentence it just heard.

saying these in an interview costs you the question

  • Says generative means the model writes text or paints images
  • Calls logistic regression generative because it outputs probabilities
  • Believes generative classifiers are unsupervised and need no labels
  • Claims linear discriminant analysis is discriminative because of its name
  • Thinks discriminative models are more accurate in every situation

context