What can a model of the joint p(x, y) do that a model of p(y|x) alone cannot?
answer
- what the model knows about x itself
- you can draw new records from it
- an empty slot can be integrated over
- swap p(y) and leave p(x|y) alone
basics
~20 sModelling the joint gives you a distribution over the features themselves, so you can draw synthetic records, integrate out a feature missing at scoring time, and judge how unusual an input is. A p(y|x) model answers only the labelling question.
solid answer
~50 sA joint model contains `p(x|y)`, a distribution over the features, and that buys capabilities a `p(y|x)` model has no machinery for. You can **sample** from it: fit the joint distribution of credit applicants and you can draw plausible synthetic applicants to stress-test a downstream policy. You can **marginalise**: if income is missing at scoring time, integrate it out of the class-conditional and score on what was actually observed, instead of inventing a value to fill the slot. You can evaluate `p(x)` and flag an input unlike anything in training. And because the score factorises into prior times class-conditional, a change in class balance is absorbed by replacing `p(y)` and renormalising, leaving the fitted densities untouched. The cost is real: density estimation in many dimensions is hard, and every wrong assumption about `p(x)` shows up in the decisions.
go deeper
Remember the one-line version: a joint model knows what the data looks like, so it can produce new examples; a p(y|x) model only labels examples it is handed.
Explain how the label score is recovered from prior times class-conditional, and why that same factorisation is what lets you integrate out an unobserved feature or swap in a different class prior.
Argue the tradeoff on a real system: which of sampling, marginalising, novelty scoring and prior swapping a requirement actually asks for, and whether they justify the assumptions and accuracy you give up.
Decide whether the organisation should carry a density model at all, or meet those needs with separate tooling - imputation, drift monitors, a synthetic-data process - and be explicit about who maintains the extra assumptions.
## What the joint actually contains A discriminative model is a function from features to a label score. Feed it a complete feature vector and it answers; that is the whole of its competence. A generative model is a description of the data-generating process: a class prior `p(y)` and, for each class, a distribution `p(x|y)` over feature vectors. Labelling is one query you can put to that description. It is not the only one, and the extra queries are the reason to consider paying for a joint model at all. ## Sampling Because `p(x|y)` is a distribution over feature vectors, you can draw from it. Fit the joint distribution of credit applicants and you can manufacture plausible synthetic applicants: draw a class, then draw an applicant profile from that class's distribution. That is useful for stress-testing a downstream decision policy against volumes of realistic-but-not-real records, for exercising a pipeline before real data exists, and for sharing a dataset's shape without sharing its rows. No `p(y|x)` model can do any of this - it has no notion of what a plausible applicant looks like, only of how a given applicant should be labelled. The usual caveat applies: samples are only as good as the density. If the fitted `p(x|y)` misses a dependency between features, the synthetic records will miss it too, and any conclusion that depends on that dependency is worthless. ## Marginalising a missing feature This is the sharpest interview version of the distinction. At scoring time a feature is unavailable - the income field did not come back from the upstream service. A discriminative model needs a number in that slot; the architecture leaves it no choice. So you impute: substitute a mean, or predict the missing value from the other features. Notice what imputation is - a model of how features relate to each other, bolted on from the outside. You have re-introduced a piece of `p(x)` through the back door, and worse, you have committed to a single value where you actually have uncertainty. A generative model integrates the unobserved part out of the class-conditional and scores against what was genuinely observed: ``` p(y | x_observed) is proportional to p(y) * integral over x_missing of p(x_observed, x_missing | y) ``` No value is invented. The uncertainty about the missing field is carried through into the label score, which is the statistically honest answer rather than a patched one. For some class-conditional forms the integral has a closed form and the computation is trivial; for others it needs numerical work. Either way the *capability* exists, which is the point being tested. ## Knowing when an input is strange Summing the joint over labels gives `p(x) = sum over y of p(y)p(x|y)` - the model's opinion on how typical a feature vector is, independent of any label. A very low value says this input is unlike anything the model was fitted on, so its label score should not be trusted. A discriminative model produces a confident-looking score for absurd inputs with no way to tell you it is extrapolating. Getting a novelty signal out of the same object you use for classification, rather than standing up a separate detector, is a genuine operational advantage. ## Absorbing a change in class balance Deployment class balance drifts - fraud rises, an offer's take-up doubles - while the way features are distributed within each class stays put. Because the score factorises as prior times class-conditional, that situation has an exact fix: replace `p(y)` with the new balance, renormalise, keep the fitted `p(x|y)` untouched. Related, adding an entirely new class means fitting one more class-conditional and updating the priors, with the existing ones left as they are, rather than refitting a model end to end. ## The bill None of this is free. - **Density estimation in many dimensions is hard.** Describing a whole cloud of points is a far more ambitious task than describing where its edge with another cloud lies, and it needs more assumptions. - **Wrong assumptions land in the decisions.** Capacity spent describing feature structure irrelevant to the boundary is capacity not spent on the boundary, and misspecification shows up as asymptotically worse classification. - **You may not need any of it.** If every scoring call carries a complete feature vector, the class balance is stable, and nobody wants synthetic data, the joint model is a cost with no matching benefit. ## How to answer when asked Name the capability, not the taxonomy. "A joint model lets me score a record with a missing field by integrating that field out, and lets me sample synthetic records for a stress test; a `p(y|x)` model has to be handed a complete row and can only label it." Then price it: those capabilities cost you assumptions about `p(x)` and usually some accuracy, so take them only when a real requirement is asking for them.
- How do you handle a feature missing at scoring time if you already have a discriminative model?Impute it, or encode absence explicitly. Imputation means predicting the missing field from the other features, or substituting a summary value; the explicit route adds a missing-indicator feature so the model learns what absence means, which needs the same pattern present in training. Note what imputation really is - a model of how features relate to one another, added from outside, that commits to one value where you actually have uncertainty.
- What is the cost of modelling p(x) when the label is all you need?Assumptions and capacity spent on structure the decision never uses. Describing each class's whole feature distribution is a much harder estimation problem than locating the boundary between classes, especially with many features, and any part of that description that is wrong propagates into the label scores. If nothing downstream asks for sampling, marginalising or novelty scores, that spend buys nothing.
- Can a generative classifier absorb a new class without refitting everything?Largely, yes. Each class-conditional is fitted from that class's examples alone, so adding a class means fitting one new `p(x|y)` and updating the priors to include it; the existing class-conditionals are unchanged. A discriminative model has no such separation - its parameters are fitted jointly against all labels at once, so a new class normally means refitting the whole thing.
saying these in an interview costs you the question
- Thinks imputing a mean is the same as marginalising a feature out
- Says a discriminative model can also sample new feature vectors
- Claims generative models are always the safer production choice
- Cannot name any use for p(x) beyond classification
- Ignores that density estimation in many dimensions is hard