When a flow beats a GAN on exact likelihood but its samples look worse, what do you conclude?
answer
- the two numbers optimise different things
- expectation taken over the data, not the model
- zero-avoiding, so the model hedges
- wasted mass is nearly free
- most bits live in local texture
basics
~10 sLikelihood and sample quality measure different things. Maximum likelihood minimises a mode-covering divergence that punishes missing data but not wasted mass, so a better likelihood is evidence of density fit, not of better-looking samples.
solid answer
~50 sNothing is broken — likelihood and perceptual quality are close to independent in high dimensions. Fitting by maximum likelihood minimises the forward KL from data to model, which is mode-covering: assigning near-zero density to a real example is catastrophic, while spending mass on regions containing no real data is barely penalised. So the exact-likelihood model hedges, covers everything, and produces blurry or implausible samples. An implicit generator trained adversarially has no density to report but is pushed toward whatever the discriminator cannot distinguish, which rewards sharpness and tolerates dropped modes. Bits-per-dimension is also dominated by local texture statistics, so a model can win bits without ever getting global structure right. The practical conclusion: pick the metric matching the deliverable — likelihood if you consume `log p(x)` for compression or anomaly scoring, human judgment if you ship samples — and never treat one as a proxy for the other.
go deeper
Be ready to say that likelihood measures how much density the model puts on real data, and that this is not the same as whether a generated sample looks convincing.
Explain the direction of the KL that maximum likelihood minimises and why an expectation taken over the data makes the model avoid zeros and hedge rather than commit.
Diagnose it as expected behaviour, not a training bug, and be careful about comparability: preprocessing and bit depth shift bits-per-dimension independently of model quality.
Own the family decision — decide first whether the deliverable is a density value or an image a person judges, and refuse to let the easily-computed number become the team's target.
## Two numbers, two objectives The scenario is real and well documented, and it is a favourite senior question because it separates candidates who understand *what an objective optimises* from those who treat every number as "model quality". **Bits-per-dimension** is the standard exact-likelihood report for image models: the model's negative log-likelihood expressed in base two and divided by the number of dimensions. Read literally, it is the average number of bits an optimal coder would need per dimension using the model's density. Lower is better. Autoregressive models and flows can report it exactly. Implicit generators such as GANs cannot report it at all — they define a sampler, not a density — which is itself part of why the comparison is awkward. **Sample quality** is a different quantity altogether: whether an individual draw looks like a plausible member of the data distribution. ## Why maximum likelihood does not chase sharpness Maximising `E_data[log p_model(x)]` is equivalent to minimising the forward Kullback-Leibler divergence `KL(p_data || p_model)`. Look at where the penalty lives. The expectation is over the *data*, so wherever real data exists and the model puts near-zero density, `log p_model` goes to minus infinity and the loss explodes. This is a **mode-covering**, zero-avoiding objective: the model is compelled to place density on every real example. Now look at the other error. If the model puts substantial density on a region containing no data at all, the data expectation never visits that region, so the term contributes essentially nothing. Wasted mass is nearly free. A finite-capacity model under this objective therefore hedges — it smears density broadly to make sure it covers everything — and smeared density produces samples that are averages of plausible things rather than any one plausible thing. In image terms, blur. The adversarial family sits at the opposite end. A GAN never evaluates a density; it is trained so a discriminator cannot separate its samples from real ones. Every generated sample is judged, so implausible samples are punished directly, which rewards sharpness. Nothing in the game punishes *failing to generate* a whole region of the data — hence mode dropping. Sharper samples, worse coverage, no reportable likelihood. ## Why likelihood is dominated by local structure The second half of the answer is dimensional. A bits-per-dimension score sums a contribution from every dimension of the data. In a high-resolution image, the overwhelming majority of those dimensions are locally predictable from their neighbours — smooth gradients, texture, sensor noise statistics. A model that nails those low-level correlations captures most of the available bits, whether or not it has any grasp of global composition. Conversely, getting global structure right moves comparatively few bits. The well-known consequence, argued formally in the generative-model evaluation literature: log-likelihood and sample quality can be made almost arbitrarily independent in high dimensions. A model can be constructed with excellent likelihood and terrible samples, and another with beautiful samples and appalling likelihood. Neither number is a proxy for the other, and a small likelihood improvement should never be reported as a visual improvement. ## What a principal-level answer does with this 1. **Refuse the single-number framing.** Ask what the model is for. If downstream consumers read `log p(x)` — lossless compression, likelihood-ratio anomaly scoring, model comparison, importance weighting — then likelihood *is* the deliverable and its exactness is the reason to pick this family. If the deliverable is images a person looks at, likelihood is a weak proxy and should not drive the decision. 2. **Compare like with like.** Bits-per-dimension is only comparable across models trained on identically preprocessed data at the same bit depth and with the same treatment of discretisation. Different preprocessing silently shifts the number, and cross-paper comparisons are frequently invalid for exactly this reason. 3. **Report both, and say which is the objective.** The failure mode to avoid organisationally is a team optimising the number that is easy to compute while shipping the property nobody measured. 4. **Recognise it is not a bug to fix.** The candidate who says "the flow must be undertrained" has missed the point. The gap is the expected behaviour of a mode-covering objective under finite capacity, not a training defect. ## Where the exactness genuinely pays It is worth being concrete about the upside so the answer is not one-sided. An exact likelihood gives you: a principled model-selection criterion that does not require a learned evaluator; density values usable as anomaly scores; a training signal with no adversarial instability, no minimax equilibrium and no bound gap; and a number you can hold constant while you change the architecture. Those are real, and they are why the family survives despite losing sample-quality comparisons. The conclusion to state out loud: a better likelihood is genuine evidence about density fit and nothing more. Judging generative families requires deciding first which property you are actually buying.
- Which divergence does maximum-likelihood training actually minimise, and why does the direction matter?It minimises `KL(p_data || p_model)`, the forward direction, whose expectation runs over the data. That makes it zero-avoiding: near-zero model density where data exists is unbounded loss, so the model covers everything. The reverse direction, `KL(p_model || p_data)`, takes the expectation over model samples and is mode-seeking — it would rather concentrate on one region and ignore the rest.
- When is an exact likelihood the deliverable rather than a diagnostic?Whenever something downstream consumes the density value: lossless compression, where code length is the likelihood; anomaly and out-of-distribution scoring by density or density ratios; importance weighting; and model comparison without a learned evaluator. In those uses sample quality is irrelevant and the family's sampling cost is never paid.
- Is bits-per-dimension comparable across two papers reporting on the same dataset?Often not. The number depends on bit depth, on how continuous density is reconciled with discrete pixel values, and on any preprocessing that rescales the data — each shifts the score by a constant that has nothing to do with model quality. Comparisons are safe only under an identical data pipeline, which is why teams reproduce baselines internally.
- A colleague proposes fixing the blurry samples by training the flow longer. What do you say?That it treats a property of the objective as a bug. Under finite capacity a mode-covering objective spends capacity on coverage, and more steps on the same objective moves further in the same direction. If sharper samples are the goal, the answer is a different objective or family, or an added perceptual criterion — not more of the same training.
A weather forecaster who always predicts a wide range is never caught out by the actual weather, and scores well on never-being-wrong, while sounding useless to anyone who wanted a single confident forecast.
saying these in an interview costs you the question
- Says the flow must simply be undertrained
- Treats bits-per-dimension as a perceptual quality score
- Claims a GAN also reports an exact likelihood
- Confuses mode-covering with mode-seeking behaviour
- Compares bits-per-dimension across different preprocessing pipelines