skip to content

How do precision and recall for generative models separate fidelity from coverage?

level: middleimportance: should knowfreq 36%

answer

  1. two numbers, two different failures
  2. generated-into-real versus real-into-generated
  3. manifolds from k-th neighbour radii
  4. clean and narrow versus broad and sloppy
  5. one outlier inflates a hypersphere

basics

~20 s

Precision is the share of generated samples falling inside the real data's feature manifold, which measures fidelity. Recall is the share of real samples falling inside the generated manifold, which measures coverage. Opposite failures cannot cancel.

solid answer

~50 s

A single scalar cannot tell you *why* a generator is bad. Precision and recall for generative models split the question in two. Both sets are embedded in a feature space; the real manifold is approximated by placing around each real feature vector a hypersphere whose radius is its distance to its k-th nearest real neighbour, and the generated manifold the same way. Precision is the fraction of generated samples landing inside the real manifold — how often the model produces something plausible. Recall is the fraction of real samples landing inside the generated manifold — how much of the real data it can produce at all. A cautious generator emitting a narrow band of clean samples scores high precision, low recall; a broad one that also emits junk scores the reverse. One scalar mixes those and can rank the two equal for opposite reasons.

go deeper

for a junior

Recall the split: one number says whether generated samples look real, the other says whether the model can produce everything the real data contains.

for a middle

Be ready to state both directions correctly and describe the neighbourhood-radius manifold estimate, including what the neighbour count k does to how permissive it is.

for a senior

Show that you read the pair diagnostically — which failure a given profile implies, and what a truncation knob does to the frontier — instead of quoting one operating point.

for a principal

Decide which side your product actually needs: a customer-facing generator may demand fidelity at the cost of coverage, and that choice should be written into the evaluation contract rather than left to whoever picks the checkpoint.

## Why one number is not enough A generative model can fail in two orthogonal ways. - **Fidelity failure**: it emits samples that do not look like real data at all — artefacts, incoherent structure, implausible combinations. - **Coverage failure**: every sample it emits is plausible, but it only ever produces a slice of what the real data contains. Whole regions of the real distribution are simply never generated. These call for different fixes and have different product consequences, and a single scalar collapses them. Two models can land on the same FID for opposite reasons: one broad and sloppy, one narrow and clean. If your evaluation cannot distinguish them, your next experiment is a guess. Inception Score makes this especially visible. It rewards each sample having a confident predicted label while the labels across the set are spread out. A generator that produces exactly one flawless canonical image per category, over and over, satisfies both conditions and scores excellently — the metric never asks how much variety exists *within* a category, and it never looks at the real data at all. ## The construction Embed both the real set and the generated set into a feature space with a fixed extractor, exactly as a distributional score would. To approximate the support of a set, take each of its feature vectors and draw a hypersphere around it whose radius is the distance to that point's k-th nearest neighbour within the same set. The union of those hyperspheres is the estimated manifold. A point counts as "inside" if it falls within any of them. The parameter k controls how permissive the estimate is: small k gives a tight, fragmented manifold, large k an over-smoothed one that swallows regions that contain no data. With the two manifolds in hand: - **Precision** = fraction of *generated* samples that fall inside the *real* manifold. High precision means the model rarely produces something outside the real data's support: fidelity. - **Recall** = fraction of *real* samples that fall inside the *generated* manifold. High recall means the model's support reaches most of the real data: coverage. The naming mirrors classification precision and recall, but nothing here involves labels or a decision threshold — the "positives" are set memberships in feature space. ## Reading the pair - **High precision, low recall**: clean but narrow. The generator has learned a plausible sub-region and stays in it. Truncating or otherwise restricting the sampling distribution moves a model along this axis deliberately, trading coverage for fidelity, and a precision-recall pair makes the trade explicit instead of hidden inside a scalar. - **Low precision, high recall**: broad but unreliable. The generator reaches everywhere, including places nothing real lives. - **Both low**: the model is not working. - **Both high**: the interesting case, and the one where you still have to check for memorisation, because copying the training data maximises both. Because you get a pair, you can also draw a curve by sweeping a sampling-truncation knob and compare models by their whole frontier rather than at one operating point. ## Known weaknesses The k-th-neighbour radius is fragile. A single real outlier sitting far from everything gets an enormous hypersphere, and that balloon inflates precision by declaring large empty regions to be part of the real manifold. **Density and coverage** were proposed to address exactly this: density counts, for each generated sample, how many real neighbourhood spheres contain it (so a sample in a genuinely dense region counts for more, and one lone outlier cannot single-handedly admit it), and coverage measures the fraction of *real* samples that have at least one generated sample inside their own neighbourhood sphere, which avoids building a manifold out of possibly wild generated points. Both families still depend entirely on the feature extractor: they measure fidelity and coverage *as seen by that network*. And both need enough samples for neighbour distances to be meaningful, with the same sample-count sensitivity that afflicts any high-dimensional estimate. ## What to say in an interview Name the two quantities in the right direction — generated-into-real is fidelity, real-into-generated is coverage — describe the neighbourhood-radius manifold estimate, give the two opposite failure profiles, and mention that the outlier-driven radius problem motivated the density and coverage variants. Getting the direction backwards is the mistake interviewers listen for.

  • Which of the two drops when a generator only ever produces a narrow slice of the data?
    Recall drops, because most real samples fall outside the generated manifold. Precision can stay high or even rise, since everything the model does emit sits comfortably inside the real support. That combination — high precision, low recall — is the signature of a model that has traded coverage for safety.
  • Why did density and coverage get proposed as replacements?
    Because the k-th-nearest-neighbour radius is dominated by outliers. One isolated real point gets a huge hypersphere that admits large empty regions into the real manifold, inflating precision. Density counts how many real neighbourhoods contain each generated sample instead of using a binary membership test, and coverage checks real points' own neighbourhoods rather than trusting a manifold built from generated samples.
  • What does a high precision and high recall pair still fail to rule out?
    Memorisation. A generator that re-emits its training images sits exactly on the real manifold and reaches all of it, so both numbers look excellent. Neither quantity inspects distance to the training data, so a separate nearest-neighbour audit against the training set is required before believing the result.

Precision asks what fraction of the postcards you painted depict places that actually exist; recall asks what fraction of the real places you ever painted at all.

saying these in an interview costs you the question

  • Reverses the two: calls real-into-generated fidelity
  • Thinks these are classification precision and recall with labels
  • Claims a single scalar can separate fidelity from coverage
  • Ignores that the manifold estimate depends on k
  • Believes high precision and recall rule out copying

context