skip to content

Sample Fidelity Metrics

Generators with no likelihood to report get scored by comparing feature statistics of real and fake samples, from a two-player game or a denoising chain alike. Interviewers ask what one number hides.

on this pageshow

questions

5

What does the Fréchet Inception Distance measure between real and generated images?

level: middleimportance: must knowfreq 66%

answer

  1. compares sets, never image pairs
  2. features from a frozen pretrained classifier
  3. only means and covariances survive
  4. one Gaussian fitted per set
  5. lower is better; zero is weak proof

basics

~20 s

FID embeds real and generated images with a fixed pretrained classifier, fits one Gaussian to each set of feature vectors, and reports the Fréchet distance between those two Gaussians. Lower means the feature distributions are closer.

solid answer

~50 s

FID is a distributional score, not a per-image one. You push a large set of real images and a large set of generated images through a fixed pretrained image classifier and keep the pooled feature vector of each image. You then summarise each set by a mean vector and a covariance matrix — a single multivariate Gaussian per set — and compute the Fréchet distance between them: `||mu_r - mu_g||^2 + Tr(S_r + S_g - 2*(S_r*S_g)^(1/2))`. The first term punishes a shift in average feature content, the second a mismatch in feature spread and correlation. Lower is better, and zero only means the two fitted Gaussians coincide, which is much weaker than the two distributions being equal. It is popular because it responds to both degraded samples and missing variety, but it is one scalar and it inherits every blind spot of the borrowed feature space.

go deeper

for a junior

Recall the shape of it: real and generated images become feature vectors from a pretrained classifier, the two sets are compared as distributions, and lower is better.

for a middle

Be ready to write the two terms and say what each punishes — a shift in mean features versus a mismatch in covariance — and to name the single-Gaussian assumption as the reason a low score is not proof.

for a senior

Show the operating discipline: a frozen protocol, enough samples, several seeds, and the real-versus-real floor before you call a small delta an improvement.

for a principal

Own the framing that a borrowed feature space encodes someone else's notion of similarity, and decide when a distributional proxy is enough to steer a roadmap versus when downstream task utility has to be the decision metric.

## The problem FID is trying to solve An implicit generator — one you can sample from but whose density you cannot evaluate — gives you a bag of samples and nothing else. There is no held-out label to score against and no per-sample notion of correctness: a generated face is not "wrong", it is either plausible or not, and the set as a whole is either as varied as the real data or not. So the evaluation question has to be reframed from "is this sample right?" to "does the distribution of generated samples look like the distribution of real samples?" Comparing distributions directly in pixel space is hopeless. Two images of the same scene shifted by one pixel are far apart in pixel distance while being perceptually identical, and pixel statistics are dominated by low-level structure nobody cares about. FID's move is to compare the distributions somewhere more semantic. ## The recipe 1. Take N real images and N generated images. The conventional N is large — tens of thousands — for reasons covered below. 2. Push every image through a fixed, pretrained image classifier (conventionally an Inception network trained on ImageNet) and keep the activations of the pooling layer just before the classification head. Each image becomes a feature vector of a couple of thousand dimensions. These features encode object-ish, texture-ish content rather than raw pixels. 3. Summarise each of the two feature clouds by its sample mean vector `mu` and its sample covariance matrix `S`. This is the same as fitting one multivariate Gaussian to each cloud by moment matching. 4. Report the Fréchet distance (also called the 2-Wasserstein distance) between those two Gaussians, which has a closed form: `FID = ||mu_r - mu_g||^2 + Tr(S_r + S_g - 2*(S_r*S_g)^(1/2))` where the subscripts are real and generated, `Tr` is the matrix trace, and the square root is the matrix square root, not an elementwise one. ## Reading the two terms The squared-norm term measures how far apart the *average* feature vectors are. A generator whose samples systematically have the wrong colour balance, the wrong object mix, or a global artefact shifts the mean and this term grows. The trace term measures disagreement in the second moment: how spread out each cloud is and how its dimensions co-vary. A generator that produces beautiful but narrow output — say it has quietly stopped covering whole regions of the real data — has a feature covariance much smaller than the real one, and the trace term grows even if the means happen to line up. This is why FID reacts to lost variety while a metric that only checked realism per sample would not. ## What the number does and does not license Lower is better. FID is non-negative, and it is zero exactly when the two fitted Gaussians coincide. That last sentence is the whole caveat. FID only ever looks at the first two moments of the feature distribution. Two genuinely different distributions with matching means and covariances in that feature space score zero against each other. Nothing about the third moment, multimodality, or per-sample plausibility enters the formula. The Gaussian fit is a modelling assumption about deep feature activations that is convenient rather than true. A second limit: the score has no absolute meaning. "FID 14" is not a quality grade — it is a number relative to a particular real reference set, sample count, feature network and preprocessing pipeline. Useful FID work reports deltas within one frozen protocol, and often reports the floor obtained by scoring one half of the real data against the other half, which is greater than zero and tells you how much of your number is estimation noise. A third limit worth stating out loud: FID never inspects individual samples. It cannot tell you *which* images are bad, and it will happily reward a generator that reproduces its training data, because a copy of the training set matches the real distribution by construction. ## Practical framing for an interview Say what it compares (feature distributions, not images), what it assumes (one Gaussian per set, so only means and covariances), which direction is good (lower), and one honest limitation. Candidates who describe FID as "comparing each generated image to the closest real image" or as a pixel-space distance have not understood the object being measured, and that mistake propagates into every downstream decision about how many samples to score and what a difference of two points means.

  • Why fit a Gaussian at all rather than compare the two feature clouds directly?
    Because the Gaussian gives a closed form. The Fréchet distance between two arbitrary high-dimensional distributions has no cheap estimator, but between two Gaussians it reduces to a mean difference plus a covariance trace term computable from sample moments. The cost is that everything beyond the first two moments is discarded, which is precisely the assumption that makes a near-zero score weak evidence.
  • What does FID see when a generator produces sharp but nearly identical samples?
    The mean term can stay small because the average feature content is still plausible, but the generated covariance collapses relative to the real one, so the trace term blows up and FID rises. That sensitivity to shrunken spread is the main reason FID displaced scores that only judged per-sample realism.
  • Does a two-point FID improvement mean the model is better?
    Not on its own. FID is an estimate with variance and a sample-size-dependent bias, so a small gap can be noise or a protocol difference. Re-score with the same protocol, several sampling seeds, and compare against the real-versus-real floor obtained by splitting the real data in half before calling a small delta a win.

It is like judging two orchestras by the average and spread of their frequency spectra rather than matching them note for note — it catches a whole missing section, but two very different pieces can share the same statistics.

saying these in an interview costs you the question

  • Says FID matches each generated image to a real one
  • Thinks a higher FID means better samples
  • Claims FID near zero proves the distributions are identical
  • Believes FID is computed on raw pixel values
  • Describes FID as measuring only image sharpness

context

open as a page

How do precision and recall for generative models separate fidelity from coverage?

level: middleimportance: should knowfreq 36%

basics

~20 s

Precision is the share of generated samples falling inside the real data's feature manifold, which measures fidelity. Recall is the share of real samples falling inside the generated manifold, which measures coverage. Opposite failures cannot cancel.

open as a page

Why can't you compare your FID number against the one reported in a paper?

level: seniorimportance: should knowfreq 44%

basics

~20 s

FID is only meaningful inside one fixed protocol. The number moves with sample count, the feature network's weights, the resizing applied before it, and which real split you score against. Change any of them and the comparison is void.

open as a page

How can a generator that memorises its training images still score an excellent FID?

level: seniorimportance: should knowfreq 28%

basics

~20 s

Because distributional scores only ask whether the generated distribution matches the real one, and a copy of the training set matches it exactly. Catching memorisation needs a separate nearest-neighbour audit against the training data, calibrated against a held-out baseline.

open as a page

How do you evaluate a generator of sensor time series when no standard feature extractor exists?

level: principalimportance: nice to knowfreq 18%

basics

~20 s

You build the yardstick yourself: train a domain encoder on real data and measure distribution distance in its features, back that with train-on-synthetic-test-on-real utility and physically meaningful summary statistics, and accept that the numbers are comparable only inside your own project.

open as a page