What does the Fréchet Inception Distance measure between real and generated images?
answer
- compares sets, never image pairs
- features from a frozen pretrained classifier
- only means and covariances survive
- one Gaussian fitted per set
- lower is better; zero is weak proof
basics
~20 sFID embeds real and generated images with a fixed pretrained classifier, fits one Gaussian to each set of feature vectors, and reports the Fréchet distance between those two Gaussians. Lower means the feature distributions are closer.
solid answer
~50 sFID is a distributional score, not a per-image one. You push a large set of real images and a large set of generated images through a fixed pretrained image classifier and keep the pooled feature vector of each image. You then summarise each set by a mean vector and a covariance matrix — a single multivariate Gaussian per set — and compute the Fréchet distance between them: `||mu_r - mu_g||^2 + Tr(S_r + S_g - 2*(S_r*S_g)^(1/2))`. The first term punishes a shift in average feature content, the second a mismatch in feature spread and correlation. Lower is better, and zero only means the two fitted Gaussians coincide, which is much weaker than the two distributions being equal. It is popular because it responds to both degraded samples and missing variety, but it is one scalar and it inherits every blind spot of the borrowed feature space.
go deeper
Recall the shape of it: real and generated images become feature vectors from a pretrained classifier, the two sets are compared as distributions, and lower is better.
Be ready to write the two terms and say what each punishes — a shift in mean features versus a mismatch in covariance — and to name the single-Gaussian assumption as the reason a low score is not proof.
Show the operating discipline: a frozen protocol, enough samples, several seeds, and the real-versus-real floor before you call a small delta an improvement.
Own the framing that a borrowed feature space encodes someone else's notion of similarity, and decide when a distributional proxy is enough to steer a roadmap versus when downstream task utility has to be the decision metric.
## The problem FID is trying to solve An implicit generator — one you can sample from but whose density you cannot evaluate — gives you a bag of samples and nothing else. There is no held-out label to score against and no per-sample notion of correctness: a generated face is not "wrong", it is either plausible or not, and the set as a whole is either as varied as the real data or not. So the evaluation question has to be reframed from "is this sample right?" to "does the distribution of generated samples look like the distribution of real samples?" Comparing distributions directly in pixel space is hopeless. Two images of the same scene shifted by one pixel are far apart in pixel distance while being perceptually identical, and pixel statistics are dominated by low-level structure nobody cares about. FID's move is to compare the distributions somewhere more semantic. ## The recipe 1. Take N real images and N generated images. The conventional N is large — tens of thousands — for reasons covered below. 2. Push every image through a fixed, pretrained image classifier (conventionally an Inception network trained on ImageNet) and keep the activations of the pooling layer just before the classification head. Each image becomes a feature vector of a couple of thousand dimensions. These features encode object-ish, texture-ish content rather than raw pixels. 3. Summarise each of the two feature clouds by its sample mean vector `mu` and its sample covariance matrix `S`. This is the same as fitting one multivariate Gaussian to each cloud by moment matching. 4. Report the Fréchet distance (also called the 2-Wasserstein distance) between those two Gaussians, which has a closed form: `FID = ||mu_r - mu_g||^2 + Tr(S_r + S_g - 2*(S_r*S_g)^(1/2))` where the subscripts are real and generated, `Tr` is the matrix trace, and the square root is the matrix square root, not an elementwise one. ## Reading the two terms The squared-norm term measures how far apart the *average* feature vectors are. A generator whose samples systematically have the wrong colour balance, the wrong object mix, or a global artefact shifts the mean and this term grows. The trace term measures disagreement in the second moment: how spread out each cloud is and how its dimensions co-vary. A generator that produces beautiful but narrow output — say it has quietly stopped covering whole regions of the real data — has a feature covariance much smaller than the real one, and the trace term grows even if the means happen to line up. This is why FID reacts to lost variety while a metric that only checked realism per sample would not. ## What the number does and does not license Lower is better. FID is non-negative, and it is zero exactly when the two fitted Gaussians coincide. That last sentence is the whole caveat. FID only ever looks at the first two moments of the feature distribution. Two genuinely different distributions with matching means and covariances in that feature space score zero against each other. Nothing about the third moment, multimodality, or per-sample plausibility enters the formula. The Gaussian fit is a modelling assumption about deep feature activations that is convenient rather than true. A second limit: the score has no absolute meaning. "FID 14" is not a quality grade — it is a number relative to a particular real reference set, sample count, feature network and preprocessing pipeline. Useful FID work reports deltas within one frozen protocol, and often reports the floor obtained by scoring one half of the real data against the other half, which is greater than zero and tells you how much of your number is estimation noise. A third limit worth stating out loud: FID never inspects individual samples. It cannot tell you *which* images are bad, and it will happily reward a generator that reproduces its training data, because a copy of the training set matches the real distribution by construction. ## Practical framing for an interview Say what it compares (feature distributions, not images), what it assumes (one Gaussian per set, so only means and covariances), which direction is good (lower), and one honest limitation. Candidates who describe FID as "comparing each generated image to the closest real image" or as a pixel-space distance have not understood the object being measured, and that mistake propagates into every downstream decision about how many samples to score and what a difference of two points means.
- Why fit a Gaussian at all rather than compare the two feature clouds directly?Because the Gaussian gives a closed form. The Fréchet distance between two arbitrary high-dimensional distributions has no cheap estimator, but between two Gaussians it reduces to a mean difference plus a covariance trace term computable from sample moments. The cost is that everything beyond the first two moments is discarded, which is precisely the assumption that makes a near-zero score weak evidence.
- What does FID see when a generator produces sharp but nearly identical samples?The mean term can stay small because the average feature content is still plausible, but the generated covariance collapses relative to the real one, so the trace term blows up and FID rises. That sensitivity to shrunken spread is the main reason FID displaced scores that only judged per-sample realism.
- Does a two-point FID improvement mean the model is better?Not on its own. FID is an estimate with variance and a sample-size-dependent bias, so a small gap can be noise or a protocol difference. Re-score with the same protocol, several sampling seeds, and compare against the real-versus-real floor obtained by splitting the real data in half before calling a small delta a win.
It is like judging two orchestras by the average and spread of their frequency spectra rather than matching them note for note — it catches a whole missing section, but two very different pieces can share the same statistics.
saying these in an interview costs you the question
- Says FID matches each generated image to a real one
- Thinks a higher FID means better samples
- Claims FID near zero proves the distributions are identical
- Believes FID is computed on raw pixel values
- Describes FID as measuring only image sharpness