How can a generator that memorises its training images still score an excellent FID?
answer
- distribution matching is not novelty
- a copied training set matches perfectly
- retrieve nearest neighbours in feature space
- compare against held-out real distances
- deduplicate, then inspect the worst tail
basics
~20 sBecause distributional scores only ask whether the generated distribution matches the real one, and a copy of the training set matches it exactly. Catching memorisation needs a separate nearest-neighbour audit against the training data, calibrated against a held-out baseline.
solid answer
~50 sEvery feature-space fidelity score is a comparison of two distributions. A generator that reproduces training examples produces a distribution identical to the training distribution, so it scores near the floor by construction. The metric is not fooled; it is answering the question it was asked, and that question does not include "did you invent this?". The audit is separate: embed generated and training samples in a feature space, retrieve each generated sample's nearest training neighbours, and look at the distribution of those distances. The critical part is calibration — compare against the distances from *held-out real* images to their nearest training neighbours. Real data has near-duplicates too, so a small raw distance proves nothing; what indicts a model is generated samples being systematically closer than a genuine unseen image is. Deduplicate first, and inspect the worst tail by hand.
go deeper
Remember that a good fidelity score says the generated samples resemble the real data as a set; it never says they were invented rather than copied.
Be able to explain why copying drives distributional scores down rather than up, and name nearest-neighbour retrieval against the training set as the separate check.
Show the working audit: feature-space retrieval, a held-out-real calibration baseline, deduplication beforehand, invariance to the augmentations used, and manual inspection of the worst tail.
Own the release policy: which corpora make memorisation a blocker rather than a curiosity, what the gate threshold is, and who signs off when the distance distribution sits close to the baseline.
## Why the metric cannot see it Start from what a distributional fidelity score is asking: *how far is the distribution of generated samples from the distribution of real samples, as measured in some feature space?* Now consider the degenerate generator that stores the training set and samples uniformly from it. Its output distribution **is** the training distribution. The distance to the real reference is as small as the estimator can report. This is not a bug in FID, and no amount of increasing the sample count fixes it. Matching the distribution and producing novel content are different properties, and only the first is being measured. The same holds for precision and recall for generative models: copied samples sit exactly on the real manifold, so precision is maximal, and a model that reproduces the whole training set also reaches all of the real support, so recall is maximal too. Partial memorisation is the realistic version. Models trained on small datasets, or on datasets with many duplicated images, reproduce some fraction of their training data closely while generating the rest. The distributional score improves as the copying increases, so the metric actively rewards the behaviour you want to prevent. ## Why it matters beyond metric hygiene A generator that emits near-copies of its training data carries the licensing, privacy and confidentiality properties of that data into every output. For medical, personal or licensed corpora that turns a modelling curiosity into a release blocker, which is why the audit belongs in the evaluation suite and not in a research appendix. ## The audit **Retrieve nearest neighbours in a feature space.** For each generated sample, find its closest training examples using embeddings rather than raw pixel distance. Pixel distance is defeated by a one-pixel shift, a small crop, a horizontal flip or a brightness change, all of which leave a copy perceptually identical while moving it far away in pixel space. **Calibrate against held-out real data.** This is the step people skip and the reason naive audits produce nonsense in both directions. Real datasets contain genuinely similar images, so some small nearest-neighbour distances are normal. Build the reference distribution by taking held-out real images — images the model never saw — and measuring *their* distances to their nearest training neighbours. Then compare the generated set's distance distribution against that reference. If generated samples are systematically closer to training images than a genuine unseen real image is, the model is copying. If the two distributions overlap, the retrieved "matches" are just the ordinary similarity present in the domain. **Deduplicate first.** A training set with many near-duplicates makes the model far likelier to reproduce those images and simultaneously makes the baseline distances tighter, so both sides of the comparison are distorted. Deduplicating in feature space before training and before the audit removes that confound. **Look at the tail with your own eyes.** Aggregate statistics tell you whether a problem exists; the worst hundred retrieved pairs tell you what it is. Side-by-side inspection of the most-suspicious generated samples and their nearest training neighbours is the step that turns a distance histogram into a decision. **Check the invariances that matter.** Retrieve under flips and crops if the training pipeline used them; otherwise an augmented copy passes the audit untouched. ## Reporting it A usable memorisation report is not a single number. It is: the distance distribution for generated samples overlaid on the held-out-real baseline, the fraction of generated samples below the baseline's low percentile, and the inspected worst cases. That gives a reviewer something they can argue with, and it gives you a threshold you can gate a release on. ## The framing that scores well in interviews The headline is that distribution matching and novelty are different questions, so no fidelity metric — however good — can answer both. Anyone who proposes "a lower FID" as the response to a memorisation concern has missed that the metric moves the wrong way. The correct answer names the separate audit, and the calibrated baseline is the detail that separates candidates who have actually run one from those who have read about it.
- Why not run the nearest-neighbour search in pixel space?Because pixel distance is not invariant to the transformations that preserve identity. A one-pixel shift, a small crop, a horizontal flip or a brightness change leaves an image perceptually identical while moving it far away in pixel space, so a copy sails through the audit. Feature-space retrieval matches on content and catches those cases.
- What does the held-out-real baseline protect you against?Both false alarms and false comfort. Real data contains genuinely similar images, so small nearest-neighbour distances are normal and a raw threshold flags innocent samples. The baseline says what distance an unseen real image achieves; only generated samples systematically closer than that indicate copying, and the comparison also exposes mild copying that a fixed threshold would miss.
- Do precision and recall for generative models catch memorisation instead?No, they reward it. Copied samples lie exactly on the real manifold, so precision is maximal, and a model reproducing the full training set also covers the real support, so recall is maximal. Neither quantity looks at distance to the training data, so memorisation remains invisible to both.
saying these in an interview costs you the question
- Says a low FID proves the samples are novel
- Proposes fixing memorisation by scoring more samples
- Runs the nearest-neighbour audit in raw pixel space
- Uses a fixed distance threshold with no baseline
- Ignores duplicated training images before auditing