Why can't you compare your FID number against the one reported in a paper?
answer
- an estimator with knobs, not a constant
- sample count biases it in one direction
- different features, different geometry
- resize and re-encode change activations
- score the real half against the other half
basics
~20 sFID is only meaningful inside one fixed protocol. The number moves with sample count, the feature network's weights, the resizing applied before it, and which real split you score against. Change any of them and the comparison is void.
solid answer
~40 sFID is an estimator with knobs, and the knobs shift it by more than most reported model improvements. Sample count matters most: the estimator is biased upward at small counts, so scoring 5,000 generated samples gives a systematically worse number than scoring 50,000 from the very same model — bias with a direction, not noise. The feature network matters: different extractors, or the same architecture with different weights, define a different geometry and therefore a different distance. Preprocessing matters: the resize filter, antialiasing, the crop and any lossy re-encoding all change the activations. So does the real reference split. Compare only within a protocol you freeze and publish, across several sampling seeds, and against a real-versus-real floor from two disjoint halves of the real data.
go deeper
Remember that FID has no absolute meaning: it is a distance to a particular real set measured with a particular pipeline, so only numbers produced the same way can be compared.
Be able to list the knobs — sample count, feature network, preprocessing, real reference split — and state that the sample-count effect is a directional bias, not symmetric noise.
Demonstrate the protocol you would impose: frozen pipeline, baselines re-scored in house, multiple seeds, and a real-versus-real floor below which a gap is not a result.
Own the evaluation contract for the org: one published protocol everyone reports against, an explicit rule for when a delta counts, and a policy that external leaderboard numbers never enter a model-selection decision unreproduced.
## Why the question comes up Someone reports that their generator reaches a better score than a published one, and the reviewer's first job is to ask whether the two numbers were produced by the same measuring instrument. Very often they were not, and the difference in protocol is larger than the difference in models. Understanding this is the difference between using FID as an instrument and using it as a scoreboard. ## Knob one: how many samples you score FID is computed from sample estimates of a mean and a covariance in a space of a couple of thousand dimensions. A covariance in that many dimensions needs a great many samples to estimate well, and the error does not average out symmetrically: the trace term is systematically inflated when the covariance estimates are noisy. The consequence is a **bias with a direction** — fewer samples give a *higher*, i.e. worse, FID, on the same model with no change to it whatsoever. The practical size of this effect is large enough to invert model rankings. Scoring one model with 5,000 generated samples and another with 50,000 hands the second model an advantage it did not earn. It also means FID cannot be treated as an unbiased estimate of a population quantity that you can compute cheaply during training and compare against a paper's headline number. What to do: fix the count, use the largest you can afford, use the same count for every model in a comparison, and if you must use a small count during training, treat it as a within-run trend line and never as a cross-paper figure. ## Knob two: the feature space FID is a distance in whatever feature space you borrowed. Two different extractors define two different geometries, so the numbers are not on the same scale and not even guaranteed to rank models the same way. Even the same architecture with different weights — a different training run, a different training set, a different input resolution — is a different feature space. There is no conversion factor between them. ## Knob three: preprocessing Everything between an image file and the network input is part of the metric. Resizing to the network's expected resolution can use different interpolation filters and may or may not antialias; the difference shows up in the high-frequency content, which deep features are sensitive to. Cropping choices, value ranges, and lossy re-encoding of generated images before scoring all move the number. Two people using the same extractor and the same sample count can still disagree because one saved samples as a compressed format and the other did not. ## Knob four: the real reference set FID is a distance to *a specific set of real images*. Scoring against the training split and scoring against a held-out split give different numbers, and so does changing how many real images are in the reference. A paper that scores against the full training set and a project that scores against a small validation slice are not measuring the same thing. ## The discipline that makes the number usable - **Freeze and publish the protocol**: extractor and weights, preprocessing chain, generated sample count, real reference set and its size. - **Re-score baselines yourself** under your protocol rather than copying numbers out of a table. A relative comparison you produced is worth more than an absolute number you inherited. - **Report variance**: several sampling seeds, since generated samples are random draws. - **Report the floor**: split the real data into two disjoint halves and score one against the other under the identical protocol. This real-versus-real value is not zero, and it tells you the resolution of your instrument. A model-to-model gap smaller than that floor is not a result. - **Prefer deltas to levels**: "this change moved FID from 21.4 to 18.9 under our protocol, floor 3.1, across three seeds" is a claim; "we achieved FID 18.9" is a number without an instrument attached. ## The senior point behind all of this FID has no absolute zero and no units anyone agrees on. It is a comparator, and a comparator is only valid across measurements taken with the same instrument. Treating a leaderboard column as if it were a physical constant is the failure mode, and it produces model selections that do not survive contact with a re-measurement.
- Which direction does the sample-count bias push FID, and why?Fewer samples give a higher, worse FID. The mean and especially the covariance are estimated from a finite draw in a high-dimensional feature space, and the noise in those estimates inflates the trace term rather than cancelling out. The bias is systematic, so any comparison across different sample counts favours whoever scored more samples.
- What does a real-versus-real FID floor tell you?Split the real data into two disjoint halves and score them against each other under your exact protocol. The result is greater than zero and represents pure estimation error plus split variation. It is the resolution of your instrument: any model-to-model gap of comparable size is not evidence of a better model.
- Is it ever acceptable to track FID on a few thousand samples?Yes, as an in-run trend. A cheap, small-sample FID computed identically at every checkpoint tells you whether training is moving in the right direction. It is not comparable to anyone else's number, and it should not be the figure that selects a final model — recompute the finalists at full sample count under the frozen protocol.
saying these in an interview costs you the question
- Treats FID as an absolute quality grade with fixed meaning
- Says sample count only adds symmetric random noise
- Copies baseline FID numbers out of a paper's table
- Thinks resizing and re-encoding cannot affect the score
- Declares a sub-point FID gap a real improvement