How do you evaluate a generator of sensor time series when no standard feature extractor exists?
answer
- the feature space is a borrowed instrument
- no default extractor outside images
- train your own encoder, freeze it
- train on synthetic, test on real
- spectra and physical bounds are checkable
basics
~20 sYou build the yardstick yourself: train a domain encoder on real data and measure distribution distance in its features, back that with train-on-synthetic-test-on-real utility and physically meaningful summary statistics, and accept that the numbers are comparable only inside your own project.
solid answer
~40 sEvery feature-space fidelity score is a borrowed instrument, and outside image domains there is nothing to borrow. Three substitutes, used together. First, train your own encoder on real data — a supervised classifier for a domain task, or a self-supervised model — and compute a distribution distance in its embedding space; the recipe survives, but the score now depends on an encoder you own and is comparable only within your project. Second, evaluate by utility: train a downstream model on synthetic data, test it on real, and compare against training on real. Third, compare interpretable summary statistics — spectra, autocorrelation, event rates, cross-channel correlations, physical bounds — because domain experts can argue with those and a learned score is opaque. Report all three and treat any absolute number as internal.
go deeper
Remember that a fidelity score is a distance measured inside some pretrained network's feature space, so outside the domain that network was trained on it may not mean anything.
Be ready to name the substitutes — a domain encoder you train yourself, downstream utility, interpretable summary statistics — and say what each one measures.
Show the operating detail: encoder trained on a disjoint real split, frozen and versioned, utility measured in both directions, hard physical constraints checked, memorisation audited.
Own the evaluation contract before modelling starts: which measurement gates a release, which are diagnostics, what delta counts, and who maintains the instrument once it is a piece of infrastructure.
## The hidden dependency in every fidelity score A feature-space score answers "how different are these two distributions?" *in the geometry of some pretrained network*. In natural images that network is a shared convention, so numbers travel between teams. Outside that setting the convention does not exist, and the borrowed instrument problem becomes visible. It is worth seeing the intermediate case first, because it is the same failure in a milder form: scoring generated chest X-rays with a feature network trained to classify everyday object photographs. A network does exist, so a number comes out — but its features were optimised to separate categories of ordinary objects, and nothing forced it to encode the fine structure that determines whether a radiographic image is plausible. The score can be stable, reproducible, and still be measuring the wrong similarity. Rankings from it need not agree with anything domain-relevant. A number that exists is not the same as a number that means something. For sensor time series, audio, tabular records or event logs, there is no default at all, and the recipe has to be rebuilt. ## Substitute one: an encoder you train Keep the structure of the metric and replace the borrowed network with one trained on your own real data — a supervised model for a task that matters in the domain, or a self-supervised encoder. Then compute a distribution distance between real and generated embeddings. What you gain: a semantic space adapted to your domain instead of an irrelevant one. What you pay: - **No external comparability.** Nobody else can reproduce or interpret your number, and it is meaningful only as a relative comparison inside your project. - **A moving instrument.** Retraining the encoder changes the metric, so it must be frozen and versioned like a dataset, with any change forcing every historical number to be recomputed. - **Blind spots inherited from the encoder's objective.** If it was trained to classify activity type, it may be insensitive to exactly the artefacts you care about — an unphysical spike or a broken sampling interval — because those never mattered for its task. - **Circularity risk.** An encoder trained on the same data the generator saw can flatter the generator; train it on a disjoint real split. ## Substitute two: utility Train a downstream model on the synthetic data, test it on a real held-out set, and compare with the same model trained on real data of the same size. This asks whether the samples carry the structure that makes the data useful. It is the strongest signal when the reason you generated data has a downstream task attached — augmentation, rare-event enrichment, sharing a stand-in for restricted data. It is also honest: it cannot be gamed by matching superficial statistics. Its limits are that it is expensive, noisy across seeds, and only measures the properties the downstream task happens to depend on. Run the reverse direction too — train on real, test on synthetic — since a large asymmetry between the two exposes a generated distribution that is narrower or shifted. ## Substitute three: interpretable statistics Compare quantities the domain already reasons about: power spectra, autocorrelation at relevant lags, distributions of event durations and inter-arrival times, cross-channel correlation structure, physical bounds and conservation constraints, drift and seasonality. These have three advantages: they can be checked, they are cheap, and a violated hard constraint is an immediate, non-negotiable failure that no learned score would flag as such. They are necessary rather than sufficient — a generator can match every summary statistic you thought to check and still be wrong in a way you did not think to check — which is exactly why they sit alongside the other two rather than replacing them. ## Also audit memorisation Everything above measures distribution match, so a generator that reproduces training records scores well on all three. In restricted-data domains, which is often precisely why synthetic data was wanted, the nearest-neighbour audit against the training set is mandatory rather than optional. ## The decision this is really about The principal-level call is not which score to compute; it is what the evaluation *contract* is. Which measurement is the gate, which are diagnostics, what deltas count as real, and how the instrument is versioned. Committing to a self-trained encoder means committing to freeze and maintain it. Committing to utility means paying for downstream training runs on every candidate. Deciding this before the first model is trained is what stops the evaluation being chosen retroactively by whichever number made the latest checkpoint look best.
- What is the cost of scoring a domain with a feature network trained on unrelated images?You get a reproducible number that may not measure anything relevant. The features were optimised to separate categories of the network's own training domain, so nothing guarantees they encode the structure that decides plausibility in yours. Rankings from such a score can disagree with domain-relevant judgments, and its stability is easily mistaken for validity.
- How do you keep a self-trained encoder from flattering the generator?Train it on a real split the generator never saw, freeze and version it exactly like a dataset, and recompute every historical number if it ever changes. Without that discipline the metric drifts with the encoder, and an encoder trained on the generator's own training data can score memorised or overfit output favourably.
- Why run train-on-real, test-on-synthetic as well as the reverse?The two directions expose different faults. Training on synthetic and testing on real measures whether the samples carry usable structure. Training on real and testing on synthetic measures whether the synthetic set is unrepresentatively narrow or shifted: a model that performs far better on synthetic than on real data is telling you the generated distribution is easier and smaller than the truth.
It is like grading wine with a thermometer because that is the instrument on the shelf; outside its intended domain you must build a scale, and then only your own tastings can be compared against each other.
saying these in an interview costs you the question
- Applies an object-photo feature network to any domain uncritically
- Treats a self-trained encoder's score as externally comparable
- Retrains the encoder between runs and compares old numbers
- Relies on summary statistics alone as sufficient proof
- Skips the memorisation audit on restricted-data domains