skip to content

Should 2-D t-SNE or UMAP coordinates be used as features for a downstream model?

level: seniorimportance: nice to knowfreq 26%

answer

  1. an instrument for looking, not a representation
  2. no reusable mapping for future rows
  3. coordinates depend on the other rows
  4. fit inside the fold or leak
  5. two dimensions is a severe bottleneck

basics

~20 s

For t-SNE, no: it produces no reusable mapping, its layout changes run to run, and two coordinates discard almost everything. A frozen UMAP embedding fit only on training folds can be defensible; treat it as a modelling choice.

solid answer

~50 s

Two dimensions from a neighbour-preserving layout are a viewing device, not a representation. With t-SNE it is close to indefensible: there is no function to apply to future rows, a fresh run gives a different layout, and the coordinates depend on which other rows were in the batch, so the feature is not well defined for a single record. UMAP differs in kind because it can embed new rows into a fitted embedding, and a moderate-dimensional UMAP representation — ten or twenty components, not two — occasionally helps distance-based models. Even then it must be fit inside each cross-validation fold: fitting the projection on all the data before splitting leaks held-out information into the training features and inflates the score. If labels were used to guide the projection, that leak is severe. The default answer is to keep the original features.

go deeper

for a junior

Know the default: these maps are for looking at data. Do not turn the two coordinates into model inputs, especially not from a method that cannot place a new row at all.

for a middle

Explain the reasons — no reusable mapping for t-SNE, coordinates that depend on the batch and the seed, and a two-dimensional bottleneck that discards most of the signal.

for a senior

Show the leakage discipline. Any fitted projection belongs inside the cross-validation fold, and any gain must be demonstrated against the original features with equal tuning effort.

for a principal

Own the standard: no unsupervised fitted transform enters a pipeline without a fold-safe fit, a versioned artefact, a refresh policy and a demonstrated win over the plain-features baseline.

## The question behind the question A candidate who produces a beautiful 2-D map with clean islands very often asks the next question: *the classes look separated there, so why not feed those two columns to the classifier?* Interviewers ask this because the reasoning sounds plausible and the answer requires knowing what the projection actually is. ## The case against, for t-SNE **There is no mapping.** t-SNE's fitted parameters are the coordinates of the rows it was given. There is no function to apply to a future record, so a model trained on those two columns has nothing to consume at scoring time. This alone ends it for any model that has to run in production. **The coordinates are not a property of the row.** A row's position depends on the other rows in the run, on the random initialisation and on the perplexity. Rerun with a different seed and every value changes. A feature whose value moves when unrelated rows are added is not a feature. **Two numbers are a severe bottleneck.** Compressing dozens or hundreds of columns into two throws away nearly all information; whatever the eye can see, the model could see better in the original space. And the projection is unsupervised, so nothing guarantees the retained directions are the ones the target depends on. **Distances are distorted anyway.** The layout preserves neighbourhoods and deliberately exaggerates separations, so a distance-based model fitted on the coordinates is consuming a distorted geometry. ## The narrower case for, with UMAP UMAP changes two of those objections. It can embed new rows into a fitted embedding, so a scoring path exists; and it is usable at more than two or three output dimensions, so the bottleneck can be widened. A ten- or twenty-dimensional UMAP representation, fit once and frozen, is a genuine option, and it sometimes helps models whose behaviour depends on local distance structure, on data with strong non-linear manifold structure. The conditions are strict. - **Fit inside the fold.** The projection is fit on data, which makes it part of the model. Fitting it on the whole dataset before splitting lets the geometry of the validation rows shape the training features. The score you then measure is optimistic and will not survive deployment. Fit on the training fold, apply to the validation fold, repeat. - **If labels guided the projection, the leak is worse.** UMAP can use labels to shape the embedding. Doing that on all rows and then evaluating on some of them is target leakage in a particularly convincing disguise — the map will look wonderful and the estimate will be worthless. - **Earn it against a baseline.** Compare against the model on the original features, with the same tuning effort. If the projection does not beat that, it is added machinery, added fit-time cost, added non-determinism and one more artefact to version. - **Two dimensions is the wrong number.** If you are choosing this for modelling rather than for looking, there is no reason to stop at the number of dimensions a screen has. ## What to do instead If the goal is fewer columns, a projection with an explicit reusable linear mapping is the ordinary answer: it applies deterministically to new rows and can be inspected. If the goal is better performance, a gradient-boosted tree ensemble on the original features is a stronger default than any 2-D sketch, and regularisation handles wide inputs without a projection step. If the goal is understanding, that is precisely what the map is for — keep it in the analysis, not in the pipeline. ## A caveat worth voicing There is one legitimate use of the picture that is not feature engineering: as a diagnostic. If the classes are visibly separated in a projection built without labels, that is encouraging evidence that a simple model on the original features can do the job. If they are hopelessly mixed, that is weak evidence and not a verdict, since a projection can hide separation that a supervised model would find in directions the layout discarded. Use it to form expectations, not to justify a design. ## The line to remember A t-SNE or UMAP map is an instrument for looking at data, and instruments do not go into the pipeline. If you catch yourself wanting the coordinates as columns, the real question is what structure you saw in the map and whether it can be expressed directly in the original features — a ratio, a threshold, an interaction — which is a feature you can define, explain and reproduce.

  • If you did use a UMAP representation as features, how would you validate it honestly?
    Treat the projection as part of the model. Fit it on each training fold only, apply it to the held-out fold, and score there. Then compare against the same model on the original features with equal tuning effort. If it does not win clearly, drop it — the extra fit cost and run-to-run variation are not free.
  • Classes look cleanly separated in an unsupervised projection — what may you conclude?
    That a simple model on the original features is likely to do well; the separation exists in the data, not just in the picture, since no labels shaped the layout. The converse does not hold: a mixed-looking map does not prove the classes are inseparable, because the layout may have discarded the directions that separate them.
  • Why is using labels to guide the projection before splitting so dangerous?
    The embedding then encodes target information about every row, including the rows you intend to evaluate on. The held-out score measures a model that has already seen those labels through the geometry, so it can look excellent while the deployed model, which never had that advantage, performs far worse.

saying these in an interview costs you the question

  • Feeds t-SNE coordinates into a production model
  • Fits the projection on all data before splitting
  • Assumes two dimensions retain the signal that matters
  • Uses labels in the projection then evaluates on the same rows
  • Skips the baseline on the original features

context