Can new rows be dropped onto an existing t-SNE map, or must it be refit?
answer
- coordinates are the parameters, not a function
- nothing to apply to an unseen row
- refit gives a different, unregistered layout
- one method embeds into a frozen map
- new populations land inside old islands
basics
~20 st-SNE has no mapping from feature space to map space: it optimises the coordinates of the rows it was given, so new rows force a full refit and a different layout. UMAP can place new rows into a frozen existing embedding.
solid answer
~50 st-SNE fits coordinates, not a function. Every map point is a free parameter moved by gradient descent, so there is nothing to apply to an unseen row: adding next week's data means re-running on the combined table, and because the objective is non-convex and randomly initialised, last week's islands come back rotated, resized or rearranged. The two pictures are not registered to each other, so week-over-week comparison is invalid. UMAP does support this: it locates a new row among the training points in its neighbour graph and optimises only that row's coordinates while the existing embedding stays fixed, keeping the map comparable over time. The catch is that new points can only land where the training manifold already exists — a genuinely new population is forced into an old island rather than appearing as one. For monitoring, freeze a map, refresh it on a schedule, and detect drift on the original features.
go deeper
Remember the headline: t-SNE has no way to place an unseen row, so adding data means refitting, and the refitted picture is not comparable to the old one.
Explain why — the fitted parameters are the map coordinates themselves, so there is no function to apply, unlike a projection that stores a reusable matrix.
Turn it into a recommendation: freeze a reference embedding for comparability, refresh it on a schedule, and state plainly that novel groups will be absorbed into existing islands rather than announcing themselves.
Decide whether a projection belongs in a recurring report at all. Weigh the cost of maintaining a frozen reference map and its refresh policy against monitoring the original features directly.
## Why t-SNE cannot do it The distinction that matters is between a method that learns a **mapping** and one that learns a **layout**. A linear projection learns a matrix: multiply any new row by it and you get coordinates, forever, deterministically. t-SNE learns no such object. Its parameters *are* the 2-D coordinates of the training rows. Optimisation moves those coordinates to minimise the divergence between the data's neighbour distribution and the map's. When it finishes, you hold `n` pairs of numbers and nothing that consumes a feature vector. So the honest answer to *add next week's rows to last week's map* is that there is no operation to perform. Your options are: 1. **Refit on the union.** Run t-SNE on last week's plus this week's rows together. This works and gives a valid map — of the combined data. It does not extend the old map; it replaces it. The old islands come back rotated, mirrored, differently sized and sometimes split or merged, because the objective is non-convex and the run started from a fresh random initialisation. If the stakeholder wanted to see *where the new rows fell relative to last week*, this does not answer it: nothing on the new picture is registered to the old one. 2. **Approximate placement.** You could find each new row's nearest neighbours among the old rows and place it at their average map position. That is a reasonable eyeball trick and worth naming, but it is an interpolation you invented, not something the method guarantees, and it cannot represent anything that lies off the old data. 3. **A parametric variant.** Parametric t-SNE trains an explicit neural mapping from feature space to map space against the same objective, which does give you a reusable function. It exists, it is more work, and it changes the character of the tool from an exploratory sketch to a fitted model that has to be maintained. ## What UMAP offers instead UMAP's layout is also optimised, but the method is defined so that new points can be embedded into an existing one. The procedure is: find the new row's nearest neighbours among the original training rows, build its edges into the already-computed graph, initialise it near those neighbours, and optimise only its coordinates while every existing point is held fixed. The old map is untouched, so the two weeks are drawn on the same canvas and are directly comparable by eye. Three caveats decide whether that is actually useful. - **The manifold is frozen.** New points are placed relative to what the map already knows. A genuinely new population — a new fraud pattern, a new cell type, a new product category — has no home, so it gets pushed into whichever existing island its neighbours happen to be in. The map will look reassuring precisely when something new has appeared. This is the failure mode to state out loud. - **Placement is approximate.** The embedded point sits where its neighbours pull it; it is not the position it would have taken had the map been fit with it present. - **Drift accumulates.** As the data distribution moves, an old frozen map becomes an increasingly poor description of new rows. Refresh on a schedule, and note the refit date on the figure. ## How to answer the stakeholder Separate the two things they might want. If they want **a current picture**, refit on all the data and present it as a new picture. Do not lay it beside last week's and invite comparison of positions; the axes are not shared. If they want **week-over-week comparison**, the projection is the wrong instrument on its own. Fit a UMAP embedding once on a reference window, freeze it, and place each week's rows into it — that gives a comparable visual. Then back it with something quantitative on the original features: population counts, a distribution comparison per feature, or a drift statistic. Every number goes to the original features; the frozen map is the visual index, not the evidence. If they want **an alert when something new appears**, the frozen map is actively misleading for the reason above, and you should be monitoring the original features directly. ## The interview shape What is being probed here is whether you know the difference between fitting coordinates and fitting a function, and whether you can carry that distinction into a sensible operational recommendation instead of just saying *t-SNE cannot do out-of-sample*. Say the mechanism, name UMAP's capability accurately, then name the manifold-freeze failure mode — that last part is what separates a textbook answer from an operational one.
- What exactly breaks if you just refit the map from scratch each week?The layouts are not registered to each other. The objective is non-convex with random initialisation, so islands rotate, swap sides and change area between runs even on identical data. Any week-over-week reading of movement is then measuring the optimiser, not the population, and it is very easy to narrate that noise as a trend.
- What is the main risk of embedding new rows into a frozen UMAP map?The map can only express structure it learned during fitting. A genuinely novel group has no region of its own, so its rows are placed inside whichever existing island their nearest neighbours occupy, and the picture looks normal exactly when something new has arrived. Detect novelty on the original features, not from the map.
- Is placing a new point at the average position of its nearest old neighbours acceptable?As a rough visual aid, yes, provided you label it as interpolation. It is not part of the method, it inherits every distortion of the layout, and it cannot represent anything lying away from the original data. Never let a position obtained this way support a quantitative claim.
saying these in an interview costs you the question
- Believes t-SNE learns a reusable projection matrix
- Compares two separately fitted maps position by position
- Assumes a frozen map will reveal a new population
- Refits weekly and narrates island movement as a trend
- Cannot say what the optimiser's parameters actually are