skip to content

Linear and Manifold Projections

Turning many columns into a few, because distance loses meaning in high dimensions: PCA keeps variance, t-SNE and UMAP serve the eye only, and other projections buy speed or class separation.

on this pageshow

explore

questions

16

What does PCA maximise when it picks the first principal component?

level: juniorimportance: must knowfreq 78%

answer

  1. a direction, not a chosen column
  2. means subtracted before anything else
  3. spread of the projected points
  4. later ones must be orthogonal

basics

~20 s

PCA picks the first principal component as the unit-length direction in the mean-centred data along which the projected points have the largest variance. Each later component maximises the remaining variance while staying orthogonal to all earlier ones.

solid answer

~50 s

First you centre: subtract each column's mean so the cloud of points sits around the origin. PCA then searches over unit-length directions and returns the one along which the projected points spread out the most - that is the first principal component, and its variance is the largest of any direction. The second component repeats the search but is restricted to directions orthogonal to the first, the third to directions orthogonal to both, and so on. Equivalently, the first `k` components define the `k`-dimensional subspace that minimises the squared reconstruction error of the points, so "keeps the most variance" and "loses the least when you project" are the same objective. Two consequences matter in interviews: a component is a weighted blend of every original column, not a subset of them; and PCA never looks at a target variable, so it is purely a description of how the inputs vary.

go deeper

for a junior

Be ready to say, in one sentence, that PCA finds directions of maximum spread in mean-centred data and that each new direction is at right angles to the previous ones. Knowing that a component blends all columns rather than picking some is the point that gets checked.

for a middle

Expect to explain the actual optimisation: unit-length direction, variance of the projected scores, orthogonality constraint on each subsequent component, and the equivalent reading as minimising reconstruction error. Be able to say concretely what breaks without centring.

for a senior

Show you know where the assumptions bite in real data: linearity, scale dependence and the fact that no target is involved. Be ready to say when a rotation into uncorrelated components genuinely helps a downstream model and when it only adds a step.

for a principal

Own the framing question of whether a linear variance-based rotation is the right tool for the data at all, and what a team gives up in traceability and pipeline complexity by putting one in front of every model by default.

## The setup Start with a table of `n` rows and `p` numeric columns. PCA rewrites those `p` columns as `p` new columns - the principal components - which are ordered so that the first carries as much of the data's variance as any single direction can, the second as much of what is left as any direction orthogonal to the first can, and so on. Nothing is thrown away until you decide to keep only the first few. ## Step zero: centring PCA is defined on **mean-centred** data. For each column you subtract that column's mean, so every column now has mean zero and the cloud of points is positioned around the origin. This is not cosmetic. Variance is defined as spread *about the mean*; the optimisation PCA performs on centred data maximises the average squared projection, and on centred data the average squared projection **is** the variance. Skip the centring and the same optimisation maximises the raw second moment instead, which is variance plus the squared offset of the cloud from the origin. With columns that sit far from zero - incomes, temperatures in kelvin, test scores out of 800 - that offset dominates, so the "first component" points roughly at the mean vector, i.e. at where the data *is*, not at the direction in which it *varies*. Every later component is then forced to be orthogonal to that useless direction, and the whole decomposition is wrong. Whether you should additionally rescale the columns before running PCA is a separate preprocessing decision with its own trade-offs. ## The objective Consider any unit-length direction vector `v` in the `p`-dimensional space. Projecting a centred row `x` onto it gives a single number, the score `s = x . v` (the dot product). Do this for all `n` rows and you get a one-dimensional summary of the data with some variance. PCA asks: **over all unit-length `v`, which one makes that variance largest?** The answer is the first principal component, and the resulting variance is that component's variance. The second component solves the identical problem with one extra constraint: `v` must be orthogonal (at right angles) to the first component. The third is orthogonal to the first two, and so on. Orthogonality is what makes PCA a *rotation* of the data rather than an arbitrary re-encoding, and it is why the component scores are mutually uncorrelated across the sample - the classic reason PCA is reached for when the original columns are badly collinear. ## The equivalent reconstruction view There is a second description of the same objective that is worth having ready. Take only the first `k` components and use them to describe each row: you get an approximate version of the row, `x` reconstructed as the column means plus a weighted sum of the `k` component directions. The squared distance between the true row and this reconstruction is the reconstruction error. The subspace that **minimises** total squared reconstruction error over the whole dataset is exactly the subspace spanned by the first `k` components. "Keep the most variance" and "lose the least when you project" are two readings of one optimisation, because on centred data total variance splits cleanly into the part you keep and the part you discard. ## Loadings and scores Two different objects get muddled constantly. The **loadings** are the coefficients that define a component - one weight per original column, describing how the component is built. The **scores** are the numbers you get after projecting the rows - one value per row per component, and these are the new columns you feed downstream. A candidate who says "PCA picks the most important features" has confused this: a component is a dense weighted mixture of *all* `p` columns, and a column with a near-zero loading has not been "unselected", it simply contributes little to that particular direction. ## Variance accounting and sign Because PCA only rotates centred data, the sum of the variances of all `p` components equals the sum of the variances of the original `p` columns. Total variance is conserved and merely redistributed, concentrated into the leading directions. That conservation is what makes the per-component share of the total - the explained-variance ratio - a meaningful quantity. One practical wrinkle: the sign of a component is arbitrary. A direction `v` and its negation `-v` span the same axis and give the same variance, so different runs or implementations may return either. Loadings and scores flip together, so the geometry and any downstream model are unchanged - but do not read meaning into a sign, and do not compare signs across two separately fitted decompositions. ## Assumptions to state out loud PCA is linear: it can only find flat subspaces, so structure that curves through the space is not captured. It is variance-based, and variance is a scale-dependent quantity, so the relative units of your columns change the answer. And it is **unsupervised**: it never sees a target, so it ranks directions by how much the inputs move, not by how useful they are for predicting anything.

  • What goes wrong if you skip centring and run PCA on the raw columns?
    The optimisation then maximises the average squared projection rather than the variance, and those differ by the squared offset of the cloud from the origin. With columns far from zero that offset dominates, so the first component points roughly at the mean vector - where the data sits - instead of the direction of maximum spread. Every later component is then forced orthogonal to that, and the decomposition describes position rather than variation.
  • Are the principal component scores correlated with one another?
    No. The components are mutually orthogonal directions, and the resulting scores are uncorrelated across the sample. That is precisely why PCA is a standard response to a set of badly collinear columns: the new columns carry the same total variance with none of the redundancy between them.
  • Why can the sign of a principal component flip between two runs?
    Only the axis is determined, not its orientation: a direction and its negation give identical variance, so either may be returned. Loadings and scores flip together, so distances, variances and any downstream model are unaffected. The practical rule is to never interpret a sign on its own and never compare signs across separately fitted decompositions.

Photographing a swarm of bees: the first principal component is the camera angle from which the swarm looks widest. The second is the widest angle among all the angles at right angles to the first.

saying these in an interview costs you the question

  • Says PCA selects the most important original columns
  • Thinks centring is optional cosmetic preprocessing
  • Claims PCA uses the target variable to rank directions
  • Believes components need not be orthogonal to each other
  • Confuses the loadings with the projected scores

context

open as a page

On a t-SNE map, why is the gap between two clusters not a real distance?

level: juniorimportance: must knowfreq 70%

basics

~20 s

t-SNE only preserves each point's near neighbours. It minimises a divergence that punishes tearing neighbours apart but barely punishes moving distant points, so between-cluster gaps and island sizes are artefacts of the layout, not measured distances.

open as a page

How does linear discriminant analysis choose its projection axes differently from PCA?

level: middleimportance: must knowfreq 58%

basics

~20 s

Linear discriminant analysis uses the labels. It picks axes that maximise the spread between class means relative to the spread inside each class. PCA ignores labels and picks directions of largest total variance, which need not separate classes at all.

open as a page

How do you choose how many principal components to keep?

level: middleimportance: must knowfreq 68%

basics

~20 s

Read the explained-variance ratios - each component's share of total variance. Common rules are a cumulative threshold such as 95%, the elbow of a scree plot, or tuning the count against the downstream model's score.

open as a page

How does latent Dirichlet allocation model a document as a mixture of topics?

level: middleimportance: must knowfreq 60%

basics

~20 s

Latent Dirichlet allocation treats each document as a probability distribution over K topics, and each topic as a probability distribution over vocabulary words. Fitting infers both, so one support ticket comes out as 0.6 billing, 0.3 login, 0.1 shipping.

open as a page

In t-SNE, what does the perplexity setting actually control?

level: middleimportance: must knowfreq 58%

basics

~20 s

Perplexity sets the effective number of near neighbours each point is fitted to. Each point's Gaussian kernel width is tuned by search until its neighbour distribution has that perplexity, so the setting chooses the scale at which structure is preserved.

open as a page

When would you prefer NMF over latent semantic analysis for extracting topics?

level: middleimportance: should knowfreq 45%

basics

~10 s

Prefer non-negative matrix factorization when people must read the topics: non-negativity makes every topic an additive list of words. Latent semantic analysis produces signed components that capture synonymy well but rarely read as themes.

open as a page

In random projection, what sets the target dimension for a 100,000-column feature matrix?

level: seniorimportance: should knowfreq 33%

basics

~20 s

The Johnson-Lindenstrauss bound sets it from the number of points and the distortion you accept, roughly log(n) divided by epsilon squared. The original 100,000 columns do not appear in the bound at all, and the bound itself is conservative.

open as a page

Why would a classifier get worse after PCA retained 95% of the input variance?

level: seniorimportance: should knowfreq 46%

basics

~20 s

PCA ranks directions by how much the inputs vary, never by how much they predict the label. A discriminative direction with small variance can sit entirely in the discarded 5%, so retained variance and retained signal are different quantities.

open as a page

A topic model's coherence peaks at 18 topics while held-out perplexity keeps improving — how do you choose K?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The two metrics answer different questions. Perplexity measures held-out predictive fit and usually keeps improving as topics multiply; coherence measures whether a topic's top words really co-occur. If humans read the topics, follow coherence and inspect the top words yourself.

open as a page

Can new rows be dropped onto an existing t-SNE map, or must it be refit?

level: seniorimportance: should knowfreq 40%

basics

~20 s

t-SNE has no mapping from feature space to map space: it optimises the coordinates of the rows it was given, so new rows force a full refit and a different layout. UMAP can place new rows into a frozen existing embedding.

open as a page

Extract 100 projected components or select 100 of 1,000 original columns for a latency budget?

level: principalimportance: should knowfreq 40%

basics

~20 s

Selection cuts what you must collect and keeps every input explainable; extraction keeps signal spread thinly across all 1,000 columns but still requires computing every one of them at serving time. If the latency is upstream, only selection helps.

open as a page

When does kernel PCA find structure that ordinary linear PCA cannot?

level: middleimportance: nice to knowfreq 24%

basics

~20 s

When the structure is curved. Linear PCA can only rotate axes and drop some, so a manifold that folds back on itself is flattened into an overlapping smear. Kernel PCA works from pairwise similarities, so its components can bend.

open as a page

Should 2-D t-SNE or UMAP coordinates be used as features for a downstream model?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

For t-SNE, no: it produces no reusable mapping, its layout changes run to run, and two coordinates discard almost everything. A frozen UMAP embedding fit only on training folds can be defensible; treat it as a modelling choice.

open as a page

How do you justify a PCA-based credit score to a regulator asking which ratio drove it?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Each principal component is a weighted blend of every input ratio, so per-component reasons are useless to an applicant. If the model on top is linear, compose the loadings with its coefficients to recover one exact weight per ratio.

open as a page

Every retrain of a production topic model returns different topics — how do you deliver something stakeholders can trust?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Topic models are not identifiable: random initialisation, a chosen topic count and shifting text mean each refit gives different, differently numbered topics. Freeze one fitted model as the published taxonomy, refit on a deliberate schedule, and review changes with humans.

open as a page