skip to content

PCA and Explained Variance

Centre and usually standardise first, then keep components until the explained-variance ratio is enough, read off a scree plot or a 95% rule. Interviewers ask when PCA hurts a supervised model.

on this pageshow

questions

4

What does PCA maximise when it picks the first principal component?

level: juniorimportance: must knowfreq 78%

answer

  1. a direction, not a chosen column
  2. means subtracted before anything else
  3. spread of the projected points
  4. later ones must be orthogonal

basics

~20 s

PCA picks the first principal component as the unit-length direction in the mean-centred data along which the projected points have the largest variance. Each later component maximises the remaining variance while staying orthogonal to all earlier ones.

solid answer

~50 s

First you centre: subtract each column's mean so the cloud of points sits around the origin. PCA then searches over unit-length directions and returns the one along which the projected points spread out the most - that is the first principal component, and its variance is the largest of any direction. The second component repeats the search but is restricted to directions orthogonal to the first, the third to directions orthogonal to both, and so on. Equivalently, the first `k` components define the `k`-dimensional subspace that minimises the squared reconstruction error of the points, so "keeps the most variance" and "loses the least when you project" are the same objective. Two consequences matter in interviews: a component is a weighted blend of every original column, not a subset of them; and PCA never looks at a target variable, so it is purely a description of how the inputs vary.

go deeper

for a junior

Be ready to say, in one sentence, that PCA finds directions of maximum spread in mean-centred data and that each new direction is at right angles to the previous ones. Knowing that a component blends all columns rather than picking some is the point that gets checked.

for a middle

Expect to explain the actual optimisation: unit-length direction, variance of the projected scores, orthogonality constraint on each subsequent component, and the equivalent reading as minimising reconstruction error. Be able to say concretely what breaks without centring.

for a senior

Show you know where the assumptions bite in real data: linearity, scale dependence and the fact that no target is involved. Be ready to say when a rotation into uncorrelated components genuinely helps a downstream model and when it only adds a step.

for a principal

Own the framing question of whether a linear variance-based rotation is the right tool for the data at all, and what a team gives up in traceability and pipeline complexity by putting one in front of every model by default.

## The setup Start with a table of `n` rows and `p` numeric columns. PCA rewrites those `p` columns as `p` new columns - the principal components - which are ordered so that the first carries as much of the data's variance as any single direction can, the second as much of what is left as any direction orthogonal to the first can, and so on. Nothing is thrown away until you decide to keep only the first few. ## Step zero: centring PCA is defined on **mean-centred** data. For each column you subtract that column's mean, so every column now has mean zero and the cloud of points is positioned around the origin. This is not cosmetic. Variance is defined as spread *about the mean*; the optimisation PCA performs on centred data maximises the average squared projection, and on centred data the average squared projection **is** the variance. Skip the centring and the same optimisation maximises the raw second moment instead, which is variance plus the squared offset of the cloud from the origin. With columns that sit far from zero - incomes, temperatures in kelvin, test scores out of 800 - that offset dominates, so the "first component" points roughly at the mean vector, i.e. at where the data *is*, not at the direction in which it *varies*. Every later component is then forced to be orthogonal to that useless direction, and the whole decomposition is wrong. Whether you should additionally rescale the columns before running PCA is a separate preprocessing decision with its own trade-offs. ## The objective Consider any unit-length direction vector `v` in the `p`-dimensional space. Projecting a centred row `x` onto it gives a single number, the score `s = x . v` (the dot product). Do this for all `n` rows and you get a one-dimensional summary of the data with some variance. PCA asks: **over all unit-length `v`, which one makes that variance largest?** The answer is the first principal component, and the resulting variance is that component's variance. The second component solves the identical problem with one extra constraint: `v` must be orthogonal (at right angles) to the first component. The third is orthogonal to the first two, and so on. Orthogonality is what makes PCA a *rotation* of the data rather than an arbitrary re-encoding, and it is why the component scores are mutually uncorrelated across the sample - the classic reason PCA is reached for when the original columns are badly collinear. ## The equivalent reconstruction view There is a second description of the same objective that is worth having ready. Take only the first `k` components and use them to describe each row: you get an approximate version of the row, `x` reconstructed as the column means plus a weighted sum of the `k` component directions. The squared distance between the true row and this reconstruction is the reconstruction error. The subspace that **minimises** total squared reconstruction error over the whole dataset is exactly the subspace spanned by the first `k` components. "Keep the most variance" and "lose the least when you project" are two readings of one optimisation, because on centred data total variance splits cleanly into the part you keep and the part you discard. ## Loadings and scores Two different objects get muddled constantly. The **loadings** are the coefficients that define a component - one weight per original column, describing how the component is built. The **scores** are the numbers you get after projecting the rows - one value per row per component, and these are the new columns you feed downstream. A candidate who says "PCA picks the most important features" has confused this: a component is a dense weighted mixture of *all* `p` columns, and a column with a near-zero loading has not been "unselected", it simply contributes little to that particular direction. ## Variance accounting and sign Because PCA only rotates centred data, the sum of the variances of all `p` components equals the sum of the variances of the original `p` columns. Total variance is conserved and merely redistributed, concentrated into the leading directions. That conservation is what makes the per-component share of the total - the explained-variance ratio - a meaningful quantity. One practical wrinkle: the sign of a component is arbitrary. A direction `v` and its negation `-v` span the same axis and give the same variance, so different runs or implementations may return either. Loadings and scores flip together, so the geometry and any downstream model are unchanged - but do not read meaning into a sign, and do not compare signs across two separately fitted decompositions. ## Assumptions to state out loud PCA is linear: it can only find flat subspaces, so structure that curves through the space is not captured. It is variance-based, and variance is a scale-dependent quantity, so the relative units of your columns change the answer. And it is **unsupervised**: it never sees a target, so it ranks directions by how much the inputs move, not by how useful they are for predicting anything.

  • What goes wrong if you skip centring and run PCA on the raw columns?
    The optimisation then maximises the average squared projection rather than the variance, and those differ by the squared offset of the cloud from the origin. With columns far from zero that offset dominates, so the first component points roughly at the mean vector - where the data sits - instead of the direction of maximum spread. Every later component is then forced orthogonal to that, and the decomposition describes position rather than variation.
  • Are the principal component scores correlated with one another?
    No. The components are mutually orthogonal directions, and the resulting scores are uncorrelated across the sample. That is precisely why PCA is a standard response to a set of badly collinear columns: the new columns carry the same total variance with none of the redundancy between them.
  • Why can the sign of a principal component flip between two runs?
    Only the axis is determined, not its orientation: a direction and its negation give identical variance, so either may be returned. Loadings and scores flip together, so distances, variances and any downstream model are unaffected. The practical rule is to never interpret a sign on its own and never compare signs across separately fitted decompositions.

Photographing a swarm of bees: the first principal component is the camera angle from which the swarm looks widest. The second is the widest angle among all the angles at right angles to the first.

saying these in an interview costs you the question

  • Says PCA selects the most important original columns
  • Thinks centring is optional cosmetic preprocessing
  • Claims PCA uses the target variable to rank directions
  • Believes components need not be orthogonal to each other
  • Confuses the loadings with the projected scores

context

open as a page

How do you choose how many principal components to keep?

level: middleimportance: must knowfreq 68%

basics

~20 s

Read the explained-variance ratios - each component's share of total variance. Common rules are a cumulative threshold such as 95%, the elbow of a scree plot, or tuning the count against the downstream model's score.

open as a page

Why would a classifier get worse after PCA retained 95% of the input variance?

level: seniorimportance: should knowfreq 46%

basics

~20 s

PCA ranks directions by how much the inputs vary, never by how much they predict the label. A discriminative direction with small variance can sit entirely in the discarded 5%, so retained variance and retained signal are different quantities.

open as a page

How do you justify a PCA-based credit score to a regulator asking which ratio drove it?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Each principal component is a weighted blend of every input ratio, so per-component reasons are useless to an applicant. If the model on top is linear, compose the loadings with its coefficients to recover one exact weight per ratio.

open as a page