How do you choose how many principal components to keep?
answer
- each component's share of total variance
- cumulative curve, not a single number
- look for the elbow, then the tail
- it is really a tuned hyperparameter
basics
~20 sRead the explained-variance ratios - each component's share of total variance. Common rules are a cumulative threshold such as 95%, the elbow of a scree plot, or tuning the count against the downstream model's score.
solid answer
~50 sEvery component has an explained-variance ratio: its variance divided by the total variance of all components. Because PCA only rotates centred data, those ratios sum to one, so the cumulative curve tells you what fraction of the spread the first `k` components retain. Three rules get used. A cumulative threshold - the smallest `k` reaching 90%, 95% or 99% - is a convention, not a statistical result. A scree plot of ratio against component index, cut at the elbow, is the visual version of the same judgment. The honest answer in a modelling context is that `k` is a hyperparameter: cross-validate it against the metric you care about, because retained variance is not retained usefulness. I would also read the curve's shape - 3 of 200 spectrometer channels reaching 95% says the data is genuinely low-rank; a flat curve says PCA buys almost nothing.
go deeper
Know that each component reports a share of the total variance and that you add those shares in order until you reach a threshold such as 95%. Being able to describe a scree plot - ratio against component index, look for the elbow - is enough at this level.
Explain why the ratios sum to one, what the elbow actually signals about noise versus structure, and why a threshold is a convention rather than a result. Be ready to say that the count is a hyperparameter you would cross-validate when a downstream model exists.
Demonstrate that you decide with the downstream metric and a training-only fit, and that you can read a flat scree plot as evidence that the reduction is not worth having. Latency and memory budgets are legitimate drivers - name them explicitly when they are what decided it.
Own the reporting standard: what the team must record so a component count is reproducible and auditable, and how often a refit is allowed to change it. Decide when the compression is worth the extra pipeline stage at all.
## What the explained-variance ratio is Fit PCA and each component comes with a variance - the variance of the scores along that direction. The **explained-variance ratio** of a component is that variance divided by the sum of the variances of all components. Because PCA is a rotation of mean-centred data, total variance is conserved: the sum over all components equals the sum of the variances of the original columns, and the ratios therefore sum to exactly one. Add the ratios of the first `k` components in order and you get the **cumulative explained variance**, the fraction of the data's total spread that survives if you keep `k` components and discard the rest. Because the components are ordered by variance, this curve is always increasing and always concave-ish in practice: big gains early, diminishing returns later. Choosing `k` means choosing where on that curve to stop. ## Rule 1: a cumulative threshold The most common answer is "keep the smallest `k` whose cumulative explained variance reaches 95%." It is simple, reproducible and defensible as a convention. What it is *not* is a statistical criterion - there is no theory that says 95% is where signal ends and noise begins. 90%, 95% and 99% all appear in practice, and the right one depends on how much reconstruction error the downstream use can tolerate. State it as a convention when you use it; a candidate who defends 95% as if it were derived is over-claiming. ## Rule 2: the scree plot and its elbow Plot the explained-variance ratio (or the raw component variance) against the component index, in order. On data with real low-rank structure this drops steeply and then flattens into a long, nearly horizontal tail. The **elbow** - the index where the steep part gives way to the flat tail - is a natural cut: components in the tail are all contributing about the same tiny amount, which is the signature of noise spread evenly across directions rather than structure concentrated in a few. The canonical clean case: a spectrometer producing 200 wavelength channels per sample. Neighbouring channels are enormously correlated, because a real absorbance peak spans many of them, and three components can carry 95% of the variance while the remaining 197 sit in a flat tail of instrument noise. Here the elbow and the 95% rule agree and both are convincing. The informative failure case is a scree plot with no elbow at all - a slow, even decay. That says variance is spread almost uniformly over directions, the data is not low-rank in any linear sense, and you will need most of the components to reach any threshold. The right conclusion is usually that PCA is not buying you much here, not that you should force a small `k` anyway. ## Rule 3: treat `k` as a hyperparameter If the components feed a model, the criterion that matters is the model's performance, not the retained variance. Cross-validate over a grid of `k`, score each with the metric you actually optimise, and pick the value that wins - with the usual preference for the simpler model when scores are within noise. This is the answer that separates candidates who have used PCA in a pipeline from candidates who have read about it: variance retained is a property of the inputs alone, and the inputs alone cannot tell you what the model needs. Fit the decomposition and choose `k` using training data only, then apply the same fitted components to held-out data, so the choice is not informed by the data you are scoring on. ## Two named criteria you may hear **Kaiser's criterion** - keep components whose eigenvalue exceeds 1 - applies when PCA is run on a correlation matrix, where each standardised variable contributes exactly 1 unit of variance, so "eigenvalue above 1" means "this component summarises more than a single variable's worth of information". It is crude and known to over-retain, but you will meet it in psychometrics-flavoured settings. **Parallel analysis** is the more serious version of the same idea: generate datasets of the same shape with no real structure, compute their component variances, and keep only the components whose variance exceeds what pure noise of that shape produces. It is more work and considerably more defensible than a fixed threshold. ## Practical constraints that often decide it for you Sometimes `k` is set by something other than the curve. A latency or memory budget caps how many features the downstream system can carry. A visualisation forces `k` = 2 or 3 regardless of what the variance says. A downstream method may become unstable above some dimension. These are legitimate reasons - just be explicit that the constraint, not the variance profile, drove the choice, and report how much variance that costs. ## What to report Whichever rule you use, report the same three things: the `k` you chose, the cumulative explained variance at that `k`, and the rule that produced it. "We kept 8 of 60 components, retaining 91% of the variance, chosen by cross-validating the downstream AUC over `k` in 2 to 20" is a complete and auditable statement. "We used PCA" is not.
- What does a scree plot with no visible elbow tell you?That variance is spread almost evenly across directions, so the data has no low-rank linear structure. You would need most of the components to reach any threshold, which means PCA is compressing almost nothing. The right response is usually to conclude that a linear variance-based reduction is the wrong tool here, not to force a small component count anyway.
- Do the explained-variance ratios of all the components sum to one?Yes. PCA is a rotation of mean-centred data, so the total variance is unchanged and merely redistributed across the new directions. Each ratio is one component's variance over that conserved total, so the full set sums to one and the kept components' ratios sum to the fraction retained.
- Is keeping 95% of the variance a principled threshold?No, it is a convention. Nothing guarantees the discarded 5% is noise or that the retained 95% is useful - variance is measured on the inputs alone and has no knowledge of what you are predicting. Use it as a starting point and a reporting convention, and let a cross-validated downstream score make the real decision where one exists.
- How would you report the choice so a reviewer can audit it?Give the component count, the cumulative explained variance at that count, and the rule that produced it - for example, eight components retaining 91% of the variance, selected by cross-validating the downstream metric over a range of counts. All three are needed: the number alone hides how much was discarded, and the variance alone hides how the number was chosen.
saying these in an interview costs you the question
- Treats 95% as a statistically derived threshold
- Fixes the component count before inspecting the variance curve
- Reads only the first component's ratio, never the cumulative
- Selects the component count using test-set performance
- Forces a small count when the scree plot is flat