skip to content

Why would a classifier get worse after PCA retained 95% of the input variance?

level: seniorimportance: should knowfreq 46%

answer

  1. PCA never looks at the target
  2. variance ranking is not signal ranking
  3. the signal can live in the discarded tail
  4. trees split on axes, components are blends

basics

~20 s

PCA ranks directions by how much the inputs vary, never by how much they predict the label. A discriminative direction with small variance can sit entirely in the discarded 5%, so retained variance and retained signal are different quantities.

solid answer

~50 s

The key sentence is that PCA never sees the target. It orders directions by input variance, so the 95% cut keeps whatever moves the most - which may be a noisy high-amplitude sensor - and discards whatever moves the least, which may be the low-amplitude direction that separates the classes. Retaining 95% of the variance is emphatically not retaining 95% of the information about `y`. There are two secondary causes worth naming. Components are dense blends of every column, so a tree-based model that used to split on one meaningful threshold now has to approximate that with axis-aligned splits on mixtures. And the component count was probably fixed by a rule rather than tuned. My diagnosis would be: check whether any discarded component is associated with the target, then tune the count by cross-validating the actual metric, and compare against the raw features as the baseline.

go deeper

for a junior

The one fact to carry away is that PCA never sees the target, so keeping the highest-variance directions is not the same as keeping the most predictive ones. Saying that clearly is a strong answer at this level.

for a middle

Explain the mechanism: variance is computed on the inputs alone, a low-variance direction can be the discriminative one, and a threshold on retained variance is therefore blind to the label. Be able to name the component count as an untuned hyperparameter.

for a senior

Walk through a real diagnosis - raw-feature baseline on the same folds, checking whether discarded components relate to the target, sweeping the count against the production metric, and noticing when the model family is the problem. Judgment about when the compression is still worth shipping is expected.

for a principal

Own the standard that no reduction step enters a pipeline on retained variance alone: a baseline comparison and a metric-based justification are the entry requirement, and someone must decide what accuracy the team will trade for latency or stability.

## The core reason: PCA is unsupervised PCA is fitted on the input matrix alone. It searches for directions of maximum variance in the mean-centred inputs and orders components by that variance. The target never enters the computation. So the ordering answers the question "along which direction do these rows differ from one another the most?" and *not* "along which direction do the two classes differ the most?" Those can be completely different directions. A concrete way to see it: imagine one input that is a loud, high-amplitude measurement whose value is unrelated to the label, and another that varies over a narrow range but shifts consistently between the two classes. PCA puts the loud one in the leading component and may push the narrow, discriminative one into the tail. If your rule was "keep 95% of the variance", the useful direction is exactly what got cut. Nothing has gone wrong with PCA - it did what it optimises. The mistake is treating a variance ranking as a signal ranking. This also explains why "we kept 95% of the variance" is a misleading sentence in a report. Ninety-five percent of the *inputs'* spread is retained. How much of the *label-relevant* structure is retained is unmeasured, and could be 100% or could be nearly none. ## Second cause: components are dense mixtures Each component is a weighted blend of every original column. That matters differently by model family. For a linear model with strong collinearity among inputs, replacing raw columns with orthogonal components often *helps*: coefficients become stable, the problem is better conditioned, and the model has fewer parameters to fit. For tree-based models the trade usually runs the other way. A tree splits on one feature at a threshold. When a single raw column carried a clean decision boundary - a ratio above a level, a count below one - the tree found it in one split. After the rotation, that same boundary is a slanted plane in the component space, and the tree has to approximate it with a staircase of axis-aligned splits, costing depth, data and accuracy. Distance-based models sit in between: they usually benefit when the raw columns are numerous and correlated, and lose when the discarded directions mattered. ## Third cause: the count was never tuned A fixed threshold picks the number of components from a property of the inputs. If nobody cross-validated that count against the downstream metric, the pipeline is running on an untuned hyperparameter. It is common for a sweep to show that a noticeably larger count - or none at all - wins. ## Fourth cause: the structure is not linear PCA finds flat subspaces. If the class boundary curves through the input space, no linear projection is guaranteed to preserve it, and a low-variance direction may be exactly what a curved boundary needs. ## How to diagnose it 1. **Establish the baseline honestly.** Score the same model on the raw features, with the same folds. If PCA is a regression relative to that baseline, you have a real finding, not noise. 2. **Interrogate the discarded components.** Fit with all components retained and look at how the model weights each one, or measure each individual component's association with the target on the training split. If a component that the 95% rule discarded ranks high, you have located the cause directly. 3. **Sweep the count.** Cross-validate over a range of component counts against the metric you actually optimise. A curve that keeps rising as you add components tells you variance-based truncation is not the right filter for this problem. 4. **Check the model family.** If it is a tree ensemble, try the raw features; the rotation is fighting how the model represents decisions. 5. **Check the leading component.** If component one is dominated by a single loud, uninformative input, the variance ranking is being driven by an artefact. ## What to do about it If discarded directions carry signal, the fix is to stop letting an unsupervised variance criterion do your feature reduction: select or reduce features using their relationship with the target instead, or keep more components, or keep the raw features and address collinearity another way. If the loss is small and PCA buys something real - inference latency, memory, a much better-conditioned model, stability under collinearity - it can still be the right call; just make that trade explicitly, on the metric that matters, rather than on the sentence "we retained 95% of the variance". ## The interview answer in one line "Because variance is a property of `X` alone. PCA ranks directions by how much the inputs move, and the direction that separates the classes may barely move at all - so a 95% variance cut can throw away exactly the signal the classifier needed."

  • How would you check whether a discarded component actually carried signal?
    Refit keeping all components, then look at how strongly the model relies on each one, or measure each component's individual association with the target on the training split. If a component that the threshold discarded ranks near the top, you have found the cause directly rather than inferring it. Do this on training folds only, so the check does not contaminate your evaluation.
  • Which model families most often lose accuracy when raw columns are replaced by components?
    Tree-based models. They split on one feature at a threshold, and a component is a dense blend of every column, so a boundary that used to be one clean split becomes a slanted plane approximated by a staircase of splits. Linear models with heavy collinearity usually go the other way and benefit, since the components are uncorrelated and the fit is better conditioned.
  • If PCA costs you a point of accuracy, why might you still ship it?
    Inference latency, memory footprint, training cost on very wide data, or a much better-conditioned and more stable model. A one-point loss can easily be worth a large reduction in serving cost or in coefficient instability. Make that a stated trade on the metrics that matter, and record what the raw-feature baseline scored.

A recording dominated by a loud air conditioner: the loudest component of the sound carries almost all the energy, but the whispered password you needed is in the quiet part you filtered out.

saying these in an interview costs you the question

  • Assumes the high-variance directions are the informative ones
  • Believes PCA consults the labels when ordering components
  • Says 95% variance retained means 95% of the information
  • Never compares against a raw-feature baseline
  • Leaves the component count untuned against the real metric

context