skip to content

How does linear discriminant analysis choose its projection axes differently from PCA?

level: middleimportance: must knowfreq 58%

answer

  1. one uses labels, one does not
  2. signal over noise, not total spread
  3. between-class scatter over within-class scatter
  4. class count caps the axis count

basics

~20 s

Linear discriminant analysis uses the labels. It picks axes that maximise the spread between class means relative to the spread inside each class. PCA ignores labels and picks directions of largest total variance, which need not separate classes at all.

solid answer

~50 s

Both produce a low-dimensional linear projection, but they optimise different things. PCA is unsupervised: it looks for directions along which the data as a whole varies most. Linear discriminant analysis is supervised: it forms a between-class scatter matrix from how far the class means sit from the overall mean, and a within-class scatter matrix from how much each class spreads around its own mean, then finds the directions that maximise the ratio `between / within` — solved as a generalised eigenvalue problem. That is why on 30 wearable gait-sensor features you can drop to two axes that visibly separate walking, climbing and cycling, while a variance-maximising projection might spend its axes on a noisy high-amplitude sensor. The catch: with C classes the between-class scatter has rank at most C-1, so linear discriminant analysis can return at most C-1 axes — two for three classes, one for a binary problem.

go deeper

for a junior

Be ready to say plainly that one method sees the labels and the other does not, and that only the label-aware one is trying to pull the classes apart.

for a middle

Explain the objective in terms of two scatter quantities — spread between class means over spread within classes — and state the C-1 limit on the number of axes and where it comes from.

for a senior

Show you know the failure modes you have actually hit: a singular within-class scatter when features outnumber samples, classes that differ in spread rather than mean, and the leakage risk of fitting a supervised projection outside the training folds.

for a principal

Own the call about whether the pipeline should carry a projection tuned to one label at all, given that other teams and future targets consume the same features and a label-specific projection quietly discards what they need.

## The shared setup Both methods take a matrix of n rows and d columns and return a linear map to a smaller number of columns. Each new feature is a weighted sum of the original ones. What separates them is the objective the weights are chosen to optimise, and whether labels are allowed to influence that choice. ## What PCA optimises PCA is unsupervised. It searches for the direction along which the projected data has the largest variance, then the next such direction orthogonal to the first, and so on. Labels never enter the computation. This is exactly right when the goal is compression or denoising, and it is a common default before a classifier — but nothing in the objective says a high-variance direction is a class-discriminating direction. A sensor with a large dynamic range and no relationship to the label will dominate the first direction; a small but consistent offset between two classes may be relegated to a direction PCA discards. ## What linear discriminant analysis optimises (Not to be confused with latent Dirichlet allocation, which shares the initials and is a topic model — this is Fisher's discriminant.) Linear discriminant analysis is supervised and builds two matrices: - **Within-class scatter**: for each class, the spread of its points around its own class mean, summed over classes. This measures noise — variation that has nothing to do with which class a point belongs to. - **Between-class scatter**: how far the class means sit from the overall mean, each weighted by the class size. This measures signal — variation that tracks the label. The method looks for the direction w that maximises the ratio of projected between-class scatter to projected within-class scatter. Informally: push the class means apart while squeezing each class tight. The solution comes from a generalised eigenvalue problem involving the two scatter matrices, and the eigenvectors with the largest eigenvalues are the discriminant axes. Concretely, with 30 features from a wearable gait sensor and three activity labels — walking, climbing stairs, cycling — the two discriminant axes tend to produce a plot with three visibly distinct clouds, because every unit of variance spent is variance that moves the class means apart. A variance-maximising projection of the same 30 columns can put the classes on top of each other if the loudest sensor happens to be label-irrelevant. ## The C-1 ceiling With C classes there are C class means, and they are constrained by the overall mean, so the between-class scatter matrix has rank at most C-1. Only C-1 eigenvalues can be non-zero, so linear discriminant analysis returns at most C-1 axes regardless of how many features you started with. Three activity classes gives two axes; a binary problem gives exactly one. This is a hard structural limit, not a tuning choice, and it is the single most common reason the method cannot be used as a general-purpose dimensionality reducer: 100 features and 2 classes collapse to a single number. ## Assumptions and where it breaks The classical derivation is optimal when each class is Gaussian and all classes share the same covariance matrix. Deviations degrade it in predictable ways: - **Equal means, different spreads.** If two classes have the same centre but different variances, the between-class scatter is near zero and no discriminant direction exists — yet the classes are perfectly separable by other means. - **Very different covariances.** A shared-covariance assumption misplaces the boundary; a quadratic variant that estimates a covariance per class is the usual answer, at the cost of many more parameters. - **More features than samples.** The within-class scatter matrix becomes singular and cannot be inverted, so the objective is ill-posed without regularisation or a preliminary reduction step. - **Multi-modal classes.** A class made of two distant clusters has a mean that sits between them, which the objective treats as the class location. ## Choosing between them in practice Use linear discriminant analysis when you have labels, few classes, comfortably more samples than features, and you want a projection tuned to those specific labels. Use an unsupervised projection when there are no labels, when you need more components than C-1, or when the projection must serve several downstream tasks — a supervised projection is fitted to one label and may throw away exactly the variance a different target needs. And whichever you pick, fit it inside the training folds only: fitting a supervised projection on the full dataset lets the test labels shape the features and inflates every score that follows.

  • How many axes can linear discriminant analysis return for a three-class problem, and why that number?
    Two. The between-class scatter matrix is built from three class means that are tied to the overall mean, so its rank is at most C-1 = 2 and only two eigenvalues can be non-zero. The limit depends on the number of classes, never on the number of features — a binary problem yields exactly one axis no matter how wide the input.
  • When would an unsupervised projection beat linear discriminant analysis on a labelled dataset?
    When you need more than C-1 components; when features outnumber samples, so the within-class scatter matrix is singular; when classes differ in spread rather than in mean, leaving almost no between-class scatter; or when the projection has to feed several different targets, since a supervised projection is fitted to one label and discards variance the other targets may need.
  • What distribution assumption sits behind the classical version, and what happens when it fails?
    It is optimal when each class is Gaussian with the same covariance matrix. When covariances differ markedly the shared-covariance boundary is misplaced and a per-class covariance variant fits better, at the cost of many more parameters to estimate. When a class is multi-modal its mean falls between its own clusters, so the objective optimises toward a location no points occupy.

PCA photographs a crowd from wherever the crowd looks widest. Linear discriminant analysis walks around until the two teams stop overlapping, even if that angle makes the crowd look narrow.

saying these in an interview costs you the question

  • Calls linear discriminant analysis unsupervised, like PCA
  • Says it can return as many axes as there are features
  • Confuses it with latent Dirichlet allocation topic modelling
  • Assumes the highest-variance direction always separates classes
  • Claims it always beats an unsupervised projection before a classifier
  • Ignores that it needs more samples than features to be well-posed

context