What do the eigenvectors of a 2x2 covariance matrix represent geometrically?
answer
- the data cloud drawn as an ellipse
- orthogonal because covariance is symmetric
- directions of maximum and minimum spread
- eigenvalue equals variance along that axis
- the eigenvalues add up to the trace
basics
~20 sThey are the principal axes of the data's elliptical cloud: orthogonal directions along which the data spreads independently. Each eigenvalue is the variance measured along its own eigenvector, and the eigenvalues sum to the total variance, the matrix trace.
solid answer
~40 sA covariance matrix is symmetric, so the spectral theorem applies: it has real eigenvalues and an orthonormal set of eigenvectors. Those eigenvectors point along the **principal axes** of the elliptical cloud the data forms, and each eigenvalue is the variance of the data projected onto its own eigenvector. Take `C = [[5,3],[3,5]]`. Its eigenvalues are `8` and `2`, with eigenvectors `(1,1)/sqrt(2)` and `(1,-1)/sqrt(2)`. So the cloud is stretched four times as much along the 45-degree diagonal as across it, and the ellipse's semi-axis lengths scale as `sqrt(8)` and `sqrt(2)`. The eigenvalues also sum to the trace, `8 + 2 = 10 = 5 + 5`, so rotating to the eigenbasis redistributes variance among axes without creating or destroying any. In that rotated frame the covariance is diagonal, meaning the rotated coordinates are uncorrelated.
go deeper
Know that the top eigenvector points along the direction of greatest spread and that its eigenvalue is the variance along that direction, not the standard deviation.
Be able to derive it: variance along a unit direction u is u^T C u, which equals lambda for a unit eigenvector, and explain why the axes must be orthogonal for a symmetric matrix.
Show the operational cautions: near-equal eigenvalues make the axes meaningless, a near-zero eigenvalue signals collinearity and unstable inversion, and units drive the whole result unless you standardize.
Own the framing decision of whether an orthogonal-axis view is the right summary for the data at all, and whether the team should analyse covariance or correlation given the measurement units in play.
## The setup A covariance matrix `C` for `d` variables is `d x d`, with the variance of variable `i` on the diagonal and the covariance of variables `i` and `j` off the diagonal. Two structural facts drive everything below: `C` is **symmetric**, because the covariance of `i` with `j` equals the covariance of `j` with `i`; and `C` is **positive semidefinite**, because no direction can have negative variance. Symmetry lets the spectral theorem apply: `C` has real eigenvalues and an orthonormal basis of eigenvectors, and can be written `C = Q D Q^T` with `Q` orthogonal and `D` diagonal. ## What the eigenvectors mean The key identity is that the variance of the data projected onto any unit direction `u` is the quadratic form ``` Var(u) = u^T C u ``` If `u` is a unit eigenvector with `Cu = lambda u`, then `u^T C u = u^T (lambda u) = lambda (u^T u) = lambda`. So **each eigenvalue is literally the variance of the data along its own eigenvector**. Maximizing `u^T C u` over unit vectors picks out the top eigenvector, with the largest eigenvalue as the maximum; minimizing picks the bottom one. The eigenvectors are the directions of extremal spread, and because `C` is symmetric they are mutually orthogonal. Drawing the cloud of a two-variable dataset as an ellipse (a contour of constant density for roughly elliptical data), the eigenvectors point along the ellipse's axes and the semi-axis lengths are proportional to `sqrt(lambda)` — square roots, because eigenvalues are in squared units of the data while the picture is drawn in the data's own units. ## A worked example Let ``` C = [[5, 3], [3, 5]] ``` Each variable has variance 5 and they covary positively. Using `lambda^2 - trace*lambda + det = 0` with `trace = 10` and `det = 25 - 9 = 16`: ``` lambda^2 - 10*lambda + 16 = (lambda - 8)(lambda - 2) -> lambda = 8, 2 ``` For `lambda = 8`: `C - 8I = [[-3,3],[3,-3]]` gives `v1 = v2`, so the top eigenvector is `(1,1)/sqrt(2)`. For `lambda = 2`: `C - 2I = [[3,3],[3,3]]` gives `v1 = -v2`, so the second is `(1,-1)/sqrt(2)`, orthogonal to the first as symmetry promises. Interpretation: the cloud is a tilted ellipse whose long axis runs at 45 degrees. Variance along that diagonal is 8; variance across it is 2. The ellipse is `sqrt(8/2) = 2` times longer than it is wide. Because both variables happen to have equal variance, the axes land exactly on the diagonals; unequal variances tilt them off 45 degrees. ## Trace, determinant and total variance Two invariants make the eigen-view easy to sanity check: - **Trace.** `8 + 2 = 10`, and the diagonal of `C` sums to `5 + 5 = 10`. Total variance is conserved: rotating to the eigenbasis moves variance between axes but never changes the total. This is why the eigenvalues are often quoted as a share of the trace. - **Determinant.** `8 * 2 = 16 = 25 - 9`. The determinant, sometimes called the generalized variance, is proportional to the squared area of the ellipse. A determinant near zero means a nearly flat, degenerate cloud. ## Decorrelation Write the eigen-decomposition as `C = Q D Q^T`. If the data (already centred) is rotated by `Q^T`, the covariance of the rotated coordinates is `Q^T C Q = D`, which is diagonal. So expressing the data in the eigenbasis makes the new coordinates **uncorrelated**, with variances equal to the eigenvalues. That is the whole geometric content of the eigendecomposition of a covariance matrix: there is always a rotation of the axes that removes all linear correlation, and the eigenvectors name that rotation. ## Reading edge cases - **Equal eigenvalues.** The ellipse is a circle, spread is the same in every direction, and the eigenvectors are not unique — any orthonormal pair works. Do not over-interpret which pair the computation returned. - **One eigenvalue near zero.** The cloud is nearly flat: some linear combination of the variables barely varies, which is the signature of near-collinearity among the variables. - **Sign.** An eigenvector and its negative describe the same axis. An arrow pointing 'up the diagonal' or 'down the diagonal' is the same finding, so never build logic on the returned sign. - **Scale dependence.** Eigenvectors of a covariance matrix are not invariant to rescaling the variables. Change one variable from metres to millimetres and its variance grows by a factor of a million, dragging the top axis toward it. Whenever variables have incomparable units, standardize first and eigendecompose the correlation matrix instead; the two decompositions generally disagree, and which one is appropriate is a modelling decision, not a numerical one.
- Why are the eigenvectors of any covariance matrix guaranteed to be orthogonal?Because a covariance matrix is symmetric, and the spectral theorem says a real symmetric matrix has real eigenvalues and an orthonormal eigenvector basis. When eigenvalues repeat, the eigenvectors are not unique, but an orthonormal set can always be chosen within that eigenspace. Orthogonality is a consequence of symmetry, not an assumption about the data.
- What does it mean if one eigenvalue of a covariance matrix is close to zero?It means the data barely varies along that eigenvector: some linear combination of the variables is almost constant, so the cloud is nearly flat in that direction. That is near-collinearity. The matrix is then close to singular, its inverse is numerically unstable, and any procedure that divides by that small variance amplifies noise.
- Would eigendecomposing the correlation matrix instead give the same axes?Generally no. Eigenvectors of a covariance matrix are not invariant under rescaling of the variables, so a variable measured in small units dominates the top axis purely through its numerical variance. Standardizing first and decomposing the correlation matrix gives every variable equal weight. Which is correct depends on whether the original units are comparable.
A covariance matrix is a fingerprint smudge on glass. The eigenvectors are the long and short axes of the smudge, and the eigenvalues say how far it stretches along each.
saying these in an interview costs you the question
- Says the eigenvector gives the mean or centre of the data
- Reports eigenvalues as standard deviations rather than variances
- Claims the axes are orthogonal because the variables are uncorrelated
- Believes eigenvectors are unchanged when variables are rescaled
- Reads meaning into the arbitrary sign of an eigenvector
- Over-interprets the axes when the two eigenvalues are nearly equal