What do the diagonal and off-diagonal entries of a sample covariance matrix hold?
answer
- square grid, one row per variable
- a variable paired with itself
- the (i, j) and (j, i) cells agree
- diagonal is variances, off-diagonal covariances
- divide by the diagonal's square roots for correlations
basics
~20 sThe diagonal holds each variable's sample variance, because a variable's covariance with itself is its variance. Off-diagonal entries hold the pairwise sample covariances, and the matrix is symmetric since covariance does not depend on pair order.
solid answer
~50 sFor p variables -- say daily returns for three assets -- the sample covariance matrix S is p by p, with `S[i][j] = cov(xi, xj)`. On the diagonal, `S[i][i] = cov(xi, xi) = var(xi)`, so the diagonal is the variances. Off the diagonal sit the pairwise covariances, and because `cov(x, y) = cov(y, x)` the matrix is symmetric: a 3 by 3 matrix has only six distinct numbers, three variances and three covariances. Every entry carries units, in this case squared return units on the diagonal and products of return units off it, so the raw numbers are hard to compare. Standardising -- dividing entry (i, j) by `sqrt(S[i][i]) * sqrt(S[j][j])` -- turns S into the correlation matrix, with ones on the diagonal and Pearson r values elsewhere. S is also positive semi-definite when computed on complete rows, which is what makes it usable downstream.
go deeper
Recall the layout: one row and one column per variable, variances on the diagonal, pairwise covariances off it, and the grid is symmetric. Know that the correlation matrix is the standardised cousin with ones on the diagonal.
Explain why the diagonal is variances (a variable's covariance with itself) and why symmetry follows from the cross-product. Be able to state how many distinct entries a p by p matrix has and how to rescale it into correlations.
Show that you have handled real matrices: units make raw entries incomparable, missing data handled pair by pair can produce an invalid matrix, and p larger than n - 1 leaves it rank-deficient. Explain what w' S w gives you.
Own the estimation policy when p is large relative to n: how many pairwise numbers the team is really estimating, whether the sample matrix is stable enough to build on, and what is reported to consumers -- covariances, correlations, or neither.
## The construction With p variables measured on the same n rows -- daily returns for three assets A, B and C over n trading days -- you can compute a covariance for every ordered pair. Collecting them into a p by p grid gives the **sample covariance matrix**: `S[i][j] = sum_t (x_ti - xbar_i) * (x_tj - xbar_j) / (n - 1)` for variables i and j, where `xbar_i` is the sample mean of variable i. For three assets, S is 3 by 3. ## What sits where **Diagonal.** `S[i][i]` is the covariance of a variable with itself, which is exactly its sample variance. So the diagonal of S for three assets is the variance of A's returns, the variance of B's, and the variance of C's. Taking square roots along the diagonal gives the standard deviations. **Off-diagonal.** `S[i][j]` for i not equal to j is the sample covariance between variables i and j -- the co-movement of A and B, A and C, and B and C. Its sign says whether the two assets tended to be above or below their own means on the same days. **Symmetry.** The cross-product `(a - abar) * (b - bbar)` does not care about order, so `cov(A, B) = cov(B, A)` and `S = S transposed`. A p by p covariance matrix therefore holds only `p * (p + 1) / 2` distinct numbers: for p = 3 that is six -- three variances plus three pairwise covariances. Reporting the upper triangle plus the diagonal loses nothing. ## Units, and why you usually standardise Every entry carries units. Diagonal entries are squared units of their variable; off-diagonal entries are the product of two variables' units. If the three columns were measured on wildly different scales -- one in percent, one in basis points, one in currency -- the raw magnitudes in S would say more about the units than about the association. The fix is the matrix version of Pearson's r. Let D be the diagonal matrix of sample standard deviations. Then `R[i][j] = S[i][j] / (sqrt(S[i][i]) * sqrt(S[j][j]))` is the **sample correlation matrix**: unitless, ones down the diagonal (each variable correlates perfectly with itself), and every off-diagonal entry a Pearson r bounded between -1 and 1. Correlation matrices are what people read; covariance matrices are what formulas consume. ## Positive semi-definiteness Computed the honest way -- one complete data matrix, all p columns from the same n rows -- S is positive semi-definite: for any weight vector w, `w' S w >= 0`. That is not an abstract nicety. `w' S w` is the sample variance of the weighted combination `w1*x1 + ... + wp*xp`, and a variance cannot be negative. For three assets, w is a portfolio's weights and `w' S w` is that portfolio's return variance, which is precisely why covariance matrices are assembled in the first place: they let you get the variance of any linear combination without recomputing anything. S can be singular (some `w' S w = 0`) when a variable is an exact linear combination of the others, or whenever p exceeds n - 1 -- more columns than rows leaves the matrix rank-deficient by construction. ## The practical trap: pairwise deletion If rows have missing values and you compute each entry from whatever rows happen to be complete for that pair, every entry comes from a different sample. The result can violate positive semi-definiteness -- you can get a "covariance matrix" implying a negative variance for some combination, or a correlation outside -1 to 1. Using complete rows only (or an estimator designed to return a valid matrix) keeps S coherent. ## Reading one Given a 3 by 3 S for assets, the questions to ask are: which asset has the largest diagonal entry (most volatile), which pair has the strongest co-movement once you standardise to correlations (raw covariances would mislead if the assets have different volatilities), and is any off-diagonal entry negative, since a negative covariance is what makes a combination less volatile than its parts. ## In an interview State the three facts crisply: diagonal is variances, off-diagonal is pairwise covariances, matrix is symmetric so p by p holds `p * (p + 1) / 2` distinct values. Then add the two things that separate a strong answer: standardising by the diagonal's square roots produces the correlation matrix, and `w' S w` is the variance of a weighted combination, which is why the matrix form exists at all.
- How do you convert a sample covariance matrix into a correlation matrix?Divide entry (i, j) by `sqrt(S[i][i]) * sqrt(S[j][j])`, the two variables' sample standard deviations. Equivalently, pre- and post-multiply S by the inverse of the diagonal matrix of standard deviations. The diagonal becomes all ones, every off-diagonal entry becomes a Pearson r in -1 to 1, and all units disappear, which is what makes the correlation matrix the readable form.
- How many distinct numbers does a covariance matrix for 10 variables contain?Fifty-five: `p * (p + 1) / 2` with p = 10, made up of 10 variances and 45 distinct pairwise covariances. Symmetry means the lower triangle duplicates the upper one, so storing or reporting the upper triangle plus the diagonal loses no information -- and it explains why estimating a covariance matrix gets expensive fast as p grows relative to n.
- What can go wrong if you compute each entry from whatever rows are complete for that pair?Each entry then comes from a different subsample, and the assembled matrix need not be positive semi-definite. That means some weighted combination of the variables would be assigned a negative variance, which is impossible, and a standardised entry can even fall outside -1 to 1. Restricting to rows complete on all variables keeps the matrix internally consistent.
saying these in an interview costs you the question
- Says the diagonal holds ones or the variable means
- Thinks the matrix is asymmetric with direction of effect
- Counts nine independent numbers in a 3 by 3 matrix
- Compares raw covariances across differently-scaled columns
- Assumes any symmetric matrix is a valid covariance matrix