Why does the magnitude of Cov(X,Y) say little about the strength of association?
answer
- it is not standardised
- units of X times units of Y
- rescale X and the number rescales
- divide by both standard deviations
- Cauchy-Schwarz gives the [-1,1] bound
basics
~10 sCovariance carries the units of X times the units of Y, so rescaling either variable rescales it. Its sign shows the direction of linear association; only the correlation Cov(X,Y)/(sd(X)*sd(Y)), bounded in [-1,1], measures strength.
solid answer
~40 sCovariance is defined as `Cov(X,Y) = E[(X - E[X])(Y - E[Y])] = E[XY] - E[X]E[Y]`. It is bilinear, so `Cov(aX, bY) = ab*Cov(X,Y)`: measure a length in millimetres instead of metres and the covariance jumps by a factor of 1000 while the relationship is unchanged. That means a covariance of 40 is not "stronger" than one of 0.4 — the two may not even be comparable quantities. What survives rescaling by positive constants is the sign, which tells you whether above-average X tends to come with above-average Y. To talk about strength you standardise: the correlation `rho = Cov(X,Y)/(sd(X)*sd(Y))` is dimensionless and, by the Cauchy-Schwarz inequality, always lies in [-1,1], hitting the endpoints exactly when Y is an increasing or decreasing affine function of X with probability one.
go deeper
Be ready to state the definition and read the sign: positive means the two tend to be above or below their means together, negative means opposite. Know that correlation is the standardised version living in [-1,1].
Expect to derive Cov(aX+b, cY+d) = ac*Cov(X,Y) at the whiteboard and explain from it why the raw magnitude is uninterpretable. Also be able to state the Cauchy-Schwarz bound that puts correlation in [-1,1].
Show judgment about which object to use where: covariance for variance algebra and risk aggregation, correlation for reporting strength. Be able to say when a correlation of 0.9 is still operationally useless because both variances are negligible.
Own the reporting convention across teams: which dependence summary goes into a dashboard, what units are attached, and when a single scalar is the wrong abstraction for the dependence a decision actually rests on.
## The definition For two random variables X and Y on the same probability space, with finite variances, the covariance is `Cov(X,Y) = E[(X - E[X])(Y - E[Y])]` which expands to the computational form `Cov(X,Y) = E[XY] - E[X]E[Y]`. The quantity inside the expectation is a product of two signed deviations. When X and Y are on the same side of their respective means the product is positive; when they are on opposite sides it is negative. The covariance is the probability-weighted average of that product, so a positive value says the pair tends to move together and a negative value says it tends to move oppositely. Note that this is a **population** quantity: it is an average over the joint distribution, not something computed from a data file. Everything below is a statement about the distribution. ## Why the number itself is not interpretable Covariance is bilinear and shift-invariant: - `Cov(X, X) = Var(X)` - `Cov(X, Y) = Cov(Y, X)` - `Cov(aX + b, cY + d) = ac * Cov(X, Y)` - `Cov(X, Y + Z) = Cov(X, Y) + Cov(X, Z)` The third rule is the crux. Additive shifts do nothing, which is good, but multiplicative rescaling passes straight through. If X is a price in dollars and Y is a duration in seconds, `Cov(X,Y)` has units of dollar-seconds. Switch to cents and minutes and the number changes by a factor of 100/60 without any change in the underlying dependence. There is no scale on which "large" covariance is defined, and no way to compare the covariance of one pair of variables with the covariance of another pair measured in different units. A second limitation: a covariance close to zero does not mean the variables are unrelated. Covariance only detects the *linear* component of association. ## Standardising: the correlation coefficient Dividing by both standard deviations removes the units: `rho = Corr(X,Y) = Cov(X,Y) / (sd(X) * sd(Y))` Because `Corr(aX + b, cY + d) = sign(ac) * Corr(X,Y)`, correlation is invariant to any positive affine rescaling of either variable — exactly the invariance you want from a measure of strength. The Cauchy-Schwarz inequality applied to the centred variables gives `|Cov(X,Y)| <= sd(X) * sd(Y)` so `rho` always lies in [-1, 1]. Equality holds precisely when the centred variables are proportional, i.e. when `Y = a + bX` holds with probability one; `rho = +1` for `b > 0` and `rho = -1` for `b < 0`. So `|rho| = 1` is the signature of an exact straight-line relationship, and intermediate values measure how close the joint distribution sits to such a line. ## A worked example Roll two fair, independent dice. Let X be the first die and let S be the sum. Bilinearity does the work: `Cov(X, S) = Cov(X, X + Y) = Cov(X,X) + Cov(X,Y) = Var(X) + 0 = Var(X)` For a fair die, `E[X] = 3.5` and `E[X^2] = 91/6`, so `Var(X) = 91/6 - 12.25 = 35/12`, roughly 2.92. Is 2.92 a strong relationship? The raw number cannot say. Standardise: `Var(S) = 2 * 35/12 = 70/12`, so `rho = (35/12) / (sqrt(35/12) * sqrt(70/12)) = sqrt(35/70) = 1/sqrt(2) ~= 0.707` Now the answer is interpretable: about 0.71, a strong but not deterministic linear association, and it would stay 0.71 if you relabelled the die faces as 10, 20, ..., 60. ## What the sign still buys you Even unstandardised, covariance is the right object in one important place: it is the term that appears in the variance of a sum, `Var(X + Y) = Var(X) + Var(Y) + 2*Cov(X,Y)`. There the units are correct by construction, because the result has the units of X squared. So the rule of thumb is: use covariance when you are doing algebra with variances, and use correlation when you are reporting or comparing strength. ## Common traps Treating covariance as bounded (it is not), comparing covariances across differently scaled variables, and reading a large covariance as a strong relationship when one of the variables simply has a huge variance. Conversely, a tiny covariance between two variables with tiny standard deviations can still correspond to `rho` near 1.
- Roll two fair independent dice, let X be the first die and S the sum. What is Cov(X, S)?Use bilinearity: `Cov(X, X + Y) = Cov(X,X) + Cov(X,Y) = Var(X) + 0`, since the dice are independent. `Var(X) = 91/6 - 3.5^2 = 35/12`, about 2.92. Standardising gives `rho = sqrt(35/12)/sqrt(70/12) = 1/sqrt(2)`, about 0.71 — a strong but not deterministic linear link, which matches the intuition that the first die is half the sum.
- What is Cov(aX + b, cY + d) in terms of Cov(X,Y)?It equals `ac * Cov(X,Y)`. The additive constants b and d shift both variables without changing their deviations from their means, so they vanish; the multiplicative constants pull straight out by bilinearity. This is exactly why covariance is scale-dependent and correlation, which divides by `|a| * |c|` through the standard deviations, is not.
- When can the correlation between two random variables actually equal 1?Only when `Y = a + bX` with `b > 0` holds with probability one — an exact increasing straight line, no scatter at all. This is the equality case of the Cauchy-Schwarz inequality applied to the centred variables. Any genuine noise around the line strictly reduces `|rho|` below 1, and `rho = -1` is the same statement with a negative slope.
Covariance is like reporting a distance without saying whether the unit is inches or kilometres; correlation is the same journey expressed as a fraction of the whole trip.
saying these in an interview costs you the question
- Says a covariance of 50 means a strong relationship
- Believes covariance is bounded between -1 and 1
- Compares covariances of variables measured in different units
- Confuses covariance with correlation and uses the names interchangeably
- Thinks adding a constant to X changes Cov(X,Y)