skip to content

Why does Cov(X,Y) = 0 fail to guarantee that X and Y are independent?

level: middleimportance: must knowfreq 72%

answer

  1. it only sees straight lines
  2. symmetry can cancel the products
  3. think of a U-shaped relationship
  4. Y = X^2 on a symmetric X
  5. independence needs the density to factorise

basics

~20 s

Covariance detects only linear association. A pair can be perfectly dependent through a curved relationship and still have zero covariance: with X uniform on (-1,1) and Y = X^2, Cov(X,Y) = 0 even though Y is a function of X.

solid answer

~40 s

Zero covariance says the best straight-line summary of the relationship has zero slope; independence says the joint distribution factorises, `f(x,y) = f_X(x) * f_Y(y)` at every point. The second is much stronger. The standard counterexample: let X be uniform on (-1,1) and set `Y = X^2`. Then `E[X] = 0` and `E[XY] = E[X^3] = 0` by symmetry, so `Cov(X,Y) = E[XY] - E[X]E[Y] = 0`, yet knowing X pins Y exactly — you cannot get more dependent than that. The implication runs one way only: independence does imply zero covariance (when the variances are finite), never the reverse. The one important family where the reverse holds is the bivariate normal, where zero correlation really does mean independence — which is why the misconception survives.

go deeper

for a junior

Memorise the direction of the arrow: independence implies zero covariance, never the reverse. Have one counterexample ready, such as a symmetric variable and its square.

for a middle

Be able to compute the counterexample live: show E[X] = 0 and E[X^3] = 0 by odd symmetry, then conclude the covariance is zero while Y is a function of X.

for a senior

Demonstrate that you check for non-monotone structure before declaring a variable uninformative, and explain which downstream results need only uncorrelatedness versus genuine independence.

for a principal

Own the framing when a team reports 'no relationship found' from a dependence scalar alone: decide what evidence standard is required before a variable is dropped, and where the cost of a missed non-linear dependence lands.

## Two different statements **Zero covariance (uncorrelated):** `E[XY] = E[X]E[Y]`. This is a single scalar equation. It says the centred variables have no linear co-movement on average. **Independence:** the joint distribution factorises everywhere. For discrete variables, `P(X = x, Y = y) = P(X = x) * P(Y = y)` for every pair (x, y); for continuous ones, `f(x,y) = f_X(x) * f_Y(y)` for (almost) every (x, y). Equivalently, the conditional distribution of Y given X = x is the same for every x. That is an infinite family of equations. Independence implies zero covariance whenever the variances are finite, because independence gives `E[XY] = E[X]E[Y]` directly. One scalar equation cannot possibly imply the whole factorisation, so the converse fails. ## The canonical counterexample Let X be uniform on the interval (-1, 1) and define `Y = X^2`. - `E[X] = 0` by symmetry. - `E[XY] = E[X * X^2] = E[X^3]`. The density is symmetric about 0 and `x^3` is an odd function, so this integral is 0. - Therefore `Cov(X,Y) = E[XY] - E[X]E[Y] = 0 - 0 * E[Y] = 0`, and the correlation is 0 too. But Y is a deterministic function of X. Knowing `X = 0.9` tells you `Y = 0.81` with certainty; knowing `X = 0` tells you `Y = 0`. The conditional distribution of Y changes completely with x, so independence fails as badly as it can. The geometry explains it. The relationship is a parabola. For every positive x that pushes Y up, there is a mirror-image negative x that pushes Y up by the same amount. The products `(x - E[X])(y - E[Y])` come in pairs of opposite sign that cancel exactly. Covariance measures the average tilt of the cloud, and a symmetric parabola has no tilt. ## Mean dependence sits in between A useful middle rung is **mean independence**: `E[Y | X = x]` does not depend on x. Mean independence implies zero covariance and is implied by independence, but neither converse holds. In the parabola example, `E[Y | X = x] = x^2` is very much a function of x, so the pair is not even mean independent — yet the covariance still vanishes because covariance only picks up the linear part of `E[Y | X = x]`. That is the sharpest way to state the limitation: covariance sees the linear projection of the conditional mean and nothing else. ## Where the converse does hold If (X, Y) is **jointly normal** (bivariate normal), then zero correlation does imply independence. In that family the joint density factorises exactly when the correlation parameter is zero. This is a property of the family, not of normality of each variable separately. A pair can have normal marginals, zero covariance, and still be dependent: take `X ~ N(0,1)`, let S be +1 or -1 with equal probability independently of X, and set `Y = S * X`. Then Y is also `N(0,1)`, `Cov(X,Y) = E[S]*E[X^2] = 0`, but `|X| = |Y|` always — plainly dependent, and the pair is not bivariate normal. ## Why the distinction matters in practice 1. **Variance algebra.** `Var(X + Y) = Var(X) + Var(Y) + 2*Cov(X,Y)`, so the additivity of variances needs only zero covariance, not full independence. Uncorrelatedness is the weaker assumption that actually does the work here — a good thing to know, because it is easier to justify. 2. **Screening for relationships.** If your only dependence check is a covariance or correlation, symmetric non-monotone structure is invisible to you. A U-shaped relationship between a feature and an outcome will read as "no relationship" while being highly predictive. 3. **Assumption strength.** Many results (laws of large numbers, variance formulas) need only uncorrelatedness; others (factorising densities, independence of functions g(X) and h(Y)) genuinely need independence. Under independence, `g(X)` and `h(Y)` are uncorrelated for *all* well-behaved g and h — that is one clean way to characterise the gap. ## How to answer in the room State the direction of the implication, give the parabola in one line with the odd-function argument, and close with the bivariate normal as the exception that explains why so many people believe the false converse.

  • Does the implication run the other way — does independence guarantee zero covariance?
    Yes, provided both variances are finite. Independence gives `E[XY] = E[X]E[Y]` directly, so `Cov(X,Y) = 0`. The finiteness caveat matters: for heavy-tailed variables with undefined second moments the covariance may not exist at all, in which case the statement is vacuous rather than false.
  • Is there a joint distribution where zero correlation does imply independence?
    Yes — the bivariate normal. If (X, Y) is jointly normal with correlation zero, the joint density factorises into the two marginal densities, so they are independent. This is a property of the joint family, not of the marginals: normal marginals alone are not enough, and outside this elliptical family the implication fails.
  • Which is the weaker assumption behind Var(X + Y) = Var(X) + Var(Y)?
    Uncorrelatedness. The identity is `Var(X + Y) = Var(X) + Var(Y) + 2*Cov(X,Y)`, so all that is needed is a vanishing covariance. Independence is sufficient but strictly stronger than required, which is worth saying out loud whenever someone claims a variance-addition step forces an independence assumption.

Covariance is a metal detector: it finds one kind of buried object very well and walks straight over everything made of plastic.

saying these in an interview costs you the question

  • Claims zero covariance means the variables are unrelated
  • Uses uncorrelated and independent as synonyms
  • Cannot produce any counterexample when pushed
  • Thinks the counterexample requires an exotic distribution
  • Says independence does not imply zero covariance

context