skip to content

How do you derive Var(X) = E[X^2] - (E[X])^2 from the definition of variance?

level: middleimportance: should knowfreq 62%

answer

  1. expand the square inside the expectation
  2. the mean is a constant, pull it out
  3. the cross term collapses
  4. variance never negative, so a square dominates

basics

~10 s

Expand the squared deviation inside the expectation. Var(X) = E[(X - mu)^2] becomes E[X^2] - 2muE[X] + mu^2, and since E[X] = mu the last two terms collapse to -mu^2, leaving E[X^2] - (E[X])^2.

solid answer

~40 s

Variance is defined as the mean squared deviation from the mean: `Var(X) = E[(X - mu)^2]` where `mu = E[X]`. Expand the square inside the expectation to get `E[X^2 - 2*mu*X + mu^2]`. Now apply linearity, remembering that mu is a constant, not a random variable: `E[X^2] - 2*mu*E[X] + mu^2`. Substituting `E[X] = mu` gives `E[X^2] - 2*mu^2 + mu^2 = E[X^2] - mu^2`. Two consequences worth stating. First, because variance is an expectation of a square it is never negative, so `E[X^2] >= (E[X])^2` for every variable with finite second moment, with equality only when X is constant. Second, the identity says variance depends on the first two moments alone, which is why the second moment is the piece you track alongside the mean.

go deeper

for a junior

Recall both forms and that variance is the average squared distance from the mean, reported in squared units while the standard deviation is in the original units.

for a middle

Do the expansion live, justify each step by linearity, and explain why the mean can be pulled out as a constant rather than factorised as a product.

for a senior

Bring the consequences: the ordering of the second moment and the squared mean, the indicator shortcut, and why the identity is not what you evaluate in floating point.

for a principal

Own when a two-moment summary is enough at all: for skewed or heavy-tailed quantities, mean and variance can be a misleading pair to report and to build decisions on.

## The two expressions Variance measures spread around the mean. Its definition is `Var(X) = E[(X - mu)^2]`, where `mu = E[X]` Read literally: take the deviation of X from its mean, square it so that positive and negative deviations both count as spread, and average those squared deviations with probability weights. The computational form is `Var(X) = E[X^2] - (E[X])^2` sometimes phrased as "the mean of the square minus the square of the mean". ## The derivation Start from the definition and expand the square inside the expectation: `(X - mu)^2 = X^2 - 2*mu*X + mu^2` So `Var(X) = E[X^2 - 2*mu*X + mu^2]`. Apply linearity of expectation, which lets you split a sum and pull out constant multipliers. The critical observation is that `mu` is a fixed number, not a random variable: it has already been averaged. Therefore `E[X^2 - 2*mu*X + mu^2] = E[X^2] - 2*mu*E[X] + mu^2` Now substitute `E[X] = mu`: `= E[X^2] - 2*mu^2 + mu^2 = E[X^2] - mu^2` which is `E[X^2] - (E[X])^2`. The whole derivation is linearity plus the fact that the expectation of a constant is that constant. The single most common error is treating mu as random and writing `E[2*mu*X] = 2*E[mu]*E[X]` as though a product rule were being applied. No product rule is needed; mu comes out as a constant factor. ## What the identity tells you **Variance is non-negative, so the second moment dominates the squared mean.** Since `(X - mu)^2 >= 0` pointwise, its expectation is at least 0, hence `E[X^2] >= (E[X])^2` for any variable with a finite second moment. Equality holds only when the squared deviation is zero with probability 1, that is, when X is a constant. This is a fact worth having on hand: interviewers ask it as a standalone true-or-false. **Only two moments matter.** Everything about the spread of X in this sense is fixed once you know `E[X]` and `E[X^2]`. That is why moment bookkeeping is usually kept as a running pair. **Units are squared.** If X is measured in seconds, `Var(X)` is in seconds squared and is not directly comparable to the mean. The standard deviation `sd(X) = sqrt(Var(X))` restores the original units, which is why spread is usually reported as a standard deviation even when the algebra is done on variances. ## A worked check Roll a fair six-sided die. `E[X] = (1+2+3+4+5+6)/6 = 3.5`. The second moment is `E[X^2] = (1+4+9+16+25+36)/6 = 91/6 = 15.1667`. So `Var(X) = 15.1667 - 3.5^2 = 15.1667 - 12.25 = 2.9167` which is `35/12`. Computing it the long way from `E[(X - 3.5)^2]` gives the same number: the deviations are plus or minus 2.5, 1.5 and 0.5, so the average of `6.25, 2.25, 0.25` each appearing twice is `(6.25 + 2.25 + 0.25)/3 = 2.9167`. The two routes agree, as they must. ## Why the identity can still be a bad way to compute Algebraically the two forms are identical. In floating-point arithmetic they are not. When the mean is large relative to the spread, `E[X^2]` and `(E[X])^2` are two big, nearly equal numbers, and subtracting them cancels most of the significant digits, a phenomenon called catastrophic cancellation. Values clustered near a million with a spread of a few units make the difference of two numbers near 10^12, and the leading digits agree, so the small true difference can be swamped by rounding error and can even come out negative, which is impossible for a variance. The deviation form, or a shift-then-square approach that first subtracts a rough centre, is numerically stable. So the identity is the right tool for algebra and derivations, and the wrong default for numerical evaluation. ## Where it gets used The identity is the workhorse behind most variance derivations. To get the variance of a Bernoulli variable that is 1 with probability p, note that `X^2 = X` because 1 and 0 are their own squares, so `E[X^2] = p` and `Var(X) = p - p^2 = p(1-p)`. That trick, squaring an indicator leaves it unchanged, is one of the cleanest applications and shows up constantly. Similarly, knowing the mean and variance of a distribution immediately gives the second moment as `E[X^2] = Var(X) + (E[X])^2`, which is the direction you use when a problem hands you a mean and a spread and asks for an expected square.

  • What does the non-negativity of variance imply about E[X^2] and (E[X])^2?
    It forces `E[X^2] >= (E[X])^2` for every variable with a finite second moment, since the difference is the expectation of a non-negative quantity. Equality holds only when the squared deviation is zero with probability 1, that is when X is constant. If a calculation ever yields a second moment below the squared mean, there is an arithmetic error.
  • Use the identity to get the variance of an indicator variable that equals 1 with probability p.
    Because the values are 0 and 1, squaring changes nothing, so `X^2 = X` and `E[X^2] = E[X] = p`. Then `Var(X) = p - p^2 = p(1-p)`. It is maximal at p = 0.5 and shrinks to zero as p approaches 0 or 1, which matches the intuition that a near-certain event carries little spread.
  • Why is this identity a poor way to evaluate variance numerically?
    When the mean is large relative to the spread, the mean of the square and the square of the mean are two big, nearly equal numbers, so subtracting them destroys significant digits. That catastrophic cancellation can even produce a negative result. The identity stays correct algebraically; for evaluation, work from deviations around a centre instead.

saying these in an interview costs you the question

  • Treats the mean as a random variable inside the expansion
  • Writes variance as E[X]^2 minus E[X^2], reversing the order
  • Claims E[X^2] equals (E[X])^2 for symmetric distributions
  • Forgets that variance carries squared units
  • Reports a negative variance without spotting the arithmetic error

context