skip to content

Why does linearity of expectation hold even when the random variables are dependent?

level: middleimportance: must knowfreq 76%

answer

  1. sums versus products
  2. the proof only marginalises
  3. indicators need not be independent
  4. expectation of an indicator is a probability

basics

~20 s

Linearity of expectation is proved by regrouping a sum over the joint distribution, never by multiplying probabilities. E[X + Y] = E[X] + E[Y] holds for any variables with finite means. Independence is needed for products, not sums.

solid answer

~50 s

Expectation is a sum weighted by the joint distribution, and `E[X + Y] = sum over (x,y) of (x + y) * P(X = x, Y = y)` splits into two sums that marginalise back to `E[X]` and `E[Y]`. Nothing in that regrouping factorises a joint probability, so no independence assumption appears; all you need is that both means exist. Independence is required for a different identity, `E[XY] = E[X]E[Y]`, and for the variance of a sum to be additive. The practical payoff is that you can decompose a horrible variable into simple, wildly dependent pieces and add their means. In the hat-check problem, n hats are handed back uniformly at random; let `I_i` be 1 if person i gets their own hat. The indicators are heavily dependent, but each has mean 1/n, so the expected number of matches is `n * (1/n) = 1` for every n.

go deeper

for a junior

Recall that expectations of a sum always add, and that the expectation of an indicator variable is simply the probability of the event it marks.

for a middle

Be able to prove it in three lines by splitting the joint sum and marginalising, and to say precisely which neighbouring identity does need independence.

for a senior

Demonstrate the modelling move: decompose a messy count into indicators and add probabilities rather than chasing the distribution of the total.

for a principal

Own the boundary in review: know that linearity needs integrability, not independence, and push back when a colleague adds variances of clearly co-moving quantities.

## The statement For random variables X and Y on the same probability space, both with finite expectation, and for constants a, b, c: `E[aX + bY + c] = a*E[X] + b*E[Y] + c` No independence, no identical distributions, no assumption about the shape of anything. This is what people mean by linearity of expectation, and it extends by induction to any finite collection: `E[sum of X_i] = sum of E[X_i]`. ## Why dependence is irrelevant Write the discrete case out. Expectation of a function of the pair (X, Y) is computed against the joint distribution: `E[X + Y] = sum over x sum over y of (x + y) * P(X = x, Y = y)` Split the bracket into two sums: - `sum over x sum over y of x * P(X = x, Y = y) = sum over x of x * P(X = x) = E[X]`, because summing the joint probability over all y collapses it to the marginal of X. - Symmetrically the second piece gives `E[Y]`. The only fact used is that the joint probabilities marginalise correctly, which is true by definition. At no point does the argument write `P(X = x, Y = y)` as a product of marginals, and that factorisation is precisely what independence would buy you. The continuous case is the same argument with integrals. So dependence simply never enters. Contrast the product. `E[XY] = sum over x sum over y of x*y*P(X = x, Y = y)` cannot be split into separate sums over x and over y unless the joint probability factorises. That is why `E[XY] = E[X]E[Y]` genuinely requires independence (or at least the weaker condition that the two variables are uncorrelated). ## The technique this unlocks Linearity is the single most useful counting tool in probability interviews, because it lets you avoid the distribution of the sum entirely. The pattern: 1. Write the quantity you care about as a sum of simple pieces, usually indicator variables. An indicator `I_A` equals 1 when event A happens and 0 otherwise. 2. Note that `E[I_A] = 1*P(A) + 0*(1 - P(A)) = P(A)`. The expectation of an indicator is just the probability of its event. 3. Add the probabilities. You never need to know how the pieces interact. ## Worked example: the hat-check problem n people check their hats at a party and the hats are returned uniformly at random, one per person. How many people expect to get their own hat back? The distribution of the number of matches is awkward: it involves derangement counts, and the indicators are strongly dependent (if the first n-1 people all have their own hat, the last one certainly does too). Linearity does not care. Let `I_i = 1` if person i receives their own hat. By symmetry person i is equally likely to receive any of the n hats, so `P(I_i = 1) = 1/n` and `E[I_i] = 1/n`. The number of matches is `M = I_1 + ... + I_n`, so `E[M] = n * (1/n) = 1` The expected number of matches is exactly 1, for every n from 1 upward. Ten people or ten million, the answer does not move. Trying to reach that result through the distribution of M is a much longer road; linearity gets there in two lines precisely because it ignores the dependence. ## Where independence does matter Three neighbouring facts are easy to conflate with linearity, and interviewers probe the boundary: - `E[XY] = E[X]E[Y]` needs the variables to be uncorrelated, which independence implies. - Additivity of variance for a sum needs uncorrelated variables; for dependent variables the variance of a sum is not the sum of variances, and can be larger or smaller. - `E[g(X)] = g(E[X])` is false for any non-affine g, independence or not. Expectation passes through sums and constant multiples, not through squares, reciprocals, logs or maxima. This is a separate trap from the independence one and catches at least as many candidates. ## Small print worth knowing Linearity requires the individual expectations to exist. For finitely many variables with finite means there is nothing to check. For infinite sums or for heavy-tailed variables where `E|X|` is infinite, the interchange of sum and expectation needs a convergence condition, and pathological counterexamples exist. Interviewers rarely go there, but knowing that the theorem has a hypothesis, and that the hypothesis is integrability rather than independence, is a good signal.

  • Where does independence actually become necessary?
    For the expectation of a product, `E[XY] = E[X]E[Y]`, whose proof needs the joint probability to factorise, and for the variance of a sum to equal the sum of variances, which needs the variables to be uncorrelated. Independence implies uncorrelated but is strictly stronger. Sums of expectations need neither.
  • What is the expected number of people who get their own hat back when n hats are returned at random?
    Exactly 1, for every n. Each person has probability 1/n of receiving their own hat, so the indicator for each has mean 1/n, and summing n of them gives 1. The indicators are strongly dependent, which is exactly why linearity is the right tool: it never asks how they interact.
  • Does E[g(X)] equal g(E[X]) for a general function g?
    No. Expectation commutes only with affine functions, so `E[aX + b] = a*E[X] + b` holds, but `E[X^2]`, `E[1/X]` and `E[max(X, 0)]` all differ from the same function applied to the mean. This is a different failure from the independence trap and is the more common one in practice.

saying these in an interview costs you the question

  • Claims E[X + Y] = E[X] + E[Y] requires independence
  • Confuses linearity with E[XY] = E[X]E[Y]
  • Thinks the variables must be identically distributed
  • Tries to derive the distribution of the sum before taking the mean
  • Applies linearity through a square or a reciprocal

context