For two equal-variance return streams, how does their correlation change the variance of a 50/50 blend?
answer
- the cross term is the whole story
- weights squared, cross term doubled
- 2*w1*w2*Cov(X,Y)
- equal weights give 0.5*s^2*(1 + rho)
- averaging bottoms out at rho*s^2
basics
~10 sWith both streams at variance s^2 and correlation rho, the blend has variance 0.5s^2(1 + rho). Perfectly correlated streams give no reduction, uncorrelated ones halve the variance, and rho = -1 cancels it entirely.
solid answer
~40 sThe general identity is `Var(w1*X + w2*Y) = w1^2*Var(X) + w2^2*Var(Y) + 2*w1*w2*Cov(X,Y)`. The cross term is where all the interesting behaviour lives. With `Var(X) = Var(Y) = s^2`, `Cov(X,Y) = rho*s^2` and equal weights of 0.5, this collapses to `Var = 0.5*s^2*(1 + rho)`. At `rho = 1` you get `s^2` — blending two identical risks buys nothing. At `rho = 0` you get `s^2 / 2`. At `rho = -1` you get exactly 0, a perfect hedge. That is the whole mathematical content of diversification: it is not the number of streams that reduces variance, it is their imperfect correlation. The same term runs the other way too — if you assume two positively correlated streams are independent, you drop `2*w1*w2*Cov(X,Y)` and understate the risk.
go deeper
Learn the identity Var(w1X + w2Y) = w1^2Var(X) + w2^2Var(Y) + 2w1w2*Cov(X,Y) by heart, and remember the weights are squared on the first two terms.
Be able to derive the identity from bilinearity of covariance and substitute Cov = rhos^2 to reach 0.5s^2*(1 + rho) for the equal-weight case, then read off the three landmark correlations.
Show the operational consequence: the floor at rho*s^2, and the fact that assuming independence among positively correlated components understates aggregate risk in exactly the direction that hurts.
Own the assumption itself. Decide how correlation inputs are sourced, reviewed and stress-tested, and be explicit that a diversification argument built on a stable rho fails when dependence spikes under stress.
## The identity For any two random variables with finite variance and any constants w1, w2: `Var(w1*X + w2*Y) = w1^2*Var(X) + w2^2*Var(Y) + 2*w1*w2*Cov(X,Y)` This follows from bilinearity of covariance, since `Var(Z) = Cov(Z,Z)`; expanding `Cov(w1*X + w2*Y, w1*X + w2*Y)` gives four terms, and the two cross terms are equal by symmetry, producing the factor 2. Note that the weights enter *squared* on the own-variance terms but as a plain product on the cross term. Note too that only the cross term can be negative. ## The equal-variance, equal-weight case Suppose two return streams each have variance `s^2` and correlation `rho`, so `Cov(X,Y) = rho * s^2`. With `w1 = w2 = 0.5`: `Var = 0.25*s^2 + 0.25*s^2 + 2*0.25*rho*s^2 = 0.5*s^2*(1 + rho)` Read off the three landmarks: - `rho = 1`: variance is `s^2`. Blending two copies of the same risk changes nothing. - `rho = 0`: variance is `s^2 / 2`, a halving; the standard deviation falls by a factor of `sqrt(2)`. - `rho = -1`: variance is 0. The two streams cancel exactly. So diversification is not a property of *counting* streams; it is a property of the covariance between them. Two hundred perfectly correlated streams are one stream. ## Extending to n streams For `n` streams, equally weighted at `1/n`, each with variance `s^2` and common pairwise correlation `rho`: `Var = (1/n^2) * (n*s^2 + n*(n-1)*rho*s^2) = s^2 * (1/n + (1 - 1/n)*rho)` As `n` grows this does **not** go to zero — it converges to `rho * s^2`. The first piece, `s^2/n`, is the part averaging can remove; the second is a floor set entirely by the shared covariance. This is the formal statement of the distinction between idiosyncratic and systematic risk, and it is also why an ensemble of highly correlated forecasters stops improving after a handful of members. (A side constraint worth knowing: a common pairwise correlation is only admissible for `rho >= -1/(n-1)` — you cannot have many things all strongly negatively correlated with each other.) ## Choosing the weights With unequal variances the equal-weight blend is not optimal. Minimising `Var(w*X + (1-w)*Y)` over w gives `w* = (Var(Y) - Cov(X,Y)) / (Var(X) + Var(Y) - 2*Cov(X,Y))` The denominator is `Var(X - Y)`, which is positive unless the two differ by a constant. The formula makes the intuition explicit: strongly negative covariance pushes the weights toward balance, while a stream that is both low-variance and weakly related to the other attracts weight. ## The same term, read as a hazard The cross term does not only help. Turn it around: - If you aggregate `n` positively correlated quantities but compute the variance as if they were independent, you omit `2 * sum over i<j of Cov(X_i, X_j)` and understate the true variance — potentially by a large multiple when n is big, since the omitted sum has `n(n-1)/2` terms against `n` retained ones. - The error is one-directional and predictable: positive covariance means you are optimistic about the stability of the aggregate, which is the dangerous direction for anything risk-related. - Negative covariance produces the reverse: you would overstate the variance of the aggregate and leave capacity on the table. ## Judgment points a senior answer should hit 1. **The floor.** Say out loud that averaging cannot beat `rho * s^2`, so effort spent adding more correlated streams has a hard ceiling on its payoff; effort spent lowering `rho` does not. 2. **Correlation is not stable.** The identity is exact given `rho`, but the input is a modelling assumption. Dependence that is mild in ordinary conditions and strong in stressed ones makes a diversification calculation flattering exactly when it matters most. 3. **Only uncorrelatedness is needed.** Variance additivity requires `Cov(X,Y) = 0`, not independence. Saying so shows you know which assumption is actually load-bearing. 4. **Units.** The blended variance is in squared units; report the standard deviation when comparing to a threshold, and remember that halving a variance only reduces the standard deviation by about 29 percent.
- What happens to the variance of an equal-weight average of n streams as n grows, if every pair has correlation rho?It equals `s^2 * (1/n + (1 - 1/n)*rho)` and converges to `rho * s^2`, not to zero. Averaging removes only the idiosyncratic part; the shared covariance is a floor. Practically, adding members to a highly correlated ensemble hits diminishing returns quickly, and the lever with real headroom is reducing the correlation rather than increasing the count.
- If you wrongly assume two positively correlated streams are independent, which way does your variance estimate err?You understate it. The true variance carries `+2*w1*w2*Cov(X,Y)`, and dropping a positive term leaves the figure too small — the optimistic and therefore dangerous direction for a risk number. With n correlated terms the omission grows like `n^2` against `n` retained terms, so the understatement compounds badly as the aggregate widens.
- For two assets with unequal variances, which weights minimise the blend's variance?`w* = (Var(Y) - Cov(X,Y)) / (Var(X) + Var(Y) - 2*Cov(X,Y))` on X, with `1 - w*` on Y. The denominator is `Var(X - Y)`, positive unless the two differ only by a constant. Low variance and low covariance with the other stream both attract weight; note the solution is unconstrained, so it may call for a negative weight.
Two rowers pulling in exactly the same rhythm amplify every wobble; slightly out of phase, their wobbles partly cancel and the boat runs steadier.
saying these in an interview costs you the question
- Forgets the factor of 2 on the covariance term
- Applies weights linearly to the own-variance terms instead of squared
- Claims more streams always drive variance toward zero
- Assumes independence when only zero covariance is needed
- Treats correlation as a fixed constant across market regimes