In a bivariate normal distribution, what does conditioning on X = x do to the distribution of Y?
answer
- conditioning stays in the family
- the mean shifts linearly with x
- the spread is free of x
- shrink by a factor of rho
- residual variance is sd_Y^2*(1 - rho^2)
basics
~20 sY given X = x is again normal, with mean mu_Y + rho*(sd_Y/sd_X)(x - mu_X) and variance sd_Y^2(1 - rho^2). The mean moves linearly with x, and the spread does not depend on x at all.
solid answer
~50 sConditioning keeps you inside the family: `Y | X = x` is normal, with `E[Y | X = x] = mu_Y + rho*(sd_Y/sd_X)*(x - mu_X)` and `Var(Y | X = x) = sd_Y^2*(1 - rho^2)`. Two features are worth naming. First, the conditional mean is **linear** in x, with the factor `rho` pulling the prediction only part of the way toward the extremity of x — a one-standard-deviation move in X shifts the expected Y by `rho` standard deviations. Second, the conditional variance is **constant** in x: the spread around that line is the same whether you condition on a typical x or an extreme one, and it is strictly smaller than `sd_Y^2` whenever `rho` is nonzero. Setting `rho = 0` makes the conditional mean `mu_Y` and the conditional variance `sd_Y^2` — the conditional equals the marginal for every x, which is exactly independence.
go deeper
Know that in a bivariate normal the conditional distribution of Y given X is still normal, and that its mean moves along a straight line as the conditioning value changes.
Be able to write both conditional formulas, mean mu_Y + rho*(sd_Y/sd_X)(x - mu_X) and variance sd_Y^2(1 - rho^2), and to explain that the variance term does not involve x.
Show that you use this as the population model behind linear prediction and regression to the mean, and that you can produce a counterexample with normal marginals that is not jointly normal.
Own the assumption's limits: elliptical joints have no tail dependence, so a diversification or risk argument resting on bivariate normality will look safest exactly when joint extremes are the actual concern.
## The bivariate normal A pair (X, Y) is **jointly normal** (bivariate normal) when every linear combination `aX + bY` is univariate normal. It is parameterised by five numbers: the two means `mu_X, mu_Y`, the two standard deviations `sd_X, sd_Y`, and the correlation `rho` in (-1, 1). The level sets of its density are **ellipses** centred at `(mu_X, mu_Y)`, tilted upward when `rho > 0`, downward when `rho < 0`, and axis-aligned when `rho = 0`. The larger `|rho|`, the more the ellipses flatten into a cigar around a line; as `|rho|` approaches 1 the distribution degenerates onto that line. ## The conditional distribution The key structural fact: `Y | X = x ~ Normal( mu_Y + rho*(sd_Y/sd_X)*(x - mu_X), sd_Y^2 * (1 - rho^2) )` Three things to unpack. **It stays normal.** Slicing the joint density along a vertical line and renormalising gives another normal density. The family is closed under conditioning, which is a large part of why it is so tractable. **The conditional mean is linear in x.** Write it in standardised form: if `z = (x - mu_X)/sd_X` is how many standard deviations x sits above its mean, then the conditional mean of Y sits `rho * z` standard deviations above `mu_Y`. Because `|rho| < 1`, the prediction is always pulled *less far* from the centre than the conditioning value was — the shrinkage is proportional to `rho`. At `rho = 0.6`, a two-standard-deviation X corresponds to an expected Y only 1.2 standard deviations out. **The conditional variance does not depend on x.** `sd_Y^2 * (1 - rho^2)` is the same number for every conditioning value, so the scatter around the line is homoscedastic by construction. It is strictly less than the marginal variance `sd_Y^2` whenever `rho` is nonzero: learning X removes a fraction `rho^2` of the variance of Y and leaves the fraction `1 - rho^2`. At `rho = 0.8`, conditioning removes 64 percent of the variance; the residual standard deviation is `sd_Y * sqrt(1 - 0.64) = 0.6 * sd_Y`. ## Why zero correlation implies independence here Set `rho = 0` in the conditional above: the mean becomes `mu_Y` and the variance becomes `sd_Y^2`, i.e. the conditional distribution of Y is identical to its marginal for **every** x. A conditional that never changes with the conditioning value is precisely independence. Equivalently, the joint density factorises into the two marginal normal densities when `rho = 0`. This is a property of the **joint** family, not of the marginals. It is entirely possible for both marginals to be normal, for the covariance to be zero, and for the variables to be dependent — in which case the pair simply is not bivariate normal. Concretely: let `X ~ N(0,1)`, let S take the values +1 and -1 with probability 1/2 each, independently of X, and set `Y = S*X`. Then Y is also `N(0,1)`, and `Cov(X,Y) = E[S*X^2] = E[S]*E[X^2] = 0`. But `|X| = |Y|` always, so the two are strongly dependent. The pair fails joint normality, visible in the fact that `X + Y` equals either `0` or `2X` and so is not normal. ## The elliptical picture Everything above is readable off the contour plot. Draw an ellipse tilted upward. Slice it vertically at `x`: the chord you get is centred on the line `mu_Y + rho*(sd_Y/sd_X)*(x - mu_X)`, and — this is the part people find surprising — the *shape* of the density along that chord is the same everywhere, only its centre moves. That line of conditional means is not the major axis of the ellipse; it is shallower, which is the geometric statement of the shrinkage by `rho`. ## What this buys you in an interview 1. **A clean mental model of linear prediction.** The optimal predictor of Y from X under joint normality is exactly this linear conditional mean, and the leftover variance is `sd_Y^2*(1 - rho^2)`. This is where the interpretation of `rho^2` as the fraction of variance explained comes from at the population level. 2. **The regression-to-the-mean phenomenon.** Because `|rho| < 1`, extreme X values are paired with less extreme expected Y values. This is a mathematical necessity, not a causal effect, and it explains a large class of apparent "the best performers decline" observations. 3. **A precise statement of the exception.** When someone claims uncorrelated implies independent, the correct reply names the family: true under joint normality, false in general, and normal marginals alone are not sufficient. ## Cautions Joint normality is an assumption, and a strong one. Real bivariate data often has heavier joint tails than an ellipse allows: two variables can look mildly correlated in the bulk and move together violently in the extremes. Nothing in the bivariate normal can represent that — its tail dependence is zero for any `rho < 1`. Checking each variable's marginal for normality does not verify the joint assumption, precisely because normal marginals do not imply joint normality.
- With rho = 0.6, how far does the expected Y move when X is two standard deviations above its mean?By `rho * 2 = 1.2` standard deviations of Y. In standardised terms the conditional mean is `rho * z`, so the prediction is always pulled less far from the centre than the conditioning value. That shrinkage is regression to the mean: it follows from `|rho| < 1` alone and needs no causal story about why extreme cases become less extreme.
- Does normality of each marginal guarantee that the pair is bivariate normal?No. Take `X ~ N(0,1)` and `Y = S*X` where S is plus or minus one with equal probability, independent of X. Y is also `N(0,1)` and `Cov(X,Y) = 0`, yet `|X| = |Y|` so they are dependent. The pair is not jointly normal — `X + Y` is not normal — which is exactly why the zero-correlation-implies-independence result does not apply.
- Why is the conditional variance smaller than the marginal variance of Y?Because conditioning on X supplies information. The conditional variance is `sd_Y^2*(1 - rho^2)`, so a fraction `rho^2` of Y's variance is accounted for by knowing X and the fraction `1 - rho^2` remains. At `rho = 0` nothing is removed and the conditional variance equals the marginal; as `|rho|` approaches 1 the residual variance collapses toward zero.
- Is the line of conditional means the same as the major axis of the density ellipse?No — it is shallower. The major axis is the direction of greatest spread of the joint cloud, while the line of conditional means traces the centre of each vertical slice. Only when the two standard deviations are equal and `|rho|` approaches 1 do the two lines coincide; in general the conditional-mean line is flatter by exactly the shrinkage factor `rho`.
Slicing a tilted ellipse vertically always gives the same-shaped chord; only its centre slides up or down as you move the slice.
saying these in an interview costs you the question
- Says the conditional variance grows for extreme x values
- Claims normal marginals imply a bivariate normal joint
- Predicts Y as far from its mean as x is from its own
- Generalises uncorrelated-implies-independent beyond joint normality
- Confuses the conditional-mean line with the ellipse's major axis