skip to content

How is the chi-square distribution built from standard normal random variables?

level: seniorimportance: nice to knowfreq 28%

answer

  1. start from independent standard normals
  2. squares, and squares cannot be negative
  3. count how many are being added
  4. each squared standard normal has mean 1
  5. mean k, variance 2k

basics

~20 s

It is a sum of squares: if Z1 through Zk are independent standard normals, then Z1^2 + ... + Zk^2 is chi-square with k degrees of freedom, never negative, with mean k and variance 2k.

solid answer

~50 s

Square independent standard normals and add them. With `Z1, ..., Zk` independent standard normal, `V = Z1^2 + ... + Zk^2` is chi-square with `k` degrees of freedom. The construction explains its properties: it cannot be negative, it is right-skewed for small `k`, its mean is `k` because each squared standard normal has expectation 1, and its variance is `2k`. Degrees of freedom add when independent chi-squares are added, since concatenating two lists of squared normals makes a longer list. Two other distributions are assembled from these pieces: `T = Z / sqrt(V/k)`, with `Z` standard normal independent of a chi-square `V` on `k` degrees of freedom, is a t on `k` degrees of freedom; and `(V1/d1) / (V2/d2)`, a ratio of independent chi-squares each divided by its own degrees of freedom, is an F with `d1` and `d2`.

go deeper

for a junior

Recall the one-line definition: add up the squares of k independent standard normal variables and you have a chi-square with k degrees of freedom, a quantity that is never negative.

for a middle

Be ready to derive mean k and variance 2k from E[Z^2] = 1 and Var(Z^2) = 2, and to explain why degrees of freedom add when two independent chi-squares are added.

for a senior

Expect to explain the t as a normal divided by an independent noisy scale, and the F as a ratio of two such scales, and to say what those constructions imply about tail weight and limiting behaviour.

for a principal

Own the framing. Be able to explain to a mixed audience that these three distributions are one normal-based construction seen three ways, so that degrees of freedom stop being a lookup parameter and become a count of remaining free directions.

## The definition Let `Z1, ..., Zk` be independent standard normal variables — mean 0, variance 1. Define `V = Z1^2 + Z2^2 + ... + Zk^2`. The distribution of `V` is called chi-square with `k` degrees of freedom. That is the definition; everything else about the distribution follows from it, which is why it is worth memorising the construction rather than the density formula. ## What the construction immediately tells you **Support.** A sum of squares cannot be negative, so `V` lives on `[0, infinity)`. Any answer that allows a negative chi-square value is self-refuting. **Mean.** For a standard normal, `E[Z^2] = Var(Z) + (E[Z])^2 = 1 + 0 = 1`. Summing `k` of them gives `E[V] = k`. The mean of a chi-square is its degrees of freedom. **Variance.** For a standard normal, `E[Z^4] = 3`, so `Var(Z^2) = 3 - 1 = 2`. Independence lets the variances add, giving `Var(V) = 2k`. Note the asymmetry with the mean: mean `k`, variance `2k`, so the standard deviation `sqrt(2k)` grows more slowly than the mean. **Shape.** With `k = 1` the distribution piles up near zero and has a long right tail — squaring sends most of a standard normal's mass into a small interval near 0. As `k` grows the shape becomes more symmetric; the skewness is `sqrt(8/k)`, which shrinks steadily. **Additivity.** If `V1` is chi-square on `d1` degrees of freedom and `V2` is an independent chi-square on `d2`, then `V1 + V2` is chi-square on `d1 + d2`. The proof is the construction itself: concatenating a list of `d1` squared normals with an independent list of `d2` squared normals produces a list of `d1 + d2` squared normals. Degrees of freedom add — they never multiply and never average. ## The t distribution as a ratio Suppose `Z` is standard normal and `V` is an independent chi-square on `k` degrees of freedom. Then `T = Z / sqrt(V / k)` has a t distribution with `k` degrees of freedom. Reading the formula tells you what the t distribution *is*: a standard normal whose scale has itself been made random by an independent, noisy estimate. The denominator `sqrt(V/k)` has mean near 1 but wobbles, and occasionally comes out small, which flings `T` far from zero. That is precisely why the t has heavier tails than the normal, and why the heaviness depends on `k`: as `k` grows, `V/k` concentrates near 1, the denominator stops wobbling, and the t approaches the standard normal. At the other extreme, `k = 1` gives the Cauchy distribution, which is also the ratio of two independent standard normals and has no mean at all. The t has finite variance `k/(k-2)` only when `k > 2`. ## The F distribution as a ratio of two chi-squares Let `V1` and `V2` be independent chi-squares on `d1` and `d2` degrees of freedom. Then `F = (V1 / d1) / (V2 / d2)` has an F distribution with numerator degrees of freedom `d1` and denominator degrees of freedom `d2`. Each chi-square is normalised by its own degrees of freedom first, so both numerator and denominator have mean 1 and the ratio hovers near 1 when the two sources of variation are comparable. Two structural facts follow from the construction: - the reciprocal of an F with `(d1, d2)` is an F with `(d2, d1)` — swapping the ratio swaps the pair; - the square of a t with `k` degrees of freedom is an F with `(1, k)`, since squaring `Z / sqrt(V/k)` produces `(Z^2 / 1) / (V / k)` and `Z^2` is a chi-square on one degree of freedom. The F has mean `d2 / (d2 - 2)` when `d2 > 2`, which is slightly above 1 and drifts toward 1 as the denominator degrees of freedom grow. ## Where the "degrees of freedom" number comes from The phrase counts independent squared normal contributions. When a quantity is computed from data after estimating something from the same data, one such contribution is used up: fixing an estimated centre removes one direction of free variation, so `n` observations can contribute only `n - 1` independent deviations around their own average. That accounting — how many free directions remain after the constraints imposed by estimation — is what determines the degrees of freedom parameter you carry into any of the three distributions above. ## Why it is worth knowing The chi-square, t and F are not three unrelated shapes to memorise; they are one normal-based construction viewed three ways: a sum of squares, a normal over a random scale, and a ratio of two random scales. Candidates who know the construction can reason about tail behaviour, about why degrees of freedom appear at all, and about limiting cases, instead of consulting a table. It is a differentiator rather than a screening question, but it separates candidates who understand the machinery from those who have only used it.

  • How is the t distribution assembled from a normal and a chi-square?
    T = Z / sqrt(V/k), where Z is standard normal and V is an independent chi-square on k degrees of freedom. The denominator is a noisy scale with mean near 1, and its occasional small values throw T far from zero, which is why the t has heavier tails than the normal. As k grows the denominator concentrates at 1 and the t approaches the standard normal; at k = 1 it is the Cauchy.
  • How is the F distribution built, and how does it relate to the t?
    F = (V1/d1) / (V2/d2), a ratio of two independent chi-squares each divided by its own degrees of freedom, so both parts have mean 1 and the ratio sits near 1. Squaring a t with k degrees of freedom gives an F with (1, k), since the numerator becomes a one-degree-of-freedom chi-square. Taking the reciprocal of an F with (d1, d2) gives an F with (d2, d1).
  • What happens to the shape of a chi-square as its degrees of freedom grow?
    It shifts right and spreads out — mean k and variance 2k — while becoming less lopsided, since its skewness is sqrt(8/k) and shrinks toward zero. With one degree of freedom the mass piles up near zero with a long right tail; with fifty it is a broad, nearly symmetric hump centred near fifty.
  • Why can a chi-square variable never be negative?
    Because it is defined as a sum of squares of standard normal variables, and every term in that sum is at least zero. The definition, not a convention or a truncation, is what confines it to the non-negative half line, so any procedure that seems to produce a negative chi-square value contains an arithmetic error.

saying these in an interview costs you the question

  • Says a chi-square variable can take negative values
  • Gives the variance as k rather than 2k
  • Thinks degrees of freedom multiply when chi-squares are added
  • Describes the t as a ratio of two independent normals
  • Confuses squaring the normals with taking absolute values

context