skip to content

Why does the Central Limit Theorem fail for Cauchy-distributed data?

level: seniorimportance: nice to knowfreq 28%

answer

  1. the condition nobody checks
  2. the defining integral does not converge
  3. averaging changes nothing at all
  4. the ratio of two standard normals
  5. the same distribution at every n

basics

~20 s

The Cauchy has tails so heavy it has no finite mean or variance, and finite variance is exactly what the theorem requires. The average of n Cauchy draws is again exactly Cauchy, so it never narrows.

solid answer

~50 s

The classical Central Limit Theorem requires i.i.d. draws with a **finite variance**. The Cauchy distribution violates this in the strongest possible way: its tails decay slowly enough that the integral defining the mean does not converge, so it has no finite mean and no finite variance. The consequence is remarkable — the average of n i.i.d. standard Cauchy draws has *exactly* the standard Cauchy distribution, for every n. Averaging a million draws gives you a quantity distributed identically to a single draw. There is no narrowing, no bell shape and no limit to converge to. Practically you spot this when a running average keeps jumping to a new level rather than settling, dragged by one observation larger than everything seen so far. A useful check: if the largest observation stays a substantial share of the running total as data accumulates, normal-approximation reasoning is unsound.

go deeper

for a junior

Know that the theorem has a finite-variance condition and that some distributions violate it. Naming the Cauchy as the standard counterexample is enough at this level.

for a middle

Explain why the Cauchy has no finite mean or variance and state the stability result: the average of n draws has exactly the same distribution as a single draw, so averaging buys nothing.

for a senior

Show how you would spot this on real data — running averages that jump, a running variance that keeps climbing, a single observation dominating the total — and what you would report instead of a mean.

for a principal

Own the risk framing: heavy-tailed metrics quietly break normal-approximation reasoning across an org, and understated tail risk is the expensive failure. Decide which metrics may be summarised by a mean at all.

## The condition the theorem actually needs The classical Central Limit Theorem assumes i.i.d. draws with a finite mean `mu` and a **finite variance** `sigma^2`. The finite-variance requirement is not decoration. The theorem's conclusion — that `sqrt(n) * (Xbar - mu) / sigma` tends to a standard normal — mentions `sigma` explicitly, so if `sigma` is infinite the statement has nothing to say. ## What makes the Cauchy pathological The standard Cauchy has density `1 / (pi * (1 + x^2))`. Its tails decay like `1 / x^2`, which is slow. Slow enough that computing the expectation requires integrating something behaving like `1 / x` out to infinity, and that integral diverges. So the Cauchy has **no finite mean** — not a large mean, not an unknown one; the defining integral simply does not converge. With no finite mean, there is no finite variance either. The Cauchy also has a memorable construction: the ratio of two independent standard normal variables is standard Cauchy. Two of the most tractable random quantities in statistics divide into one of the least tractable. ## The stability property The striking consequence is that the average of n i.i.d. standard Cauchy variables is *itself exactly standard Cauchy*, for every n. Not approximately, not asymptotically — exactly, at n = 1, n = 100 and n = 10,000,000. Averaging accomplishes nothing at all: the distribution of the average is identical to the distribution of a single observation. This is the cleanest possible refutation of the intuition that 'more data always makes an average more reliable'. That intuition rests on the variance of the average shrinking as `sigma^2 / n`, and if `sigma^2` is infinite the argument evaporates. ## The behaviour you would actually see Compute a running average of Cauchy draws and plot it against the number of observations. It does not converge. It wanders at some level for a while, then a single enormous observation arrives and yanks it to a completely new level, where it wanders again until the next one. The plot has visible jump discontinuities that keep occurring no matter how far out you go, because the tail keeps producing observations larger than everything previously seen. ## The general picture beyond the Cauchy The Cauchy is one member of a family. A power-law (Pareto-type) tail with index `alpha` — meaning the probability of exceeding x decays like `x^(-alpha)` — behaves as follows: - `alpha > 2`: finite mean and finite variance; the classical CLT applies, though convergence can be slow. - `1 < alpha <= 2`: the mean exists but the variance is infinite. The classical CLT fails. Suitably scaled sums converge to a non-normal **stable** distribution, and the correct scaling is `n^(1/alpha)` rather than `sqrt(n)`. - `alpha <= 1`: not even the mean exists. The Cauchy sits at `alpha = 1`. So 'the CLT fails' does not always mean 'nothing converges' — often something converges, just to a heavy-tailed stable law under a different scaling, and treating it as normal understates tail risk severely. ## Diagnosing it in real data You cannot observe an infinite variance directly; every finite dataset has a finite sample variance. What you look for is instability. 1. **Running average.** Plot it against sample size. Repeated level shifts driven by single observations, rather than a settling trend, are the signature. 2. **Running variance.** If the sample variance keeps climbing as you add data instead of stabilising, the population variance may not exist. 3. **Maximum-to-sum ratio.** Track the largest observation as a fraction of the running total. Under finite variance this ratio drifts toward zero; when it stays substantial no matter how much data arrives, the tail is dominating. 4. **Tail plots.** Rank the observations and plot the tail on log-log axes. A straight line indicates a power law, and its slope estimates `alpha`, which tells you which regime you are in. ## What to do instead If the variance genuinely does not exist, the mean may be the wrong summary rather than a hard-to-estimate one. Medians and other quantiles remain well defined and stable for any distribution, including the Cauchy, and typically become the honest target. Where the mean must be reported for a business reason, the right response is to be explicit that no finite-sample average is trustworthy, and to quantify uncertainty by a method that does not assume a finite variance rather than by normal-curve arithmetic. ## Why this comes up at senior level Genuinely infinite variance is rare in practice; what is common is data heavy-tailed enough that the CLT is effectively unusable at any n you will ever have. Interviewers use the Cauchy as the clean extreme case to test whether a candidate treats the finite-variance condition as a real assumption to check, or as boilerplate to recite.

  • What happens to a running average of Cauchy draws as data accumulates?
    It never settles. It drifts around some level, then a single extreme observation shifts it to a new level, and that keeps happening indefinitely. Because the average of n draws has exactly the same distribution as one draw, there is no sense in which more data improves it.
  • If a distribution has a finite mean but infinite variance, does anything still converge?
    Yes, but not to a normal. For a power-law tail with index between 1 and 2, suitably scaled sums converge to a non-normal stable distribution, with `n^(1/alpha)` scaling rather than `sqrt(n)`. Treating such data as normal understates tail probabilities substantially.
  • How would you detect a possible infinite variance in real data?
    Watch stability rather than any single statistic. Plot the running variance — if it climbs steadily instead of levelling off, that is a warning. Also track the largest observation as a share of the running total; under finite variance that ratio should fall toward zero as data accumulates.
  • What summary would you report instead of a mean for such data?
    Quantiles. A median is well defined for any distribution, including the Cauchy, and is stable under heavy tails because it depends on ranks rather than magnitudes. If a mean is genuinely required by the business question, state explicitly that no finite-sample average of that metric is trustworthy.

Averaging Cauchy draws is like pooling money with people whose fortunes are unbounded: no matter how many join, the pot's per-head share is still decided by whoever happens to be the richest in the room.

saying these in an interview costs you the question

  • Says the CLT applies to any distribution whatsoever
  • Claims a big enough sample fixes infinite variance
  • Confuses no finite mean with a mean of zero
  • Assumes failed CLT means nothing converges at all
  • Treats a finite sample variance as proof the population variance exists

context