What does the Central Limit Theorem say about the average of many independent samples?
answer
- about the average, not the data
- the parent shape washes out
- one condition on the second moment
- the flat die roll becomes a bell
- centred at mu, variance sigma^2 over n
basics
~10 sThe Central Limit Theorem says that averaging many independent draws from almost any distribution with finite variance produces an average whose distribution is approximately normal, even when the individual observations are not remotely normal.
solid answer
~50 sThe Central Limit Theorem (CLT) is a statement about the *average*, not about the data. Take X1 ... Xn drawn independently from the same distribution with mean `mu` and finite variance `sigma^2`. The theorem says that as n grows, the distribution of the sample average `Xbar` approaches a normal distribution centred at `mu` with variance `sigma^2 / n`. Formally, the standardised quantity `sqrt(n) * (Xbar - mu) / sigma` converges in distribution to a standard normal. Two things matter. First, the shape of the parent distribution is almost irrelevant: a single roll of a fair die is flat over 1-6, yet the average of 30 rolls is close to bell-shaped. Second, the theorem is not unconditional — the draws must be independent and identically distributed and the variance must be finite. The raw observations never become normal; only the distribution of their average does.
code
python · 12 linesimport random, statistics
from collections import Counter
def mean_of_k(k):
return sum(random.random() for _ in range(k)) / k
for k in (1, 2, 30):
means = [mean_of_k(k) for _ in range(20000)]
counts = Counter(min(int(m * 10), 9) for m in means)
print("k =", k, " sd of the average =", round(statistics.pstdev(means), 3))
for b in range(10):
print(" %.1f-%.1f %s" % (b / 10, (b + 1) / 10, "#" * (counts[b] // 200)))go deeper
Be ready to state it in one sentence: the average of many independent draws is approximately normal whatever the data looks like. Know that it describes the average, not the observations.
Explain the mechanics: i.i.d. draws, finite variance, and the limit N(mu, sigma^2 / n). Interviewers here expect you to distinguish the distribution of the data from the distribution of the average out loud.
Show you know when the approximation is actually usable on real data — how skewness slows it down and what you would check before leaning on normal reasoning for a production metric.
Own the framing: the CLT is why normal-approximation arithmetic is the default across an analytics org, and where that default quietly fails. Be ready to say which metrics you would not let it be applied to.
## The statement Let `X1, X2, ..., Xn` be independent and identically distributed (i.i.d.) random variables with a finite mean `mu = E[X]` and a finite variance `sigma^2 = Var(X)`. Define the sample average `Xbar_n = (X1 + X2 + ... + Xn) / n` The Central Limit Theorem (CLT) says that as `n -> infinity`, the standardised average `Z_n = sqrt(n) * (Xbar_n - mu) / sigma` converges **in distribution** to a standard normal, `N(0, 1)`. The practical restatement, the one you use at an interview whiteboard, is: for large n, `Xbar_n` is approximately `N(mu, sigma^2 / n)`. The same theorem covers the *sum*: `X1 + ... + Xn` is approximately `N(n*mu, n*sigma^2)`, since the sum is just n times the average. ## What each piece means **i.i.d.** — independent means one draw carries no information about another; identically distributed means every draw comes from the same distribution. Both can be relaxed in more general versions of the theorem, but the classical statement assumes them. **Finite variance** — the distribution must have a variance that is a finite number. This is the condition candidates forget, and it is the one that actually fails in practice for very heavy-tailed data. **Converges in distribution** — this is a statement about *shapes of distributions*, not about individual numbers. It says the probability that `Z_n` lands below any fixed value gets closer and closer to the standard normal probability of landing below that value. No particular realised average is guaranteed to be anything. **Approximately** — the CLT is an asymptotic result. At any finite n the normal shape is an approximation, and how good it is depends on the parent distribution, especially its skewness. ## What the CLT is not The single most common misreading is that the CLT makes *the data* normal. It does not. If you collect a million session durations from a long-tailed distribution and plot a histogram of the raw values, you get a long-tailed histogram, exactly as before — a bigger sample gives you a *better picture of the true skewed shape*, not a bell. The bell appears only when you plot the distribution of averages computed from repeated samples. A second misreading is that the CLT is about the average settling down near the true mean as data accumulates. That convergence is a different result. The CLT is finer-grained: it describes the *shape of the fluctuation* around the mean, and tells you that fluctuation is normal-shaped once rescaled by `sqrt(n)`. A third is treating the theorem as unconditional. Without finite variance there is no normal limit at all — the limit may be a different, heavy-tailed distribution, or the average may fail to stabilise entirely. ## Why the bell appears Intuitively, an average blends many independent contributions. Each draw can push the average up or down, and those pushes partly cancel. Extreme outcomes require many draws to conspire in the same direction, which is exponentially unlikely; middling outcomes can be produced in enormously many ways. That combinatorial squeeze is what produces the bell, and it is why the parent shape washes out. A classic demonstration uses the flat `Uniform(0, 1)` distribution. One draw is flat across the interval. The average of two draws is triangular, peaked at 0.5 — you can reach the middle in many ways and the endpoints in only one. By the time you average thirty draws, the histogram is visually indistinguishable from a bell centred at 0.5, and it is much narrower than the original interval, because averaging shrinks the spread by a factor of `sqrt(n)`. The same happens with a fair die. A single roll is flat: each of 1 through 6 has probability 1/6, mean 3.5, variance 35/12. The average of 30 independent rolls is bell-shaped and tightly concentrated around 3.5. Nothing about the die changed; the averaging did the work. ## Why interviewers ask it The CLT underpins most of the everyday normal-approximation reasoning in analytics: it is the reason a mean computed from a messy, non-normal metric can still be reasoned about with normal-curve arithmetic. An interviewer wants to hear the three components — i.i.d. draws, finite variance, and the conclusion about the *average* rather than the data — plus an honest note that 'large n' is not a fixed number and depends on how skewed the parent distribution is.
- Does the CLT require the underlying population to be normally distributed?No — the opposite is the point. The parent distribution can be flat, skewed, discrete or bimodal; the theorem still gives an approximately normal distribution for the average. It only requires independent, identically distributed draws with a finite variance. If the parent already is normal, the average is exactly normal at every n, so the theorem adds nothing there.
- Does the CLT apply to the sum of the draws as well as their average?Yes. The sum is n times the average, so it inherits the same result with rescaled parameters: for large n the sum is approximately normal with mean `n * mu` and variance `n * sigma^2`. Note the contrast — the sum's spread grows like `sqrt(n)` while the average's spread shrinks like `1 / sqrt(n)`.
- What does convergence in distribution actually mean here?It means the cumulative probabilities converge, not the values. For any fixed number z, the probability that the standardised average falls below z approaches the standard normal probability of falling below z. It says nothing about any single realised average, and it does not claim the random variables themselves converge to anything.
Individual voices in a crowd are wildly different, but the crowd's average loudness is smooth and predictable. The CLT says the smoothness of the average is guaranteed, not the smoothness of any voice.
saying these in an interview costs you the question
- Claims the raw data becomes normal as the sample grows
- States the CLT requires a normal population to begin with
- Says it applies to every distribution, with no finite-variance condition
- Confuses it with the sample mean converging to the true mean
- Treats n = 30 as a guarantee rather than a rough guideline