Why is the n = 30 rule of thumb for the Central Limit Theorem unreliable on skewed data?
answer
- the theorem names no threshold
- asymmetry sets the pace
- the whale that is usually absent
- required n grows with the square of it
- roughly 25 times skewness squared
basics
~20 sn = 30 is a teaching guideline, not part of the theorem. The more skewed the population, the larger the sample must be — heavy right tails can need thousands of observations before the average looks normal.
solid answer
~50 sThe Central Limit Theorem is asymptotic: it promises normality as n goes to infinity and names no finite threshold. 'n = 30' is a teaching heuristic that happens to work for roughly symmetric, light-tailed populations. How fast the approximation kicks in is driven mainly by **skewness**: the leading error term in the normal approximation to the average shrinks like `skewness / sqrt(n)`, so the sample size needed to hold a given accuracy grows roughly with the *square* of the skewness. A common guideline is `n > 25 * skewness^2`. Revenue per user is the standard counterexample: most users spend nothing or a little, a handful spend enormously, and skewness of 5 or 10 is routine. Plug that in and you need hundreds to thousands of observations, not thirty. With n = 30 the distribution of the average is still visibly right-skewed, so normal-based tail probabilities are biased.
go deeper
Know that n = 30 is a rule of thumb, not part of the theorem, and that skewed data needs more. Being able to say 'it depends on the shape' already puts you ahead.
Explain why skewness sets the convergence speed and quote the rough n > 25 * skewness^2 guideline with a worked number. Describe what the distribution of averages looks like when n is too small.
Show the diagnosis you would actually run on a production metric, and be specific about the directional damage: skewed averages misprice tail probabilities the same way every time, so errors accumulate rather than cancel.
Own the organisational call: whether the team reports means at all on heavy-tailed revenue metrics, what guardrails stop normal-approximation reasoning being applied blindly, and what you trade away by switching estimand.
## Where 'n = 30' came from The Central Limit Theorem is a limit statement. It says the standardised average converges to a standard normal as the sample size grows without bound; it contains no number 30, no number 100, and no threshold of any kind. The rule of thumb is a pedagogical convenience: for populations that are roughly symmetric and light-tailed — a uniform, a flat die roll, a mild mixture — the average is already close to bell-shaped by about thirty draws, so textbooks adopted the number and it stuck. It was never a theorem, and treating it as one is one of the most consequential everyday errors in applied statistics. ## What actually controls the speed The accuracy of the normal approximation to the distribution of an average is governed by how asymmetric the parent distribution is. The leading correction term to the normal approximation is proportional to `skewness / sqrt(n)`. Setting that error below a fixed tolerance and solving for n shows the required sample size grows with the **square** of the skewness. A widely quoted operational guideline, attributed to Cochran, is: `n > 25 * skewness^2` Run some numbers through it: - A symmetric distribution has skewness 0 — no skew penalty; the rule reduces to 'a modest n is fine'. - An exponential distribution has skewness exactly 2, giving `n > 100`. - A lognormal with log-scale parameter 1 has skewness above 6, giving `n > 900`. - Revenue-per-user distributions with a handful of whales routinely show sample skewness of 10 or more, pushing the requirement into the thousands. Kurtosis (heavy tails without asymmetry) matters too, entering at the next order, but skewness dominates in practice because business metrics are bounded below at zero and unbounded above. ## The revenue-per-user case Consider per-user revenue over a month. A large share of users spend nothing at all, giving a spike at zero. Most paying users spend a small amount. A tiny fraction spend hundreds or thousands of times the median. The distribution is a spike plus a long right tail. Average thirty such users and one of two things happens. Usually no whale is in the sample, and the average lands well below the true mean. Occasionally a whale is included, and the average jumps far above it. The resulting distribution of averages is not a bell: it is itself right-skewed, with a dense body below the true mean and a stretched upper tail. That is the CLT working correctly and slowly — it *will* become normal, just not at thirty. The practical damage is directional. Because the distribution of the average is right-skewed at small n, normal-curve reasoning misprices both tails: it understates how often the average lands moderately low and understates how far the average can travel high. Any tail probability read off a normal curve is then systematically wrong, and the error does not average out across repeated analyses — it biases the same way every time. ## How to check rather than assume The honest answer in an interview is that you diagnose rather than defer to a number. 1. **Look at the parent distribution.** Plot it. A visible spike-plus-tail shape immediately kills the n = 30 assumption. 2. **Compute the sample skewness** and put it through the `25 * skewness^2` guideline for an order-of-magnitude sample-size requirement. 3. **Check the tail's concentration.** If the top 1% of observations contribute a large share of the total, a small sample's average is effectively a lottery on whether that 1% appears. 4. **Resample.** Draw many samples of your actual n from the data you have and plot the resulting averages. If that histogram is visibly skewed, the approximation is not there yet at that n. ## Remedies When n cannot be increased far enough, the options are to change the quantity or change the method. Analysing a transformed scale (a log, for instance) tames the skew, but note that the mean of the logs is not the log of the mean, so the business question must genuinely be about the transformed quantity. Trimming or winsorising the extreme tail reduces skew at the cost of changing what you are estimating and needs an explicit, pre-registered rule. Reporting a different summary — a median or a trimmed mean — sidesteps the issue when the business question tolerates it. Alternatively, aggregate at a coarser unit, since averages of already-averaged groups are far better behaved. ## What the interviewer is listening for They want to hear that n = 30 is a heuristic and not part of the theorem, that skewness is the driver, that the required n scales roughly with skewness squared, and that you would inspect the data rather than assert a threshold. A candidate who says 'we had 30 samples so the CLT applies' has revealed that they learned a slogan rather than a theorem.
- What sample size does the 25 * skewness^2 guideline suggest for data with skewness 6?Roughly `25 * 36 = 900` observations, thirty times the usual rule of thumb. The quadratic term is what makes heavy skew so expensive: doubling the skewness quadruples the sample size you need for the same approximation quality.
- How would you check empirically whether your n is large enough?Repeatedly draw samples of your actual size from the observed data and plot the distribution of the resulting averages. If that histogram is visibly right-skewed rather than symmetric, the normal approximation is not usable at that n. It is a direct look at the object the CLT makes a claim about.
- Why is a mild left or right skew less dangerous than a heavy right tail?Mild skew produces a small leading error that `sqrt(n)` erases quickly. A heavy right tail means a large share of the total sits in rare observations, so whether the sample happens to contain them dominates the average. That lottery is what keeps the distribution of averages non-normal for a long time.
- Does averaging on the log scale fix the problem?It fixes the *approximation* — logs of a right-skewed positive metric are usually near-symmetric, so the average of the logs behaves well at modest n. But it changes the estimand: the average of the logs corresponds to a geometric mean, not the arithmetic mean, so it only helps if the business question genuinely concerns the transformed quantity.
Thirty is a speed limit copied from a quiet suburb and posted on every road. On a straight, flat street it is fine; on a mountain switchback with a long drop it will get you killed.
saying these in an interview costs you the question
- Treats n = 30 as a condition stated in the theorem
- Ignores skewness when judging sample-size adequacy
- Assumes more data always fixes it without checking the tail
- Says the approximation error is symmetric in both tails
- Transforms to logs without noticing the estimand changed