Why are sample skewness and kurtosis unreliable estimates on a small sample?
answer
- third and fourth powers, not first
- one far-out point dominates the sum
- the estimate is a random variable
- standard error shrinks only as sqrt(n)
- sqrt(6/n) and sqrt(24/n) under normality
basics
~20 sBoth statistics raise deviations to the third or fourth power, so a handful of extreme points dominate them. Their standard errors are large on small samples, roughly sqrt(6/n) for skewness and sqrt(24/n) for excess kurtosis under normality.
solid answer
~50 sSkewness cubes the z-scores and kurtosis raises them to the fourth power, so a single point 4 standard deviations out contributes 64 or 256 units while a typical point contributes about 1. That means both estimates are driven by the few observations you have least information about — the tail. Under normality the approximate standard errors are `sqrt(6/n)` for skewness and `sqrt(24/n)` for excess kurtosis, so at `n = 30` they are roughly 0.45 and 0.89: perfectly normal data routinely produces a sample skewness of 0.8 or an excess kurtosis of 1.5 by chance alone. There is also a hard ceiling — the sample kurtosis cannot exceed a bound that grows with `n`, so a sample of 20 literally cannot display the kurtosis of a genuinely heavy-tailed process. Report both statistics with the sample size, treat small-sample values as directional, and never build a decision rule on a shape statistic from a few dozen points.
code
python · 16 linesimport random, statistics
def excess_kurtosis(xs):
n = len(xs)
m = statistics.fmean(xs)
m2 = sum((x - m) ** 2 for x in xs) / n
m4 = sum((x - m) ** 4 for x in xs) / n
return m4 / m2 ** 2 - 3.0
random.seed(0)
vals = sorted(
excess_kurtosis([random.gauss(0.0, 1.0) for _ in range(30)])
for _ in range(2000)
)
print(round(statistics.fmean(vals), 2)) # average estimate
print(round(vals[100], 2), round(vals[1900], 2)) # 5th and 95th percentilesgo deeper
Know that these are estimates from a sample, not properties you have measured exactly, and that a small sample gives an unreliable answer.
Explain the mechanism: third and fourth powers give extreme observations huge leverage, and the standard errors shrink only with the square root of the sample size.
Show operational habit — quoting the sample size, attaching a bootstrap or normal-theory interval, and running the drop-the-largest-point sensitivity check before acting on a shape statistic.
Own the policy question: whether automated shape checks belong in a data pipeline at all, what minimum sample size makes them meaningful, and how you stop teams from reading noise as a finding.
## The leverage problem Sample skewness averages cubed z-scores and sample kurtosis averages z-scores raised to the fourth power. Those exponents are the whole story. A point 1 SD from the mean contributes 1 to either sum. A point 4 SD out contributes 64 to the skewness sum and 256 to the kurtosis sum. So both statistics are, in effect, measurements of the extreme tail — and the extreme tail is exactly the part of the distribution where a small sample carries the least information. With 30 observations you have perhaps one or two points beyond 2 SD, and whether those points landed at 2.1 or 3.4 will swing your estimate substantially. Contrast this with the mean and the variance, which average first and second powers and therefore pool information from every observation roughly evenly. ## The standard errors Under the assumption that the data really is normal, the large-sample approximate standard errors are `SE(skewness) ~ sqrt(6 / n)` `SE(excess kurtosis) ~ sqrt(24 / n)` Put numbers on those: | n | SE(skewness) | SE(excess kurtosis) | |---|---|---| | 30 | 0.45 | 0.89 | | 100 | 0.24 | 0.49 | | 1000 | 0.077 | 0.15 | | 10000 | 0.024 | 0.049 | At `n = 30`, a two-standard-error band around a true skewness of zero runs from about -0.9 to +0.9. Since a common rule of thumb calls anything past 0.5 "moderately skewed" and anything past 1 "strongly skewed," perfectly normal data will regularly hand you a sample skewness that looks meaningfully skewed. The same at `n = 30` for excess kurtosis: a value of 1.5 or -1.2 is within ordinary sampling variation. Only at the thousands does the estimate become sharp enough to distinguish mild shape departures. Note also that these standard errors are themselves derived under normality — on genuinely heavy-tailed data the sampling variability of the kurtosis estimate is *worse*, sometimes far worse, because the estimator's own variance depends on the eighth moment, which may barely exist. ## The algebraic ceiling There is a second, less-known trap: sample skewness and kurtosis are **bounded by the sample size**. The magnitude of the sample skewness cannot exceed roughly `sqrt(n)`, and the sample kurtosis cannot exceed a bound of order `n`. A process whose true excess kurtosis is 30 simply cannot show that value in a sample of 15 — the arithmetic forbids it. So on a small sample, a modest kurtosis reading is not evidence that the tails are modest; it may be the ceiling of what the sample could have reported. ## Bias The plain moment-ratio estimators are also biased. For normal data, the average sample excess kurtosis computed as `m4 / m2^2 - 3` sits below zero on small samples, and the plain sample skewness is attenuated toward zero. The usual reported forms apply small-sample corrections — the adjusted Fisher-Pearson skewness multiplies by `sqrt(n(n-1)) / (n-2)`, and the corresponding kurtosis correction inflates the estimate — but those corrections address the centring, not the enormous variance. A corrected estimate of a wildly noisy quantity is still wildly noisy. ## What to do instead - **Always report `n` alongside the statistic.** A skewness of 0.7 means something completely different at 40 rows and at 400,000. - **Attach an interval.** Either the normal-theory standard error above, or a bootstrap interval, which handles the non-normal case better. A wide interval crossing zero is the honest answer. - **Look at the whole shape.** A histogram or an empirical distribution over the same data carries far more information than two scalars and makes it obvious when one record is driving everything. - **Recompute without the largest observation.** If dropping one point moves the kurtosis by an order of magnitude, you have learned that you are measuring that point, not the distribution. - **Never gate a decision on a small-sample shape statistic.** Automated rules of the form "if |skewness| > 1 then transform" applied to segments with 20 rows each will fire essentially at random. ## What a strong answer sounds like An interviewer is checking whether you distinguish the *parameter* from the *estimate*. The population skewness of a symmetric distribution is exactly zero; the sample skewness computed from 30 draws is a random variable with a standard deviation near 0.45. Confusing those two — reading a sample value as though it were the truth — is the failure mode this question is designed to catch.
- At n = 30, what sample skewness would you consider unremarkable for genuinely normal data?Under normality the standard error is about `sqrt(6/30)` = 0.45, so a two-standard-error band runs to roughly plus or minus 0.9. Anything inside that is ordinary sampling variation, even though a common rule of thumb would label 0.7 as moderately skewed. The sample size has to be in the hundreds before values that size mean much.
- How would you attach uncertainty to a kurtosis estimate on non-normal data?The normal-theory formula `sqrt(24/n)` understates the variability when the tails are genuinely heavy, because the estimator's own variance depends on very high moments. A bootstrap interval is the practical alternative: resample the data with replacement, recompute the statistic each time, and report the spread of those values.
- Why can a small sample fail to reveal a genuinely heavy-tailed distribution?Two reasons. You may simply not have drawn any of the rare extreme values that create the heavy tail. And the sample kurtosis is algebraically bounded by a limit that grows with `n`, so a sample of 15 cannot report the kurtosis of a process whose true value is 30 no matter what it contains.
- What is a quick sensitivity check on a shape statistic?Recompute it after removing the single most extreme observation. If the kurtosis moves by an order of magnitude, the statistic is describing that one record rather than the distribution, and you should look at the record itself and at the full shape before drawing any conclusion.
saying these in an interview costs you the question
- Reads a sample shape statistic as the population value
- Ignores sample size when quoting a skewness figure
- Applies fixed thresholds to samples of a few dozen
- Thinks the small-sample correction removes the noise
- Assumes a low kurtosis reading rules out heavy tails