What does a positive excess kurtosis tell you about a distribution's tails?
answer
- fourth power of the z-scores
- the normal is the reference point
- three is subtracted for a reason
- it is about the tails, not the peak
- above zero means extremes are more common
basics
~20 sPositive excess kurtosis means heavier tails than a normal distribution with the same standard deviation: extreme deviations happen more often than normal probabilities predict. Excess kurtosis subtracts 3, the kurtosis of a normal, so a normal scores exactly 0.
solid answer
~50 sKurtosis is the fourth standardised moment: the average of the z-scores raised to the fourth power. A normal distribution has kurtosis 3, so excess kurtosis is defined as kurtosis minus 3, which puts the normal at zero as a reference point. Positive excess kurtosis (leptokurtic) means the distribution puts more mass far from the centre than a normal with the same standard deviation — extreme observations are relatively more common. Negative excess kurtosis (platykurtic) means lighter tails; a uniform distribution has excess kurtosis -1.2. Daily stock returns are the standard illustration: they routinely show excess kurtosis well above zero, and moves beyond 5 standard deviations turn up several times a decade, whereas a normal distribution puts a move that large at roughly one day in a few million. It is a statement about tail weight, not about how pointy the peak looks.
go deeper
Recall the reference points: a normal has kurtosis 3, excess kurtosis subtracts that 3, and a positive excess value means more extreme observations than a normal would produce.
Explain the mechanics — z-scores to the fourth power, why that weighting makes the statistic a tail measure, and why the peakedness description is misleading. Expect to be asked which convention your number uses.
Demonstrate the operational consequence: normal-based tail thresholds under-count extremes, estimators get noisier, and variance stops being an adequate risk summary. Bring a real heavy-tailed series you have handled.
Own the modelling call. Argue when a heavy-tailed reality should be met with a different model, a different loss function, or an explicit tail budget, and how you would communicate tail risk to people who only see a mean and a standard deviation.
## The definition Kurtosis is the **fourth standardised moment**. For a sample with mean `xbar` and standard deviation `s`, `b2 = (1/n) * sum( ((xi - xbar) / s)^4 )` Every deviation is converted to a z-score and raised to the fourth power. The fourth power is always positive, so unlike skewness there is no sign to read — kurtosis only reports *how much* of the total variation comes from far-out observations, not which side they are on. A normal distribution has kurtosis exactly 3, whatever its mean and standard deviation. Because that constant is such a natural reference, the quantity almost always reported is **excess kurtosis**: `excess kurtosis = kurtosis - 3` so a normal distribution scores 0, heavier-tailed shapes score above 0, and lighter-tailed shapes score below 0. When someone says "the kurtosis is 4," always establish which convention they mean — 4 on the raw scale is mildly heavy-tailed, while 4 on the excess scale is dramatically heavy-tailed. ## Heavy versus light tails - **Positive excess kurtosis — leptokurtic.** More probability sits far from the centre than a normal with the same standard deviation would place there. Extreme deviations are relatively frequent. Financial returns, insurance claim sizes, and network latency records typically look like this. - **Zero excess kurtosis — mesokurtic.** The normal reference itself, and any other distribution that happens to match its fourth moment. - **Negative excess kurtosis — platykurtic.** Lighter tails than a normal: extremes are rarer and the variation is more evenly spread across a bounded middle. A continuous uniform distribution has excess kurtosis -1.2. The theoretical floor is -2, attained by a fair coin flip coded as 0/1, where every observation sits at one of two points. ## Why "peakedness" is the wrong intuition Older textbooks describe kurtosis as the peakedness of a distribution. This is misleading and interviewers often probe it. The fourth power gives an observation 3 SD from the centre a weight of 81 while an observation 0.5 SD from the centre contributes 0.0625 — near-centre points contribute essentially nothing to the sum. Kurtosis is therefore dominated almost entirely by the far tails; the height of the peak is at best a by-product of the fact that a fixed standard deviation forces a trade between tail mass and central mass. The defensible one-line description is *tail weight relative to a normal*, or, more precisely, the propensity to produce values far from the centre. ## The concrete example: daily equity returns Standardise a long history of daily returns and compute the fourth moment: the result is reliably and substantially above 3, so excess kurtosis is comfortably positive. The practical consequence is easiest to see at the extreme. Under a normal distribution, a move beyond 5 standard deviations in either direction has probability of roughly 6 in 10 million; at about 250 trading days a year, you would expect one such day roughly once in several thousand years. Real return series contain several 5-sigma days per decade. Nothing is wrong with the arithmetic — the normal model is simply the wrong tail model, and excess kurtosis is the single number that flags that mismatch. ## Kurtosis and skewness are different questions A distribution can be perfectly symmetric (skewness 0) and still have enormous excess kurtosis: a symmetric shape with a dense centre and two long tails. Conversely, a strongly skewed distribution normally *also* has positive excess kurtosis, because the long tail on one side contributes fourth-power terms. In fact the two are linked by an inequality that always holds: `excess kurtosis >= skewness^2 - 2`. Highly skewed data therefore cannot be strongly platykurtic. ## Practical consequences Excess kurtosis is a warning that procedures calibrated on a normal reference will misprice the tail. Concretely: - Interval and threshold rules that assume normal tail probabilities ("three standard deviations covers 99.7%") under-state how often you will see a value outside the band. - Sample means still converge, but slowly: heavy tails mean a single observation can carry a large share of the total, so estimates of the mean and of the variance are themselves noisier than their formulas suggest at a given sample size. - Variance itself becomes a poor sufficient summary of risk, because two distributions with the same standard deviation can have very different exposure to a rare large value. ## Reporting it honestly Always say which convention you used, always report the sample size next to the number, and remember that the estimate itself has a large standard error on small samples — under normality it is roughly `sqrt(24/n)`, which is about 0.9 at `n = 30`. A single number without that context invites over-reading.
- Why is it wrong to describe kurtosis as the peakedness of a distribution?Because the fourth power crushes the contribution of points near the centre. A z-score of 0.5 contributes 0.0625 while a z-score of 3 contributes 81, so the statistic is driven almost entirely by far-out observations. The correct description is tail weight relative to a normal; any apparent peak change is a side effect of holding the standard deviation fixed.
- What kind of distribution has a negative excess kurtosis?One with lighter tails than a normal, where extremes are rarer and the mass is spread across a bounded middle. A continuous uniform distribution has excess kurtosis -1.2. The theoretical minimum is -2, reached by a two-point distribution such as a fair coin coded 0/1.
- Can a distribution have zero skewness and very large excess kurtosis at the same time?Yes. The two measure different things. A perfectly symmetric distribution with a dense centre and two long tails has skewness 0 by symmetry and large positive excess kurtosis from the fourth-power terms in both tails. The reverse constraint does exist: excess kurtosis is always at least `skewness^2 - 2`.
Two roads can have the same average speed and the same variability, yet one has occasional 200 km/h outliers. Excess kurtosis is the number that separates them.
saying these in an interview costs you the question
- Says kurtosis measures how pointy the peak is
- Forgets that a normal has kurtosis 3, not 0
- Reports kurtosis without saying whether 3 was subtracted
- Thinks high kurtosis implies the data is skewed
- Claims heavy tails also mean a larger standard deviation