When do skewness and kurtosis change a decision that mean and standard deviation alone would drive?
answer
- two moments only pin down a normal
- ask what the loss function looks like
- linear payoff versus convex payoff
- three-sigma sizing assumes a normal tail
- shape as a trigger, not a dashboard column
basics
~20 sWhenever the decision is sensitive to rare large values rather than to typical ones. Two distributions can share a mean and a standard deviation while differing sharply in their third and fourth standardised moments, so a buffer or threshold sized on the first two moments alone can be badly wrong.
solid answer
~50 sThe first two moments fully describe a normal distribution and nothing else. Two variables with identical mean 100 and standard deviation 20 can differ enormously in skewness and kurtosis, which is exactly where the tail probability lives. So the test I apply is whether the payoff is roughly linear in the value or convex in it. If the decision is "what is the typical outcome" or "what is the total across many independent draws," the mean and standard deviation carry the load and shape matters little. If the decision is a capacity buffer, a stock-out threshold, a solvency margin, an alerting cut-off, or anything where one large realisation is qualitatively worse than several medium ones, then tail weight decides the answer and mean-plus-three-SD sizing systematically under-provisions on a heavy-tailed variable. In those cases I stop summarising with two numbers and instead simulate from the empirical distribution or report the tail directly, with the shape statistics as the flag that told me to.
go deeper
Know the core fact: two datasets can have the same mean and the same standard deviation and still behave very differently at the extremes.
Be able to explain why the first two moments pin down a distribution only within the normal family, and why the third and fourth moments carry the information about rare large values.
Demonstrate the operational consequence with something you have sized — a buffer, a threshold, an alert — and explain how you checked the empirical exceedance rate rather than trusting a normal-theory band.
Own the portfolio decision: which metrics warrant treatment beyond two moments, how you keep that from turning into complexity everywhere, and how you get a team off a uniform template that is silently wrong for its heavy-tailed metrics.
## The premise A mean and a standard deviation are a complete description of a distribution only if you have already committed to the normal family. For everything else, they are two summaries out of infinitely many, and the two they capture are the ones least connected to rare events. The cleanest way to make this concrete is to construct two distributions with identical mean and identical standard deviation but different third and fourth standardised moments. One can be symmetric with light tails, bounded, and almost never far from centre. The other can be symmetric with heavy tails, mostly hugging the middle but occasionally producing a value six or eight standard deviations out. Every summary table built on mean and SD reports them as the same variable. Every decision that depends on the far tail treats them completely differently. ## The test I actually apply Ask what the loss function looks like as a function of the realised value. **Roughly linear payoff** — the total cost across many draws, average handle time across a quarter, aggregate spend. Here the mean does most of the work and the standard deviation tells you how precisely you know it. Shape matters mainly through estimation noise: heavy tails mean your estimate of the mean converges more slowly, so you need more data than the usual formulas suggest. But the decision itself is not shape-driven. **Convex or threshold payoff** — a capacity buffer that either holds or fails, a solvency margin, an alert threshold, a queue that collapses past a limit, an SLA breach. Here almost all of the expected loss is concentrated in the tail, and the first two moments are close to silent about it. This is the regime where skewness and kurtosis change the answer. **Asymmetric payoff** — where being wrong high and wrong low cost different amounts. Skewness, specifically, decides which side you should be conservative on. ## Why three-sigma sizing fails on heavy tails "Mean plus three standard deviations" is an implicitly normal rule: under a normal, that band covers about 99.87% of the upper tail, so the buffer is breached about one time in 750. On a distribution with substantial positive excess kurtosis, the same nominal band can be breached one time in 50 or one time in 20. The standard deviation has not lied — it correctly summarises the average squared deviation — but the mapping from standard deviations to probabilities is a property of the shape, not of the spread, and that mapping is exactly what excess kurtosis is telling you has changed. The error compounds because the failure is silent: the buffer looks conservative on paper, and the shortfall only reveals itself on the rare day it matters. ## What I do instead of adding a third and fourth number to the dashboard The temptation is to put skewness and kurtosis into the standard reporting template. I generally resist that, for three reasons. They are noisy statistics that invite over-reading; most consumers of a dashboard cannot convert a kurtosis of 4 into an action; and their value is diagnostic rather than operational. A better division of labour: - **Use the shape statistics as a trigger, not as an output.** Compute them once during characterisation of a metric. If the excess kurtosis is large, that is a signal that the metric needs a different treatment, and the finding gets written down as a property of the metric. - **Replace the parametric summary with a direct tail statement.** Rather than mean plus three SD, state the empirical exceedance level and the size of the exceedances when they happen. That is what the decision actually consumes. - **Simulate rather than derive.** For a convex payoff, resample from the observed distribution and evaluate the loss directly. This sidesteps the need to summarise shape into a scalar at all. - **Record the sample size and the vintage.** A shape characterisation from a period that contained no stress event is a characterisation of a calm regime, not of the process. ## The organisational failure mode The recurring pattern is a team that has standardised on a mean-and-SD template because it is uniform and easy to automate, applied uniformly to metrics whose shapes are wildly different. The template is not wrong for the well-behaved metrics; it is silently wrong for the heavy-tailed ones, and nobody notices because the template gives an answer either way. The lead's job is to know which metrics in the portfolio are in the second category and make sure those have a different treatment, rather than to demand that every metric carry four moments. ## What to say when challenged A reasonable pushback is that adding shape complexity is over-engineering. The honest answer is that it usually is, and the discipline is knowing where it is not: name the small set of decisions in your system whose cost is dominated by rare large values, and confine the extra rigour to those. Everywhere else, two moments and a plot are enough.
- Would you add skewness and kurtosis to a standard metrics dashboard?Usually not. They are noisy, hard for most readers to convert into an action, and diagnostic rather than operational. I would compute them once when characterising a metric, record the finding as a property of that metric, and instead surface a direct tail statement — how often the threshold is exceeded and by how much — which is what the decision actually consumes.
- How would you convince a team that a mean-plus-three-standard-deviations buffer is unsafe for their metric?By showing the empirical exceedance rate rather than arguing about moments. Count how often the historical data crossed that band and compare it to the roughly one-in-750 rate the rule implicitly assumes. A concrete count of breaches lands where a kurtosis figure does not, and the shape statistic then explains why.
- When does shape genuinely not matter to a decision?When the payoff is close to linear in the realised value and you are aggregating many draws — total cost over a quarter, or an average across a large population. There the mean carries the decision and the standard deviation tells you how precisely you know it. Heavy tails then affect only how much data you need, not the structure of the answer.
Two bridges rated for the same average and the same variability of load are not the same bridge if one occasionally sees a load ten times the average.
saying these in an interview costs you the question
- Treats mean and standard deviation as a complete summary
- Sizes every buffer at three standard deviations by habit
- Adds shape statistics everywhere without a decision behind them
- Assumes equal standard deviation means equal tail risk
- Characterises a heavy-tailed metric from a calm period only