skip to content

Distribution Shape & Outliers

Reading the whole shape instead of one number: histograms and the empirical CDF, quantiles, skew and heavy tails, and the rules that flag a point as extreme. Interviewers ask what an average hides.

on this pageshow

explore

questions

15

Why does one extreme salary shift a dataset's mean but leave its median almost unchanged?

level: juniorimportance: must knowfreq 82%

answer

  1. sum versus position
  2. each value enters the average in full
  3. the middle rank hardly moves
  4. breakdown point: 0% versus 50%

basics

~20 s

The mean adds every value in full, so one enormous salary drags the average up. The median depends only on which value sits in the middle position, so an extreme point changes the ordering but barely moves the centre.

solid answer

~50 s

The mean gives every observation weight `1/n`, so raising one value by `d` raises the mean by exactly `d/n` — with no ceiling on `d`, a single point can push the mean anywhere. The median depends only on rank: making the largest value ten times larger reshuffles nothing below the middle, so the middle observation stays the same ordinary number. Formally this is the breakdown point — the smallest fraction of the sample you must corrupt to send the estimator to infinity. It is `1/n` (effectively 0%) for the mean and 50% for the median. The classic illustration: when a billionaire walks into a bar, the mean net worth of everyone inside becomes absurd while the median is still roughly what the regulars have. Reporting both, and looking hard when they disagree, is the practical habit.

go deeper

for a junior

Be ready to state plainly that the mean uses every value's size while the median only uses position, and to give a two-line numeric example where they diverge.

for a middle

Explain the mechanism with the arithmetic: raising one point by d moves the mean by d/n and moves the median not at all. Name the breakdown point and give both numbers.

for a senior

Show judgment about which summary to report and when. Interviewers expect you to note that a mean-median gap is a signal to investigate, not an automatic instruction to switch statistics.

for a principal

Own the tradeoff explicitly: robustness costs efficiency, and the mean is irreplaceable when the quantity of interest is a total. Be able to argue when a team should standardise on which summary and why.

## Two definitions of a centre Given `n` numbers, the arithmetic mean is `(x1 + x2 + ... + xn) / n`. The median is what you get after sorting: the middle observation when `n` is odd, and the average of the two middle observations when `n` is even. Both claim to describe a typical value, and on clean symmetric data they agree closely. Under a single extreme value they behave completely differently, and the reason is structural rather than accidental. ## Why the mean moves Every observation enters the mean with weight `1/n`. Increase one observation by an amount `d` and the mean increases by exactly `d/n`. Because `d` is unbounded, the mean is unbounded: no matter how large `n` is, there exists a single value large enough to move the mean wherever you want. Concretely, take eleven salaries — ten at 40,000 and one at 10,000,000. The sum is 10,400,000, so the mean is about 945,000: a number that describes nobody in the room. The median is the sixth value in sorted order, which is 40,000, and it describes ten of the eleven people. ## Why the median does not The median is a function of ranks, not of magnitudes. Replacing the largest value with a number a thousand times larger leaves it the largest value; the sorted order below the middle is untouched, so the middle observation is unchanged. To move the median you must change which observation occupies the middle, and that requires corrupting enough points to push the middle position past a genuinely different value. ## Breakdown point The breakdown point of an estimator is the smallest fraction of the sample that an adversary must replace with arbitrary values in order to drive the estimator arbitrarily far from the truth. It is the standard one-number summary of robustness. - Mean: `1/n`, which tends to 0 as the sample grows. One bad value is always enough. A larger sample does not help — it only makes each individual point cheaper to spot, not less damaging in aggregate. - Median: 50%. You must corrupt half the data before the median can be sent anywhere. - Trimmed mean with a fraction removed from each tail: breakdown point equal to that trimming fraction, which is why it sits between the two. A common misconception is that a large sample makes the mean safe. It does not: contamination in real data usually scales with the data, so a million-row table typically contains far more bad rows than a thousand-row one. ## What robustness costs The median is not free. On clean data drawn from a normal distribution, the sample median is a less precise estimate of the centre than the mean: its asymptotic relative efficiency is `2/pi`, about 0.64. In practical terms the median needs roughly 57% more observations to match the mean's precision. That is a real price, and it is why the mean remains the default on well-behaved data. The tradeoff is the whole point: you buy protection against contamination with precision on clean data. There is also a use case the median simply cannot serve. The mean multiplies back into a total: mean salary times headcount is the payroll bill. The median has no such property. If the quantity you care about is a sum — total spend, total load, total revenue — the mean is the correct summary even when it is unrepresentative, and the honest move is to report it alongside the median rather than to substitute one for the other. ## How to use this in practice Compute both and compare. A large gap between mean and median is a flag, not a diagnosis: it can mean genuine asymmetry in the population, or it can mean a handful of corrupt records. Distinguishing the two requires looking at the actual extreme values and where they came from. What the comparison does buy you is the knowledge that a single-number summary is about to mislead someone, which is usually the moment to show a distribution instead of a statistic. ## What interviewers listen for They want the mechanism (weight `1/n` versus rank), the vocabulary (breakdown point, 0% versus 50%), and the honesty to name the cost of robustness rather than declaring the median universally better.

  • What does it cost you to report the median instead of the mean on clean, symmetric data?
    Precision. On data from a normal distribution the sample median's asymptotic relative efficiency against the mean is `2/pi`, roughly 0.64, so you need about 57% more observations to estimate the centre as tightly. Robustness is bought with sample size. On clean, symmetric data the mean is the better estimator; the median only wins once contamination or heavy tails are plausible.
  • How many observations must an adversary corrupt to make the sample median arbitrarily large?
    About half of them. Until half the values have been replaced, the middle position in sorted order is still occupied by an uncorrupted observation, so the median stays inside the range of the real data. That 50% figure is the median's breakdown point, and it is the maximum any sensible location estimator can achieve — beyond half, the contaminated values are the majority and no estimator can tell which half is real.
  • Does dropping the single largest value make the mean robust?
    No. Dropping one maximum gives a breakdown point of `1/n`, which still tends to 0 — two bad values defeat it, and contaminated data rarely arrives one point at a time. It also makes the estimator asymmetric and discards a genuine extreme just as readily as an error. If you want a tunable middle ground, trim a fixed fraction from each tail instead, which buys a breakdown point equal to that fraction.

The mean is a seesaw that must physically balance every passenger, so one very heavy person tips it however far they like. The median just asks who is standing in the middle of the queue.

saying these in an interview costs you the question

  • Says the median is better than the mean in all situations
  • Claims a large enough sample makes the mean robust to outliers
  • Thinks the median works by deleting the extreme values
  • Confuses the median with the mode or the midrange
  • Cannot say what robustness costs on clean data

context

open as a page

What is the five-number summary of a dataset, and what does it tell you that the mean hides?

level: juniorimportance: must knowfreq 78%

basics

~20 s

The five-number summary is the minimum, first quartile, median, third quartile and maximum. Together they show where the bulk of the data sits and how far each tail reaches, which a single average cannot express.

open as a page

What does the sign of sample skewness tell you about the shape of a distribution?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Positive sample skewness means the long tail stretches to the right, toward large values. Negative means the long tail stretches to the left. A value near zero indicates a roughly symmetric shape, like a normal bell.

open as a page

How does a boxplot's 1.5 x IQR rule decide that a data point is an outlier?

level: middleimportance: must knowfreq 66%

basics

~20 s

It takes the first and third quartiles Q1 and Q3, sets IQR = Q3 - Q1, and marks any point below Q1 - 1.5 x IQR or above Q3 + 1.5 x IQR. Those cutoffs are called the fences.

open as a page

Why do latency SLOs report p95 and p99 instead of the mean response time?

level: middleimportance: must knowfreq 70%

basics

~20 s

Mean latency averages away the slow requests that users actually notice. The 95th and 99th percentiles report the experience of the worst 5 percent and 1 percent of requests, which is where dissatisfaction, timeouts and retries live.

open as a page

What does a positive excess kurtosis tell you about a distribution's tails?

level: middleimportance: must knowfreq 63%

basics

~20 s

Positive excess kurtosis means heavier tails than a normal distribution with the same standard deviation: extreme deviations happen more often than normal probabilities predict. Excess kurtosis subtracts 3, the kurtosis of a normal, so a normal scores exactly 0.

open as a page

Why can a plain z-score miss an outlier that a MAD-based modified z-score flags?

level: middleimportance: should knowfreq 48%

basics

~20 s

The plain z-score divides by the sample standard deviation, which the outlier itself inflates, so the point hides behind the spread it created. The modified z-score uses the median and the median absolute deviation, neither of which the outlier can move.

open as a page

How can a histogram's bin width change the shape the data appear to have?

level: middleimportance: should knowfreq 54%

basics

~20 s

Bin width is a smoothing dial the analyst chooses. Wide bins merge separate clusters into one hump; narrow bins turn random sampling noise into spurious spikes and gaps. Where the bin edges start can also split or merge a cluster.

open as a page

Why does taking logs of right-skewed values like house prices reduce the skewness?

level: middleimportance: should knowfreq 48%

basics

~20 s

The logarithm compresses large values far more than small ones, so a long right tail is pulled in toward the body of the data. Applied to strictly positive right-skewed values such as house prices, it can bring skewness close to zero.

open as a page

A sensor logs occasional 250C readings; how do you decide if they are contamination or genuine extremes?

level: seniorimportance: should knowfreq 56%

basics

~20 s

No statistical rule can answer this. Check provenance: the instrument's rated range, sentinel codes, unit mix-ups and an independent source. A statistic says a value is unusual; only evidence about its origin says it is wrong.

open as a page

Why do two reporting pipelines compute different 90th percentiles from the same 10 numbers?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Because there is no single definition of a sample quantile. With 10 sorted values the requested rank falls between two observations, and the competing conventions either take a specific observation or interpolate between neighbours, which yields different numbers from identical data.

open as a page

Why are sample skewness and kurtosis unreliable estimates on a small sample?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Both statistics raise deviations to the third or fourth power, so a handful of extreme points dominate them. Their standard errors are large on small samples, roughly sqrt(6/n) for skewness and sqrt(24/n) for excess kurtosis under normality.

open as a page

What changes when you trim 10% from each tail instead of winsorising at the 5th and 95th percentiles?

level: seniorimportance: nice to knowfreq 34%

basics

~10 s

Trimming removes the extreme observations, shrinking the sample. Winsorising keeps every row but replaces extreme values with the percentile cutoffs, so the count is unchanged and ties pile up at both caps.

open as a page

What does the bandwidth of a kernel density estimate control, and how does it mislead?

level: seniorimportance: nice to knowfreq 26%

basics

~10 s

Bandwidth sets how wide the smooth bump placed at each observation is, so it controls smoothness. Too large flattens two genuine humps into one bell; too small turns individual observations into spurious peaks.

open as a page

When do skewness and kurtosis change a decision that mean and standard deviation alone would drive?

level: principalimportance: nice to knowfreq 29%

basics

~20 s

Whenever the decision is sensitive to rare large values rather than to typical ones. Two distributions can share a mean and a standard deviation while differing sharply in their third and fourth standardised moments, so a buffer or threshold sized on the first two moments alone can be badly wrong.

open as a page