skip to content

Why does a sample median's precision depend on distribution shape, not just on n?

level: seniorimportance: should knowfreq 32%

answer

  1. sqrt(n) is only half the story
  2. shape enters through a local quantity
  3. think spacing between neighbouring sorted values
  4. spacing is one over the density
  5. peaked centre pins it, hollow centre lets it wander

basics

~20 s

A sample median's standard error is about 1 divided by (2 times the density at the median times sqrt(n)). A peaked centre pins it down; a flat or hollow centre lets it wander at the same n.

solid answer

~50 s

The asymptotic standard error of a sample median is `1 / (2 * f(m) * sqrt(n))`, where `f(m)` is the probability density at the population median. Sample size enters only through `sqrt(n)`; everything else is shape. If the distribution is sharply peaked at its centre, many observations crowd around the median, so moving the middle observation up or down by one rank barely changes the value, and the estimate is tight. If the centre is a flat plateau - or worse, a valley between two modes, as in a latency metric with a cache-hit cluster and a cache-miss cluster - the density there is low, neighbouring order statistics are far apart in value, and the median swings across samples. It is also why neither estimator is universally more precise: under normality the median's variance is about 1.57 times the mean's, but on a sharply peaked heavy-tailed distribution the median wins.

go deeper

for a junior

Know that a median computed from a sample varies from sample to sample, and that robustness to outliers is not the same thing as being precisely estimated. Being able to say shape matters is enough at this level.

for a middle

Be able to state that the median's standard error is one over twice the density at the median times sqrt(n), and explain the spacing intuition: high density means tightly packed sorted values, so the middle one moves little.

for a senior

Demonstrate the diagnostic habit - inspect the distribution near the centre, check for a plateau, a hollow between modes, or heavy ties, and refuse to quote a stable median when a split-half recomputation moves it.

for a principal

Own the metric definition itself. Decide when a median is the wrong summary for a multimodal population and push for segmentation or an explicitly tail-aware metric rather than defending a central number nobody can reproduce.

## The formula, and what each piece does For independent draws from a distribution with density `f` continuous and positive at the population median `m`, the sample median is asymptotically normal with standard error `SE = 1 / (2 * f(m) * sqrt(n))` This is the general sample-quantile result at `p = 0.5`: `sqrt(p(1-p)/n) / f(q_p)` with `p(1-p) = 0.25`, whose square root is 0.5, giving the 1/2 in the numerator. Two levers set the precision. `sqrt(n)` is the one everybody expects: quadruple the data, halve the uncertainty. `f(m)` is the one candidates miss. It is the height of the density curve exactly where the median sits, and it is pure shape - it has nothing to do with how many observations you collected. ## Why the density controls it Think about what the sample median is: the middle value of the sorted sample. Perturb the sample slightly and the middle *rank* moves by a step or two. How much the *value* moves depends on how far apart consecutive sorted observations are near the centre - the local spacing. High density means points are packed tightly there, so a rank step is a tiny value step, and the estimator barely moves. Low density means the sorted values are strung far apart near the middle, so the same rank step is a large jump in value. Density and spacing are reciprocals of each other: the expected gap between neighbouring order statistics around `m` is about `1 / (n * f(m))`. That reciprocal is exactly why `f(m)` lands in the denominator of the standard error. ## The concrete cases **Sharply peaked centre.** A metric that clusters tightly around a typical value has a high `f(m)`. Its median is precise from a modest sample, even if the distribution has an ugly tail - tail observations change ranks but not which value sits in the middle. **Flat plateau.** A near-uniform distribution over a wide range has a low, constant density. The median is still unbiased but it wanders, because a large stretch of values is roughly equally likely to be the middle one. **Bimodal with a hollow centre.** This is the pathological case worth naming in an interview. Suppose request latency has one cluster of fast responses and one cluster of slow responses, with few requests in between. The population median falls in the sparse valley. `f(m)` is small, so the standard error is large, and across samples the median can flip between the neighbourhoods of the two modes. Reporting "the median is stable at 40ms" from one sample is then close to meaningless, and worse, the median is a poor summary of a distribution that has no typical value. ## Mean versus median, without the folklore Candidates often recite "the median is more robust, so use it". Robust to contamination, yes; more *precise*, not always. Under a normal distribution the mean's variance is `sigma^2 / n` while the median's is `pi * sigma^2 / (2n)`, about 1.57 times larger - the median throws away information the normal distribution puts in the tails, and its asymptotic relative efficiency is `2/pi`, roughly 0.64. Flip to a sharply peaked, heavy-tailed distribution and the ordering reverses: the mean's variance is inflated by rare extreme values while the median sits on a high-density centre and is estimated well. The right answer to "mean or median" is therefore about the shape of the distribution and the estimand you actually care about, not a slogan. ## What this changes in practice Before trusting a median from a fixed sample size, look at the histogram near the centre. Ask three things: is there a peak or a plateau there, is the metric heavily tied or rounded (ties break the continuity the formula assumes), and is the distribution multimodal in a way that makes a single central summary the wrong tool. A precision claim that quotes only `n` and never mentions shape is incomplete for a median in a way it would not be for a mean. ## The one-line version A median is estimated well when the data crowd around it; the same `n` buys you far less precision when the centre of the distribution is flat or hollow.

  • Under a normal distribution, which is estimated more precisely, the mean or the median?
    The mean. Its variance is `sigma^2/n`; the median's is about `pi/2` times that, roughly 1.57 times larger, an asymptotic relative efficiency of about 0.64. The median only wins when the distribution has enough tail mass to inflate the mean's variance, or enough contamination that the mean is estimating the wrong thing.
  • How would you spot the bimodal failure case before quoting a median?
    Plot the distribution rather than summary numbers. Look for a low-density region where the median falls, and check stability by recomputing the median on random halves of the data - if it flips between two neighbourhoods, the median is sitting in a valley and no single central summary describes the metric. Report the two modes or segment the population instead.
  • Does this density term also affect the precision of a p90 or p99?
    Yes, and usually more severely. The general form is `sqrt(p(1-p)/n) / f(q_p)`, so any quantile is governed by the density where it falls. In a right-skewed metric the density thins out as you move up the tail, so p90 is less precisely estimated than the median and p99 much less so at the same sample size.

Finding the middle of a crowd is easy when everyone is packed shoulder to shoulder and hard when the middle of the room is empty and the people are split against two walls.

saying these in an interview costs you the question

  • Says precision depends only on sample size
  • Claims the median is always more precise because it is robust
  • Confuses robustness to outliers with low sampling variance
  • Quotes a median from a clearly bimodal metric without comment
  • Thinks the median's standard error is the standard deviation over sqrt(n)

context