skip to content

For a median from n = 20, how does an order-statistic interval get its coverage?

level: seniorimportance: nice to knowfreq 16%

answer

  1. trade a value question for a count question
  2. how many points fall below the median
  3. that count is binomial with p one half
  4. shape of the data never enters
  5. sixth and fifteenth of twenty sorted values

basics

~20 s

Count how many observations fall below the population median: for continuous data that count is Binomial(20, 0.5) whatever the distribution. The 6th and 15th sorted values fail only when the count leaves 6 to 14, giving about 95.9% coverage.

solid answer

~50 s

Sort the 20 observations and let `B` be the number of them below the true population median `m`. For any continuous distribution each observation is below `m` with probability exactly 0.5 and independently of the others, so `B` follows a Binomial(20, 0.5) distribution - the shape of the data never enters, which is what makes the method distribution-free. The interval `[X_(6), X_(15)]` misses `m` only if fewer than 6 observations fall below it or more than 14 do, so its coverage is `P(6 <= B <= 14) = 1 - 2 * P(B <= 5)`, about 0.959. The next interval in, `[X_(7), X_(14)]`, drops to about 0.885. So you cannot hit exactly 95% - coverage comes in discrete jumps and you take the conservative side. The price is that the endpoints are always observed data values and the achievable levels are lumpy; the payoff is no density estimate, no normality assumption, and exact coverage.

go deeper

for a junior

Know that a confidence interval for a median can be read off two sorted values of the sample, and that this needs no assumption about the shape of the data.

for a middle

Be able to explain the count argument: the number of observations below the true median is Binomial(n, 0.5) for continuous data, and interval coverage is the probability that count stays within the chosen ranks.

for a senior

Show that you know the practical limits - discrete achievable levels, endpoints pinned to observed values, and the failure of the exactness guarantee under heavy ties or rounding.

for a principal

Own the tradeoff between assumption-free and efficient. Decide when the organisation should pay the width penalty for exact distribution-free coverage rather than lean on a parametric interval nobody has validated.

## The trick in one sentence Turn a question about a value into a question about a count, because the count has a distribution you know exactly. ## Setting it up Take 20 independent observations from a continuous distribution and sort them: `X_(1) <= X_(2) <= ... <= X_(20)`. These sorted values are the order statistics. Let `m` be the unknown population median, defined by `P(X < m) = 0.5`. Now define `B` as the number of observations that land below `m`. Each observation is below `m` with probability 0.5 by the definition of the median, and the observations are independent, so `B ~ Binomial(20, 0.5)` Notice what is *not* in that statement: any mention of the shape, spread, skewness or density of the underlying distribution. That is the whole point. Continuity (so ties have probability zero) and independence are the only assumptions. ## From a count to an interval Consider the candidate interval `[X_(r), X_(s)]` with `r < s`. When does it fail to contain `m`? - It sits entirely above `m` when fewer than `r` observations are below `m`, that is `B <= r - 1`. - It sits entirely below `m` when at least `s` observations are below `m`, that is `B >= s`. So the coverage probability is `P(r <= B <= s - 1)`. Choose the endpoints symmetrically, `s = n + 1 - r`, and this becomes `P(r <= B <= n - r)`. ## Doing the arithmetic for n = 20 Take `r = 6`, so `s = 15` and the interval is `[X_(6), X_(15)]`. Coverage is `P(6 <= B <= 14) = 1 - 2 * P(B <= 5)` by the symmetry of the Binomial(20, 0.5) distribution. `P(B <= 5)` sums the binomial coefficients 1, 20, 190, 1140, 4845 and 15504, giving 21700 out of `2^20 = 1048576`, or about 0.0207. Coverage is therefore about `1 - 0.0414 = 0.959`. Step one rank inward, `r = 7`, and the interval `[X_(7), X_(14)]` covers with probability about 0.885. There is nothing between 0.885 and 0.959, so at n = 20 there is no interval of this family with exactly 95% coverage. The convention is to take the conservative one: 95.9% coverage from the 6th and 15th sorted values. ## Why the coverage is exact, not asymptotic Every step above used a finite binomial probability. No central limit theorem was invoked, no density was estimated, no normal approximation was made. The stated coverage holds at n = 20 exactly, which is unusual and valuable - most small-sample intervals only promise approximate coverage. This is the clean counterpoint to the fact that a sample quantile's standard error depends on an unknown density: here you never need that density, because you traded the value question for a count question. ## The costs **Discreteness.** Achievable levels jump. You cannot dial in 95% precisely, and at very small n the gaps are large - the family may offer only a few levels at all. **Endpoints are data points.** The interval can only start and end at observed values, so it inherits their granularity and cannot extend beyond the sample range. **Ties break it.** The Binomial(20, 0.5) argument assumes `P(X = m) = 0` . With heavily rounded or discrete data, observations land exactly on the median and the probability of being strictly below is no longer 0.5, so the stated coverage is wrong - typically conservative, but no longer exact. **Efficiency.** Because it makes no assumptions, it cannot exploit any. When you genuinely know the distributional family, a parametric interval will be narrower at the same level. ## Generalising past the median The same construction works for any quantile `q_p`: the count below `q_p` is Binomial(n, p), and you pick ranks whose binomial tail probabilities sum to your target error. For extreme `p` and small `n` the achievable ranks run off the end of the sample, which is another way of seeing that a p99 from a few hundred points is barely supported by data at all. ## The one-line version Rank order carries distribution-free information: how many points beat the true median is binomial no matter what the data look like, and that count is enough to build an exact interval from two sorted values.

  • Why can't you get exactly 95% coverage from this construction at n = 20?
    Because the coverage is a sum of binomial point probabilities and moving an endpoint changes it in a discrete jump. The interval `[X_(6), X_(15)]` gives about 95.9% and the next one in, `[X_(7), X_(14)]`, gives about 88.5%. Nothing lies between, so you take the conservative 95.9%.
  • What happens to this method when the metric is heavily rounded?
    It stops being exact. The binomial argument assumes an observation lands exactly on the population median with probability zero, so that being below has probability 0.5. With ties that is no longer true, the effective probability shifts, and the stated coverage no longer holds - usually erring conservative, but it is no longer a guarantee you can quote.
  • How would you extend the idea to a p90 rather than a median?
    Replace the fair coin with a biased one: the count of observations below the population p90 is Binomial(n, 0.9). Pick ranks whose lower and upper binomial tail probabilities sum to your error budget. For extreme quantiles and small n the required ranks run past the ends of the sample, which is the method telling you there is not enough data.

It is like calling a coin flip for each observation - above or below the true median - and reading the interval off how lopsided twenty flips are allowed to get.

saying these in an interview costs you the question

  • Thinks the method assumes normality somewhere
  • Believes it needs an estimate of the density at the median
  • Expects an exact 95% level to always be achievable
  • Applies it unchanged to heavily tied or discrete data
  • Calls the coverage asymptotic rather than exact

context