skip to content

Why does Anderson-Darling catch a tail-only difference that a Kolmogorov-Smirnov test misses?

level: seniorimportance: nice to knowfreq 26%

answer

  1. where each statistic puts its attention
  2. the achievable gap shrinks near 0 and 1
  3. an unweighted maximum lands mid-distribution
  4. weight is one over F(1-F)
  5. the weight explodes at both extremes

basics

~20 s

Kolmogorov-Smirnov takes an unweighted maximum CDF gap, and such gaps are naturally tiny in the tails, so its maximum lands near the middle. Anderson-Darling weights the squared gap by one over F(1-F), which explodes at the extremes.

solid answer

~50 s

The two statistics weight the same discrepancy differently. Kolmogorov-Smirnov uses `D = max_t |F_n(t) - F(t)|`, an unweighted maximum. Near the extremes the cumulative proportion is pinned close to 0 or 1, so the achievable gap there is intrinsically small and the maximum is nearly always attained near the centre - the test is close to blind beyond the 95th percentile. Anderson-Darling instead integrates the squared discrepancy with weight `1 / (F(t) * (1 - F(t)))`, which grows without bound as `F` approaches 0 or 1. A gap of the same absolute size out in the tail therefore contributes far more to `A^2` than the same gap at the median. If two response-time samples agree everywhere up to the 95th percentile and diverge only beyond it, Anderson-Darling can reject while KS shrugs. The cost is that a handful of extreme observations - the sparsest, noisiest part of the data - can drive the whole statistic.

go deeper

for a junior

Recall the headline: one statistic takes a plain maximum gap and is strongest in the middle, the other weights discrepancies so that the extreme ends of the range count much more heavily.

for a middle

Explain the weight one over F times one minus F, note that it is the reciprocal of the variance of the empirical CDF at that point, and say why the plain maximum almost always occurs near the centre.

for a senior

Show you choose the statistic from where the decision lives. If tail latency drives the call, say a non-significant plain-maximum result proves nothing about the tail and re-test with a tail-weighted statistic.

for a principal

Own the framing that a test statistic encodes what the organisation considers a meaningful difference. Decide which region of the distribution the business is actually exposed to before anyone picks a test.

## Two ways to score the same disagreement Both statistics start from the same picture: the empirical CDF `F_n` of the sample - the step function that jumps by `1/n` at each observation - laid against a comparison CDF `F`. They differ only in how they turn the vertical discrepancy `F_n(t) - F(t)` into a single number. **Kolmogorov-Smirnov** takes the unweighted supremum: ``` D = max over t of | F_n(t) - F(t) | ``` **Anderson-Darling** takes a weighted integral of the squared discrepancy: ``` A^2 = n * integral of [ (F_n(t) - F(t))^2 / (F(t) * (1 - F(t))) ] dF(t) ``` The whole difference in behaviour lives in that denominator. ## Why the unweighted maximum ignores the tails Think about how much room there is to disagree at a given point. At a value where the true cumulative proportion `F(t)` is 0.98, the empirical CDF is squeezed between 0 and 1 and, in any plausible sample, sits within a couple of percentage points of 0.98. The discrepancy there simply **cannot** be large in absolute terms. Statistically, the standard deviation of `F_n(t)` around `F(t)` is `sqrt(F(t)(1 - F(t)) / n)`, which is maximised at `F(t) = 0.5` and collapses toward zero at either extreme. So an unweighted maximum is structurally a middle-of-the-distribution statistic. Two distributions can differ dramatically in shape beyond the 95th percentile - one with a fat tail of slow requests, one that cuts off - and still produce a `D` no larger than ordinary sampling noise near the median. ## What the weight does The factor `1 / (F(1 - F))` is the reciprocal of that same variance term. It equals 4 at the median and rises without bound toward either end: at `F = 0.99` it is about 101, at `F = 0.999` about 1001. Anderson-Darling therefore measures each discrepancy **relative to how big a discrepancy could plausibly occur there**. A two-percentage-point gap at the 99th percentile is enormous on that scale; the same gap at the median is unremarkable. The consequence is a statistic whose sensitivity is spread across the whole range rather than concentrated in the centre - and, in practice, one that is markedly more sensitive to tail behaviour than KS. That is why it is the usual recommendation when the difference you care about lives in the extremes: the slow requests, the large losses, the rare big values. ## The squaring and the integral Two further design choices matter. Squaring rather than taking absolute values makes large local discrepancies count disproportionately. Integrating rather than maximising means a **persistent** moderate discrepancy across a wide region accumulates, whereas KS would report only the single worst point. So Anderson-Darling also picks up a broad, consistent shift that never produces one dramatic gap. ## The price 1. **Tail data is the sparsest and noisiest.** Multiplying it by a large weight amplifies noise as well as signal. In a sample of 500, the region beyond the 99th percentile holds about five observations; a single unusually extreme point can move `A^2` a lot. 2. **No interpretable effect size.** `D` reads directly as 'the largest difference in cumulative proportion' - a 0-to-1 number you can quote to a product owner. `A^2` has no such reading; it is a test statistic and nothing more. 3. **No localisation.** Neither test tells you where the difference is, but at least `D` comes with a specific point where the maximum occurred. 4. **Critical values are less portable.** Anderson-Darling's null distribution in the one-sample case depends on whether and which parameters were estimated, so the thresholds are procedure-specific. A two-sample form exists and is judged against its own null distribution. 5. **Sensitivity to contamination.** A few mis-recorded extreme values - a timeout written as a huge number, a sentinel value - can dominate the statistic where KS would barely notice. ## Choosing between them Ask where the difference that would change a decision lives. - If the decision hinges on the middle - the typical user, the median latency - KS is adequate, more robust to a few bad extreme values, and gives you a quotable effect size. - If the decision hinges on the extremes - tail latency, rare large costs, the worst 1% of outcomes - the unweighted maximum is the wrong instrument and Anderson-Darling is the better default. - Either way, look at the two empirical CDFs before and after the test. The picture tells you what the single number cannot: where the curves separate and by how much. ## A sanity habit When KS returns a comfortable non-significant result on data whose tails you care about, treat that as unproven rather than reassuring. Re-run with a tail-weighted statistic, or explicitly compare the high quantiles, before concluding the two samples behave the same where it counts.

  • What is the price you pay for that tail weighting?
    The tails hold the fewest observations, so weighting them heavily amplifies noise and makes the statistic sensitive to a handful of extreme or mis-recorded values. You also lose an interpretable effect size: the maximum-gap statistic reads directly as a difference in cumulative proportion, while the weighted integral is only a test statistic.
  • Which statistic would you prefer when the difference that matters sits near the median?
    The unweighted maximum gap. It is most sensitive exactly there, it is more robust to a few bad extreme values, and its value doubles as a plain effect size you can quote - the largest difference in the fraction of observations at or below a given value.
  • Does the squaring and integration change anything beyond tail sensitivity?
    Yes. Integrating accumulates a moderate discrepancy that persists over a wide region, whereas a maximum reports only the single worst point. So a broad consistent offset that never produces one dramatic gap registers on the integrated statistic but may be invisible to the maximum.

Judging two runners only by their largest gap in metres favours the middle of the race, where they can drift apart freely. Weighting by how much drift is even possible at each point makes a small gap at the finish line count as heavily as a big one at halfway.

saying these in an interview costs you the question

  • Says the maximum-gap statistic is equally sensitive everywhere
  • Thinks the tail weight is a mere constant multiplier
  • Treats the Anderson-Darling value as an interpretable effect size
  • Picks the tail-weighted test with a few dozen observations
  • Claims either statistic tells you where the distributions differ

context