skip to content

Distribution Comparison Tests

Testing whether two samples share a distribution, not just a mean: the Kolmogorov-Smirnov statistic on empirical CDFs, Anderson-Darling, binned chi-square. Drift questions rest on it.

on this pageshow

questions

5

What does a two-sample Kolmogorov-Smirnov test compare between two samples?

level: juniorimportance: must knowfreq 58%

answer

  1. compares whole distributions, not one summary
  2. step function through the sorted data
  3. the largest vertical gap
  4. sup of the absolute CDF difference

basics

~20 s

A two-sample Kolmogorov-Smirnov test compares the two samples' entire distributions. It builds a step-shaped empirical CDF for each sample and takes the largest vertical gap between them, so a difference in location, spread or shape can all trigger rejection.

solid answer

~50 s

It compares whole distributions rather than a single summary. For each sample you build the empirical CDF, `F_n(t)` = fraction of that sample at or below `t`, a step function rising from 0 to 1. The statistic is the largest vertical distance between the two step functions, `D = max_t |F_n(t) - G_m(t)|`. The null hypothesis is that both samples were drawn from the same continuous distribution; the alternative is simply that they were not, so a shift in the median, a change in spread, or a new bump in the shape can each produce a large `D`. Because the two curves only move at observed values, you compute `D` by merging the sorted samples and tracking the running difference. Rejection tells you the distributions differ somewhere - it does not tell you where or by how much, so always plot the two curves and report `D` alongside the p-value.

go deeper

for a junior

Be ready to say in one sentence that it compares two full distributions through their empirical CDFs and that the statistic is the biggest vertical gap between the two step functions.

for a middle

Explain how the statistic is actually computed by walking the merged sorted samples, why it is distribution-free under continuity, and how the critical value shrinks with sample size.

for a senior

Show the diagnosis habit: after a significant result, plot both curves, locate the gap, quote it as a percentage-point difference, and say whether that gap sits where the business decision lives.

for a principal

Own the tradeoff of an omnibus test. It buys assumption-free breadth at the cost of a vague alternative and no localisation, so decide up front whether the team needs a broad alarm or a targeted, interpretable comparison.

## The question the test answers Most comparison tests ask a narrow question: do these two groups have the same mean? the same variance? the same median? The two-sample Kolmogorov-Smirnov (KS) test asks the broad one: **could these two samples have come from the same distribution at all?** Nothing about the distribution is assumed - no normality, no equal variance, no particular shape. ## The empirical CDF Given a sample `x_1, ..., x_n`, the empirical cumulative distribution function is ``` F_n(t) = (number of observations <= t) / n ``` It is a step function: it starts at 0 to the left of the smallest value, jumps by `1/n` at each observation, and reaches 1 at the largest. It is the sample's own answer to the question 'what fraction of the data is at or below this value?' - a complete, lossless summary of the sample, unlike a mean or a variance. ## The statistic With two samples of sizes `n` and `m` and empirical CDFs `F_n` and `G_m`, ``` D = max over t of | F_n(t) - G_m(t) | ``` The two curves are flat between observed values, so the maximum gap must occur at one of the observed points. In practice you merge the two sorted samples, walk through them once, add `1/n` to a running total when the point came from sample one and subtract `1/m` when it came from sample two, and record the largest absolute value the running total reaches. A concrete reading: two builds of an API produce response-time samples. If `D = 0.18` at 400 ms, that means the fraction of requests finishing under 400 ms differs by 18 percentage points between the builds - the single worst disagreement anywhere along the range. ## Why one table of critical values works everywhere Under the null hypothesis, with a **continuous** underlying distribution, `D` is *distribution-free*: its sampling distribution depends only on `n` and `m`, not on the shared underlying distribution. The reason is that a strictly increasing transformation of the data leaves all the vertical gaps untouched - the ranks and hence the step positions are unchanged - so you can imagine both samples transformed to uniform values on the interval 0 to 1 without altering `D`. That is why one set of critical values serves any continuous data. The usual large-sample rule rejects at level `alpha` when ``` D > c(alpha) * sqrt( (n + m) / (n * m) ) ``` with `c(0.05)` about 1.36. Notice the shape: the threshold shrinks roughly like `1/sqrt(n)` when both samples grow, so bigger samples detect smaller gaps. ## What rejection does and does not tell you Rejection says: the distributions are not identical. It does **not** say which one is larger, where along the range they differ, or whether the difference matters. A significant `D` driven by a gap at the 30th percentile and one driven by a gap at the 90th percentile look identical in the p-value and mean completely different things operationally. This is why the honest report is a plot of the two empirical CDFs with the location of the maximum gap marked, plus `D` itself as an effect size on a 0-to-1 scale. A large p-value likewise does not prove the distributions are the same - it means the data did not provide enough evidence against sameness, which for small samples is a weak statement. ## Requirements and limits - **Ordered data.** The empirical CDF needs a meaningful 'less than or equal to', so the test applies to continuous or at least ordinal measurements, never to unordered category labels. - **Continuity.** The null distribution of `D` is derived assuming no exact repeats. Heavily repeated values violate that assumption and make the standard p-value untrustworthy. - **Independence.** Observations within each sample must be independent, and the two samples independent of each other. - **Where its power sits.** The unweighted maximum gap is largest where the empirical CDFs vary most, near the middle of the distribution; differences confined to the extreme tails move `D` very little. - **Sample size.** Because the threshold falls like `1/sqrt(n)`, very large samples will flag differences far too small to act on. ## One-sided variants Instead of the absolute gap you can take the signed one-sided maximum, `max_t (F_n(t) - G_m(t))`. That tests whether one distribution lies systematically above the other - one sample tending to produce smaller values than the other across the whole range. The two-sided absolute version is the default and is what most interview questions mean by 'the KS test'. ## How to present it in an interview Say what it compares (whole distributions, via empirical CDFs), give the statistic (largest vertical gap), state the null (same continuous distribution), and add the caveat that a significant result locates nothing - you still have to look at the curves.

  • What exactly is the null hypothesis, and what does rejecting it license you to say?
    The null is that both samples were drawn from the same continuous distribution. Rejecting it licenses only the claim that they differ somewhere - not where, in which direction, or by how much. To say anything more you inspect the two empirical CDFs, note where the maximum gap sits, and quote the size of that gap.
  • Why is the two-sample KS statistic called distribution-free?
    Under the null with continuous data, the sampling distribution of D depends only on the two sample sizes, not on the shared underlying distribution. Any strictly increasing transformation of the data leaves every vertical gap unchanged, so the same critical values apply to response times, prices or temperatures alike.
  • How does sample size enter the decision rule?
    The 5% threshold is roughly 1.36 * sqrt((n + m) / (n * m)), which shrinks like one over the square root of sample size. Doubling both samples lets you detect a gap about 30% smaller. With very large samples the threshold becomes tiny, so trivial differences clear it.

Lay two staircases side by side, one built from each sample, each climbing from the floor to the ceiling. The statistic is the widest height difference between them at any point along the walk.

saying these in an interview costs you the question

  • Says the KS test only detects a difference in means
  • Claims a significant result shows where the distributions differ
  • Runs it on unordered category labels with no natural ordering
  • Reads D as a probability that the distributions differ
  • Treats a large p-value as proof the distributions are identical

context

open as a page

Why are standard Kolmogorov-Smirnov critical values wrong when the reference normal's mean and SD come from the same sample?

level: middleimportance: should knowfreq 34%

basics

~20 s

Fitting the reference curve to the same data pulls it toward the sample, shrinking the maximum gap. Standard tables assume a reference fixed in advance, so p-values come out too large and the test under-rejects.

open as a page

How does a Kolmogorov-Smirnov test behave on integer star ratings with heavy ties?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A Kolmogorov-Smirnov test is unreliable on heavily tied data: its null distribution assumes continuous measurements with no exact repeats. With five rating values the empirical CDFs jump in blocks and the classical p-value comes out too conservative.

open as a page

A KS test on 1,000,000 sessions returns p below 1e-16 for a shift nobody would notice - how do you decide whether to act?

level: principalimportance: should knowfreq 36%

basics

~20 s

At a million observations the KS rejection threshold shrinks toward zero, so any real difference becomes significant. Judge the magnitude instead: read the statistic D as an effect size and compare it against a materiality threshold agreed beforehand.

open as a page

Why does Anderson-Darling catch a tail-only difference that a Kolmogorov-Smirnov test misses?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Kolmogorov-Smirnov takes an unweighted maximum CDF gap, and such gaps are naturally tiny in the tails, so its maximum lands near the middle. Anderson-Darling weights the squared gap by one over F(1-F), which explodes at the extremes.

open as a page