What does a two-sample Kolmogorov-Smirnov test compare between two samples?
answer
- compares whole distributions, not one summary
- step function through the sorted data
- the largest vertical gap
- sup of the absolute CDF difference
basics
~20 sA two-sample Kolmogorov-Smirnov test compares the two samples' entire distributions. It builds a step-shaped empirical CDF for each sample and takes the largest vertical gap between them, so a difference in location, spread or shape can all trigger rejection.
solid answer
~50 sIt compares whole distributions rather than a single summary. For each sample you build the empirical CDF, `F_n(t)` = fraction of that sample at or below `t`, a step function rising from 0 to 1. The statistic is the largest vertical distance between the two step functions, `D = max_t |F_n(t) - G_m(t)|`. The null hypothesis is that both samples were drawn from the same continuous distribution; the alternative is simply that they were not, so a shift in the median, a change in spread, or a new bump in the shape can each produce a large `D`. Because the two curves only move at observed values, you compute `D` by merging the sorted samples and tracking the running difference. Rejection tells you the distributions differ somewhere - it does not tell you where or by how much, so always plot the two curves and report `D` alongside the p-value.
go deeper
Be ready to say in one sentence that it compares two full distributions through their empirical CDFs and that the statistic is the biggest vertical gap between the two step functions.
Explain how the statistic is actually computed by walking the merged sorted samples, why it is distribution-free under continuity, and how the critical value shrinks with sample size.
Show the diagnosis habit: after a significant result, plot both curves, locate the gap, quote it as a percentage-point difference, and say whether that gap sits where the business decision lives.
Own the tradeoff of an omnibus test. It buys assumption-free breadth at the cost of a vague alternative and no localisation, so decide up front whether the team needs a broad alarm or a targeted, interpretable comparison.
## The question the test answers Most comparison tests ask a narrow question: do these two groups have the same mean? the same variance? the same median? The two-sample Kolmogorov-Smirnov (KS) test asks the broad one: **could these two samples have come from the same distribution at all?** Nothing about the distribution is assumed - no normality, no equal variance, no particular shape. ## The empirical CDF Given a sample `x_1, ..., x_n`, the empirical cumulative distribution function is ``` F_n(t) = (number of observations <= t) / n ``` It is a step function: it starts at 0 to the left of the smallest value, jumps by `1/n` at each observation, and reaches 1 at the largest. It is the sample's own answer to the question 'what fraction of the data is at or below this value?' - a complete, lossless summary of the sample, unlike a mean or a variance. ## The statistic With two samples of sizes `n` and `m` and empirical CDFs `F_n` and `G_m`, ``` D = max over t of | F_n(t) - G_m(t) | ``` The two curves are flat between observed values, so the maximum gap must occur at one of the observed points. In practice you merge the two sorted samples, walk through them once, add `1/n` to a running total when the point came from sample one and subtract `1/m` when it came from sample two, and record the largest absolute value the running total reaches. A concrete reading: two builds of an API produce response-time samples. If `D = 0.18` at 400 ms, that means the fraction of requests finishing under 400 ms differs by 18 percentage points between the builds - the single worst disagreement anywhere along the range. ## Why one table of critical values works everywhere Under the null hypothesis, with a **continuous** underlying distribution, `D` is *distribution-free*: its sampling distribution depends only on `n` and `m`, not on the shared underlying distribution. The reason is that a strictly increasing transformation of the data leaves all the vertical gaps untouched - the ranks and hence the step positions are unchanged - so you can imagine both samples transformed to uniform values on the interval 0 to 1 without altering `D`. That is why one set of critical values serves any continuous data. The usual large-sample rule rejects at level `alpha` when ``` D > c(alpha) * sqrt( (n + m) / (n * m) ) ``` with `c(0.05)` about 1.36. Notice the shape: the threshold shrinks roughly like `1/sqrt(n)` when both samples grow, so bigger samples detect smaller gaps. ## What rejection does and does not tell you Rejection says: the distributions are not identical. It does **not** say which one is larger, where along the range they differ, or whether the difference matters. A significant `D` driven by a gap at the 30th percentile and one driven by a gap at the 90th percentile look identical in the p-value and mean completely different things operationally. This is why the honest report is a plot of the two empirical CDFs with the location of the maximum gap marked, plus `D` itself as an effect size on a 0-to-1 scale. A large p-value likewise does not prove the distributions are the same - it means the data did not provide enough evidence against sameness, which for small samples is a weak statement. ## Requirements and limits - **Ordered data.** The empirical CDF needs a meaningful 'less than or equal to', so the test applies to continuous or at least ordinal measurements, never to unordered category labels. - **Continuity.** The null distribution of `D` is derived assuming no exact repeats. Heavily repeated values violate that assumption and make the standard p-value untrustworthy. - **Independence.** Observations within each sample must be independent, and the two samples independent of each other. - **Where its power sits.** The unweighted maximum gap is largest where the empirical CDFs vary most, near the middle of the distribution; differences confined to the extreme tails move `D` very little. - **Sample size.** Because the threshold falls like `1/sqrt(n)`, very large samples will flag differences far too small to act on. ## One-sided variants Instead of the absolute gap you can take the signed one-sided maximum, `max_t (F_n(t) - G_m(t))`. That tests whether one distribution lies systematically above the other - one sample tending to produce smaller values than the other across the whole range. The two-sided absolute version is the default and is what most interview questions mean by 'the KS test'. ## How to present it in an interview Say what it compares (whole distributions, via empirical CDFs), give the statistic (largest vertical gap), state the null (same continuous distribution), and add the caveat that a significant result locates nothing - you still have to look at the curves.
- What exactly is the null hypothesis, and what does rejecting it license you to say?The null is that both samples were drawn from the same continuous distribution. Rejecting it licenses only the claim that they differ somewhere - not where, in which direction, or by how much. To say anything more you inspect the two empirical CDFs, note where the maximum gap sits, and quote the size of that gap.
- Why is the two-sample KS statistic called distribution-free?Under the null with continuous data, the sampling distribution of D depends only on the two sample sizes, not on the shared underlying distribution. Any strictly increasing transformation of the data leaves every vertical gap unchanged, so the same critical values apply to response times, prices or temperatures alike.
- How does sample size enter the decision rule?The 5% threshold is roughly 1.36 * sqrt((n + m) / (n * m)), which shrinks like one over the square root of sample size. Doubling both samples lets you detect a gap about 30% smaller. With very large samples the threshold becomes tiny, so trivial differences clear it.
Lay two staircases side by side, one built from each sample, each climbing from the floor to the ceiling. The statistic is the widest height difference between them at any point along the walk.
saying these in an interview costs you the question
- Says the KS test only detects a difference in means
- Claims a significant result shows where the distributions differ
- Runs it on unordered category labels with no natural ordering
- Reads D as a probability that the distributions differ
- Treats a large p-value as proof the distributions are identical