skip to content

How does a Kolmogorov-Smirnov test behave on integer star ratings with heavy ties?

level: seniorimportance: should knowfreq 40%

answer

  1. the theory assumes a continuous distribution
  2. exact repeats should have probability zero
  3. the step function jumps in large blocks
  4. the null now depends on the tie pattern
  5. p-values come out too large

basics

~20 s

A Kolmogorov-Smirnov test is unreliable on heavily tied data: its null distribution assumes continuous measurements with no exact repeats. With five rating values the empirical CDFs jump in blocks and the classical p-value comes out too conservative.

solid answer

~50 s

The KS null distribution is derived for continuous data, where the chance of two observations being exactly equal is zero. Integer star ratings violate that outright: each empirical CDF is a step function with at most five jumps, so `D` can only take a handful of values and its true null distribution now depends on the shared rating proportions - it is no longer distribution-free. Judged against the classical critical values the test is generally **conservative**: reported p-values are too large, so it under-rejects and wastes power. Two clean fixes exist. Build the null yourself by permutation - pool the two samples, reshuffle the group labels many times, recompute `D` each time, and take the p-value as the fraction of shuffled statistics at least as large as the observed one. That preserves the tie pattern exactly and assumes no continuity. Or treat the five ratings as ordered cells and compare the two samples' cell counts against the counts expected under a common distribution.

go deeper

for a junior

Know the precondition: the test assumes continuous measurements where exact repeats essentially never happen, so a handful of possible integer values is outside what it was built for.

for a middle

Explain what breaks - the step functions jump in a few large blocks, the statistic takes only a few values, and its null distribution now depends on the shared proportions instead of just the sample sizes.

for a senior

Show the repair and the direction of the error. State that the classical p-value is conservative, then build the null by permuting group labels or compare the rating cell counts directly.

for a principal

Own the standard: when a metric is inherently discrete, the team should default to resampling or an explicitly categorical comparison rather than a continuous-theory test with a footnote nobody reads.

## Where the assumption bites Everything convenient about Kolmogorov-Smirnov rests on one assumption: the underlying distribution is **continuous**, so no two observations are ever exactly equal. That assumption is what makes the statistic distribution-free - transform the data by any strictly increasing function and every vertical gap is unchanged, so one table of critical values covers all continuous data. Integer star ratings break it completely. A rating takes one of five values; in a sample of 2,000 reviews there are 2,000 observations and five distinct values. The empirical CDF is not a fine staircase with `n` small steps but a blocky function with at most five large jumps, each jump the proportion of reviews at that rating. ## What actually goes wrong **The statistic becomes coarse.** With five possible values there are only four interior points where the two empirical CDFs can be compared. `D` is the largest of four differences in cumulative proportion. It can take only a limited set of values, so the test's decision surface is chunky - small changes in the data either move `D` not at all or move it in a jump. **Distribution-freeness is lost.** The exact null distribution of `D` now depends on the shared rating proportions. Two samples where 80% of reviews are five stars behave differently under the null from two samples spread evenly across the five ratings. There is no single table indexed by `n` and `m` that covers both. **The classical p-value is generally conservative.** The continuous theory imagines the maximum being taken over an infinitely fine grid; with ties, the maximum is taken over only a few points, so the statistic tends to be smaller than the continuous null distribution expects. Comparing it to continuous critical values gives p-values that are too large and a true rejection rate below the nominal level. The failure mode is missed differences, not false alarms - which is the dangerous direction, because a non-significant result is read as reassurance. ## Fix one: permutation Stop asking a table what the null distribution is and build it from the data: 1. Pool all observations from both samples into one collection. 2. Randomly reassign group labels, keeping the two group sizes fixed. 3. Recompute `D` for that shuffled split. 4. Repeat several thousand times to collect a null distribution of `D`. 5. The p-value is the fraction of shuffled statistics greater than or equal to the observed `D`. This is valid precisely because the tie pattern is carried along by the resampling - the shuffled datasets contain exactly the same multiset of ratings, so whatever coarseness ties introduce is present in the null distribution too. It assumes only that under the null the group label is exchangeable with respect to the rating, which is exactly the hypothesis being tested. No continuity is needed anywhere. ## Fix two: treat the ratings as ordered cells With five values you already have five natural cells, so no binning decision is required. Lay out the two samples' counts per rating, compute what each cell's count would be if both samples shared one common rating distribution, and compare observed to expected with a chi-square-style statistic. This handles discreteness natively - discreteness is the model, not a violation of it. The same manoeuvre rescues a **continuous** sample that has been mangled into a few distinct values by rounding, truncation or a sentinel: group the range into cells and compare cell counts instead of running a test that assumes no repeats. The cost is that you discard information about where within a cell the observations sat, and you need enough observations in every cell for the comparison to behave. ## What not to do - **Do not jitter.** Adding tiny random noise to break ties technically restores continuity but injects an arbitrary choice - the noise scale - into the answer, and the result changes between runs. It converts a known problem into a hidden one. - **Do not report the classical p-value with a caveat and move on.** If it is conservative, a non-significant result is uninformative, and a caveat does not make it informative. - **Do not conclude 'no difference' from a non-significant classical result.** With heavy ties that is close to a non-statement. ## Reporting Whatever route you take, show the two rating distributions side by side as proportions per star. With five categories the full picture fits in one small table, and a reader can see immediately whether the difference is a shift toward five stars, a hollowing-out of the middle, or a growth in one-star reviews - which no single test statistic will tell them.

  • How do you get a trustworthy p-value for the maximum-gap statistic when the data is full of ties?
    Permute. Pool both samples, reshuffle the group labels many times keeping the group sizes fixed, recompute the statistic on each shuffle, and take the p-value as the fraction of shuffled values at least as large as the observed one. The tie pattern rides along in every shuffle, so no continuity assumption is needed.
  • When would you bin a continuous sample into cells and compare cell counts instead?
    When rounding, truncation or sentinel values have collapsed the measurements onto a few distinct points, or when the reference is only available as cell probabilities. Discreteness then becomes the model rather than a violation. The cost is losing within-cell detail and needing enough observations in each cell.
  • Is adding tiny random noise to break the ties an acceptable fix?
    No. It restores continuity only formally, and it makes the result depend on an arbitrary noise scale and on the random seed. Two analysts get different p-values from the same data. Permutation or an explicitly discrete comparison answers the question without inventing information.

saying these in an interview costs you the question

  • Assumes the classical p-value is still valid with heavy ties
  • Thinks ties make the test reject too often
  • Jitters the data with noise to break the ties
  • Reads a non-significant tied-data result as no difference
  • Believes the statistic stays distribution-free with repeats

context