skip to content

Why are standard Kolmogorov-Smirnov critical values wrong when the reference normal's mean and SD come from the same sample?

level: middleimportance: should knowfreq 34%

answer

  1. the reference was fitted to the same data
  2. the gap gets pulled smaller
  3. table assumes a pre-specified curve
  4. p-values too large, test under-rejects
  5. simulate the estimate-then-test procedure

basics

~20 s

Fitting the reference curve to the same data pulls it toward the sample, shrinking the maximum gap. Standard tables assume a reference fixed in advance, so p-values come out too large and the test under-rejects.

solid answer

~50 s

The one-sample statistic is `D = max_t |F_n(t) - F0(t)|`, the largest vertical gap between the empirical CDF and a reference CDF `F0`. The published critical values assume `F0` was fixed in advance. If you instead estimate the mean and SD from the very sample you are testing, the fitted curve is drawn toward the data, so `D` is systematically smaller than the null table expects. Using the standard table then gives p-values that are too large - the test is conservative and misses real departures. The fix is critical values simulated under the same estimate-then-test procedure: for a normal with both parameters estimated, those are the Lilliefors values, roughly two-thirds of the standard ones at the 5% level. They are family-specific, so a different fitted family needs its own simulation, which you can do yourself by parametric bootstrap.

go deeper

for a junior

Remember the trap itself: comparing a sample to a curve whose parameters came from that same sample is not the situation the published critical values were built for.

for a middle

Explain the mechanism and the direction. Fitting pulls the reference toward the data, the maximum gap shrinks, so p-values come out too large and the test under-rejects rather than over-rejects.

for a senior

Demonstrate the repair: simulate the whole estimate-then-test loop to build the correct null, or hold out a disjoint subset so the reference is fixed, and state clearly which route the reported p-value used.

for a principal

Generalise it into a team rule - the null distribution must match the procedure actually run, estimation included - and decide when a simulated null is worth its compute versus choosing a decision rule that needs no p-value at all.

## The setup A one-sample Kolmogorov-Smirnov test compares one sample against a **reference CDF** `F0`. The statistic is ``` D = max over t of | F_n(t) - F0(t) | ``` where `F_n` is the empirical CDF - the step function that jumps by `1/n` at each observation. Large-sample critical values come from the Kolmogorov distribution: reject at 5% when `sqrt(n) * D` exceeds about 1.36. ## The hidden assumption in those critical values That null distribution is derived under a specific story: `F0` is **completely specified before you look at the data**. A test of 'is this sample drawn from a standard normal with mean 0 and SD 1?' fits the story. A test of 'is this sample drawn from *some* normal?' does not, because you do not know which normal - so the natural move is to plug in the sample mean and sample SD. The moment you do that, the reference stops being a fixed curve and becomes a curve **fitted to the data being tested**. Estimation aims the reference at the sample: the fitted curve passes through roughly the sample's own centre and matches its own spread. The maximum gap between a sample and a curve fitted to it is systematically smaller than the gap between that sample and a curve chosen without seeing it. ## Which way the error goes Because `D` is shrunk by the fitting, comparing it to a table built for un-fitted references means you are comparing a deflated statistic to an inflated threshold. Consequences: - The p-value is **too large**. - The true type I error rate is **below** the nominal level - the test is **conservative**. - Power collapses: real departures from normality go undetected, and the analyst walks away with false reassurance. Getting the direction right is what interviewers are checking. Many candidates guess that estimating parameters makes a test too eager to reject; here it makes it too reluctant. ## The correction The repair is not to change the statistic but to change the reference distribution it is judged against. Simulate the *entire procedure* under the null: draw many samples of size `n` from the fitted family, and for each one re-estimate the parameters from that simulated sample, compute `D` against its own fitted curve, and collect the results. The resulting distribution of `D` is the correct null. Its upper percentiles are noticeably smaller than the classical ones - for a normal with both mean and SD estimated, the 5% critical value is roughly two-thirds of the classical `1.36 / sqrt(n)`. Those tabulated values for the estimated-normal case are the **Lilliefors** critical values. The same idea, done by simulation on the fly, is a **parametric bootstrap**: fit, simulate from the fit, re-fit each simulated sample, build the null distribution of `D` empirically, and read the p-value off as the fraction of simulated statistics at least as large as yours. ## Why it is family-specific Once parameters are estimated, `D` is no longer distribution-free. How much the fitting shrinks `D` depends on which family was fitted and how many parameters were estimated. Values derived for an estimated normal do **not** transfer to, say, an exponential with an estimated rate - that case has its own separately derived table. Any candidate who says 'use the Lilliefors correction' for an arbitrary fitted family has missed this. There is one saving grace: for location-scale families such as the normal, the corrected null distribution does not depend on the *values* of the estimated mean and SD, only on the fact that they were estimated - which is why a single table indexed by `n` can exist at all. ## Practical guidance - If the reference is genuinely known in advance - a specification, a theoretical model with no free parameters, a distribution from an independent historical period - the classical values are correct, and this whole problem disappears. - If you must estimate, either use values derived for that exact estimate-then-test procedure or simulate them yourself. - Split-sample is a clean alternative when data is plentiful: estimate the parameters on one half, test on the other. The reference is then fixed relative to the tested half, and the classical critical values apply again. - Report which route you took. 'KS against a fitted normal, classical p-value' is a result an informed reader will discount. ## The general lesson This is one instance of a broad rule: **the reference distribution must be the distribution of your statistic under the procedure you actually ran**, estimation steps included. Reusing a table derived for a simpler procedure silently changes the error rate, and it does not always change it in the safe direction.

  • If someone uses the classical table anyway, in which direction is their error?
    Conservative. The fitted reference shrinks the maximum gap, so the statistic is compared against a threshold that is too high, the p-value is too large, and the true rejection rate sits below the nominal level. The practical harm is missed departures and unearned confidence, not false alarms.
  • Does a correction derived for an estimated normal transfer to other fitted families?
    No. Once parameters are estimated the statistic is no longer distribution-free, and the amount of shrinkage depends on which family was fitted and how many parameters. An exponential with an estimated rate has its own separately derived values. For anything else, simulate the null yourself by parametric bootstrap.
  • How would you simulate the correct null distribution from scratch?
    Fit the family to your sample. Repeatedly draw a fresh sample of the same size from the fitted distribution, re-estimate the parameters from that simulated sample, and compute the statistic against its own re-fitted curve. The p-value is the fraction of simulated statistics at least as large as the observed one.
  • Is there a way to keep the classical critical values when parameters are unknown?
    Yes, if data is plentiful: estimate the parameters on one subset and run the test on a disjoint subset. The reference is then fixed with respect to the tested data, so the classical null distribution holds. The cost is power, since each half is smaller than the whole.

It is like grading an archer against a target painted around wherever the arrows landed. The scores look excellent, but the standard scoring table was written for targets nailed up before the shot.

saying these in an interview costs you the question

  • Thinks estimating parameters makes the test reject too often
  • Applies the classical table to a data-fitted reference curve
  • Assumes the correction works for any fitted family
  • Believes the statistic stays distribution-free after estimation
  • Reads a non-significant result as proof the fitted family is right

context