Why can Shapiro-Wilk reject normality at n = 5,000 yet pass a clearly skewed n = 8 sample?
answer
- power grows with sample size
- the null is never accepted
- no real data are exactly normal
- loud when it does not matter
- silent when it does
basics
~20 sA normality test measures evidence against normality, not the size of the departure. At n = 5,000 it detects deviations far too small to matter; at n = 8 it has almost no power, so passing proves nothing about the population.
solid answer
~50 sShapiro-Wilk tests the null hypothesis that the sample came from a normal distribution, and like any hypothesis test its power rises with sample size. Real data are never exactly normal, so at n = 5,000 a departure of no practical consequence is easily detected and the p-value collapses. At n = 8 the test can barely distinguish anything, so failing to reject is absence of evidence, not evidence of normality - you never accept a null. The perverse consequence is that the test shouts precisely when the sample is large enough that the central limit theorem has already made normality irrelevant, and stays silent when the sample is small and the assumption matters most. I therefore treat it as a screen at most: I look at a Q-Q plot to see the direction and magnitude of the departure, judge whether that magnitude threatens the procedure I am running, and decide from there.
go deeper
Remember the direction of the logic: a small p-value is evidence against normality, and a large one is not evidence for it. Never report that a test proved the data are normal.
Explain power. Say why a test of a null that is never exactly true rejects eventually as n grows, and why a tiny sample cannot convict even a badly skewed population.
Show the judgment: describe how you replace the yes/no verdict with a magnitude question, tie the departure to what your specific procedure is sensitive to, and resample when it is genuinely borderline.
Own the policy. Decide whether automated pipelines should branch on normality tests at all, what evidence is recorded instead, and how to keep teams from either rubber-stamping or over-reacting to a p-value driven by sample size.
## What the test is asking Shapiro-Wilk takes a sample and tests the null hypothesis that it was drawn from a normal distribution. The statistic is built from a comparison of the sorted sample against the values expected from a normal distribution - essentially a numerical summary of how well the points would line up on a normal Q-Q plot. A small p-value is evidence against normality; a large one is not evidence for it. ## Why sample size dominates the verdict Every hypothesis test has power: the probability of rejecting the null when it is false. Power increases with sample size and with the size of the true departure. Two facts then collide. First, exact normality never holds for real measurements. Data are bounded below by zero, rounded to two decimals, censored at a cap, or generated by a mixture of processes. There is always some departure, so the null is always false in the strict sense. With enough data, any test of a null that is false by any margin will eventually reject. At n = 5,000 the amount of departure needed to produce p < 0.001 is far smaller than the amount that would change any conclusion you would draw. Second, at n = 8 the test has so little power that even a strongly skewed population usually fails to produce a significant result. Non-significance means the data could not convict, not that the population is innocent. This is the standard asymmetry of null hypothesis testing: rejection is informative, non-rejection is not. Put together, the test's sensitivity moves in exactly the opposite direction from the assumption's importance. For a comparison of means, non-normality of the raw data matters most at small n, where the central limit theorem has not had room to work - and that is where the test is blind. At large n the central limit theorem has largely absorbed the issue - and that is where the test screams. ## The right question: how big, in what direction, and does it matter? Swap the yes/no question for a magnitude question. Look at a normal Q-Q plot to see whether the departure is skew or tail weight, whether it is driven by the bulk or by a few points, and how far the ends stray. Quantify with descriptive measures such as skewness and kurtosis if you want numbers. Then ask what the procedure is sensitive to: - Comparing means at a healthy sample size: moderate skew is absorbed; what remains dangerous is a handful of extreme values, which inflate the sample standard deviation and drain power. - Small samples: the raw shape matters directly, and this is where you should be most cautious - though it is also where you have the least evidence about the shape. - Anything about the tails - a high percentile, a prediction interval, a probability of exceeding a threshold: the tail behaviour is the estimand, so heavy tails are not a technicality but the whole answer. When you genuinely cannot tell whether a departure matters, resample: draw repeatedly from your own data and see whether the interval or error rate you plan to report actually holds. That answers the operational question directly, rather than the philosophical one about whether the population is exactly normal. ## Legitimate uses of the test None of this makes the test useless. It is a cheap screen on a moderate sample. It is defensible as a documented, pre-declared step in a process that requires one. It works as triage across many series at once, where nobody can eyeball every plot. And significance on a moderate sample size, combined with a plot showing a large bend, is a real signal. What it must not be is the automatic gate that decides your analysis, because its verdict tracks n more strongly than it tracks anything you care about. ## Responding when it rejects If a large sample fails a normality test and you wanted to compare means, the usual answer is to proceed, after checking the plot for outliers or a second mode - those distort the mean itself and are a substantive finding, not a technicality. If the sample is small and the plot shows a serious bend, the honest options are to transform the outcome, to use a procedure that does not rely on the shape, or to accept wider uncertainty. And if the outcome is strongly skewed, the deeper question is whether the mean is the right summary at all, which no normality test will ever tell you.
- If a normality test misleads in both directions, when is it worth running at all?As a cheap screen on a moderate sample, as a pre-declared step when a process demands documentation, or as triage across many series where nobody can inspect every plot. Significance on a mid-sized sample, paired with a plot showing a large bend, is genuine signal. What it must not be is the automatic gate that selects your analysis, since its verdict tracks sample size more than any departure you care about.
- A sample of 5,000 fails a normality test but you want to compare group means. What now?Proceed. At that size the sampling distribution of the mean is close to normal for anything short of extreme tails, so the failed test is not a reason to abandon the comparison. I would still plot the data to check for outliers or a second mode, because those distort the mean itself and are a substantive finding. I would also ask whether the mean is the right summary of a strongly skewed outcome.
- How do you decide whether a departure from normality is big enough to matter?By tying it to the procedure and the sample size instead of to a p-value. I read the Q-Q plot for the shape and extent of the bend, quantify skew and tail weight, and when it is borderline I resample from my own data to see whether the interval or error rate I intend to report actually holds up. The question is whether my conclusion changes, not whether the distribution is textbook.
saying these in an interview costs you the question
- Reads a large p-value as proof that the data are normal
- Abandons a mean comparison because a test rejected at n = 5,000
- Believes real-world measurements can be exactly normal
- Applies the same significance threshold regardless of sample size
- Runs a normality test on eight observations and trusts the verdict