When are expected cell counts too small for a chi-square test, and what do you use instead?
answer
- the rule is about expected, not observed
- at least five in every cell
- Cochran allows a fifth below five
- condition on both margins
- hypergeometric enumeration, no approximation
basics
~20 sThe common rule wants every expected count at least 5, or more leniently no expected count below 1 with at most a fifth below 5. Below that, collapse categories, gather more data, or run Fisher's exact test.
solid answer
~50 sThe chi-square reference distribution is an asymptotic approximation to the true discrete sampling distribution of the statistic, and it degrades when cells are sparse. The rule of thumb is that all **expected** counts should be at least 5; a more permissive version allows up to 20% of cells between 1 and 5 provided none falls below 1. The key trap is that the rule is about expected counts, not observed ones — an observed zero is fine if the expectation is healthy. When it fails you have three options: merge substantively similar categories so cells fill up, collect more data, or abandon the approximation and compute an exact p-value. For a 2x2 table that exact route is **Fisher's exact test**, which conditions on both margins and enumerates every table with those margins using the hypergeometric distribution. It is exact rather than approximate, and it stays valid at any sample size.
go deeper
Know the headline rule: expected counts of at least 5 in every cell, and Fisher's exact test as the standard fallback for a small 2x2 table. Say expected, not observed.
Explain why the rule exists — the chi-square curve is a large-sample approximation to a discrete statistic — and describe how conditioning on both margins makes Fisher's test exact through the hypergeometric distribution.
Show the operational judgment: which repair fits which situation, why post-hoc category merging invalidates the p-value, and when the honest answer is that the data are too sparse to support any test.
Own the upstream call. Sparse cells usually mean the category scheme or the sample size was wrong for the question, so the lead's job is to fix the design rather than to shop for a method that returns a number.
## Why there is a rule at all The chi-square statistic is computed from counts, which are discrete, but it is compared against a continuous chi-square distribution. That comparison is justified by a limiting argument: as the expected count in each cell grows, the standardised cell deviations `(O - E)/sqrt(E)` behave increasingly like normal variables, and the sum of their squares behaves increasingly like a chi-square variable. When expectations are small the approximation is poor, and typically **anti-conservative** — the reported p-value is smaller than the truth, so you over-reject. ## The rules of thumb Two versions circulate and both are worth being able to state: - **The strict version:** every expected count is at least 5. - **Cochran's more permissive version:** no expected count is below 1, and at most about 20% of cells have expected counts below 5. Neither is a theorem. They are practical guardrails, and interviewers mainly want to see that you know the rule constrains **expected** counts. This is the single most common error: a candidate points at an observed zero and declares the test invalid. An observed zero in a cell that expected 40 is a striking finding, not a violation. Conversely a cell that observed 3 but expected 0.4 is exactly the problem case, whatever the observed number happens to be. ## What to do when it fails **Collapse categories.** If a survey has "fully on-site" answered by four people, folding it into an adjacent category raises the expectations. This must be defensible on substance, decided before you see the p-value; merging categories to manufacture significance is a form of data dredging. **Get more data.** Expected counts scale with N, so doubling the sample doubles every expectation. Often unavailable, but it is the honest first answer. **Compute an exact p-value.** Instead of leaning on the limiting distribution, enumerate the possible tables and add up the probabilities of those at least as extreme as the one observed. For a 2x2 table this is **Fisher's exact test**; for larger tables the same idea generalises through exact or permutation-based enumeration, at greater computational cost. ## Fisher's exact test Fisher's test conditions on **both** sets of margins being fixed at their observed values. Under that conditioning and the null hypothesis of no association, the count in any one cell follows a **hypergeometric distribution** — the same distribution as drawing without replacement from an urn. Because the margins pin down the whole table once one cell is known, you can enumerate every possible table, compute each one's probability exactly, and sum the probabilities of the tables at least as extreme as yours. No large-sample approximation is involved, which is why it remains valid however small the counts are. The design that motivated it is the **lady tasting tea**. A taster claims she can tell whether milk or tea was poured first. Eight cups are prepared, four each way, and she is told there are four of each and asked to identify which four had milk first. Under the null hypothesis that she is guessing, the number of correctly identified cups follows a hypergeometric distribution over the C(8,4) = 70 equally likely selections. Getting all four right happens with probability 1/70, about 0.014 — small enough to reject guessing at the 5% level, and only just. Getting three of four right is far more likely under guessing and would not be persuasive. This is also a neat illustration of the conditioning: the taster knows the margins (four and four), which is precisely what Fisher's test assumes. A nuance worth knowing: because it conditions on margins that were not truly fixed by the design in most observational tables, Fisher's test is often **conservative** — its actual rejection rate can sit below the nominal level. So "exact" means the p-value is computed exactly under its own conditional null, not that it is uniformly the most powerful choice. ## Related repairs to know about For 2x2 tables, **Yates' continuity correction** subtracts 0.5 from each absolute deviation before squaring, nudging the discrete statistic toward the continuous reference. It shrinks the statistic and enlarges the p-value, and it is widely considered over-conservative; most practitioners now prefer either the uncorrected statistic when expectations are adequate or an exact test when they are not. ## The judgment call Sparse tables are usually a signal about the study, not just the arithmetic. A cell expecting 0.3 means you have almost no information about that combination, and no test will conjure it. The right response is often to say so, present the raw counts, and either widen the categories or plan a larger collection — rather than to hunt for a method that returns a number.
- Why does the chi-square approximation tend to over-reject rather than under-reject when cells are sparse?The statistic is discrete but is read against a continuous curve. With small expectations its true distribution is lumpy and sits to the right of the smooth chi-square in the upper tail, so the tail area you look up understates the true probability of a value that extreme. The p-value comes out too small.
- Which distribution does Fisher's exact test use, and why is that the right one?The hypergeometric. Conditioning on both margins is equivalent to drawing a fixed number of items without replacement from a finite pool split into two classes, which is exactly what the hypergeometric describes. Because it is enumerated rather than approximated, the p-value is valid at any sample size.
- A colleague merges two sparse categories after seeing that the test just missed significance. What is your objection?The merge is now a data-dependent choice made in the p-value's favour, so the reported significance is no longer calibrated. Category structure should be fixed on substantive grounds before analysis. If a merge is genuinely defensible, say it was made post hoc and treat the result as exploratory.
saying these in an interview costs you the question
- Applies the minimum-count rule to observed instead of expected counts
- Declares any table with an observed zero invalid
- Thinks Fisher's exact test needs large samples
- Merges categories after seeing the p-value
- Believes a continuity correction makes the test exact