skip to content

How does a chi-square goodness-of-fit test decide whether 600 die rolls came from a fair die?

level: juniorimportance: must knowfreq 68%

answer

  1. counts, never percentages
  2. observed against a hypothesised distribution
  3. square the gap, divide by expected
  4. categories minus one, minus estimated parameters
  5. big statistic means poor fit

basics

~20 s

A chi-square goodness-of-fit test compares observed category counts with the counts a hypothesised distribution predicts. A fair die over 600 rolls predicts 100 per face; the statistic sums (observed minus expected) squared, divided by expected, across the six faces.

solid answer

~50 s

You state the null hypothesis as a specific distribution over categories: each face of a fair die has probability 1/6, so 600 rolls give an expected count of 100 per face. With observed counts 92, 108, 95, 117, 89, 99 the statistic is `sum((O - E)^2 / E)` = (64 + 64 + 25 + 289 + 121 + 1)/100 = 5.64. Degrees of freedom are the number of categories minus one, minus any parameter you estimated from the same data — here 6 - 1 = 5. Compared against the chi-square distribution on 5 df, the 5% critical value is about 11.07, so 5.64 does not reject: these counts are consistent with a fair die. Two rules matter: feed the test raw counts, never percentages, and remember a non-rejection is not proof of fairness.

go deeper

for a junior

Be ready to say what goes in the formula: observed counts, expected counts from the hypothesised probabilities, and categories minus one for degrees of freedom. Practise doing the arithmetic out loud on a small table.

for a middle

Explain why the division is by the expected count and why one degree of freedom disappears because the counts must total N. Expect to be asked what changes when a parameter is estimated from the same data.

for a senior

Show judgment about when the chi-square approximation is safe, what a non-rejection does and does not license, and how you would localise a rejection to specific categories without over-reading residuals.

for a principal

Own the framing question: whether a categorical fit test is the right instrument at all, what deviation would actually matter to the business, and how you avoid a huge sample turning trivial miscalibration into a significant result.

## What the test is for A **goodness-of-fit** test asks whether counts falling into a fixed set of categories are consistent with one specific probability distribution over those categories. The data are frequencies — how many times each outcome occurred — not measurements or averages. The die is the canonical example: six categories (faces 1 to 6), a null hypothesis that each face has probability 1/6, and 600 tosses to judge it on. ## The pieces **Observed counts (O)** are what you actually saw. For our 600 rolls: 92, 108, 95, 117, 89, 99. They sum to 600 by construction. **Expected counts (E)** are what the null hypothesis predicts *on average*: E_i = N * p_i, where p_i is the null probability of category i. With p_i = 1/6 and N = 600, every E_i = 100. Expected counts need not be whole numbers — 100.5 is a perfectly valid expected count. **The statistic** is chi-square = sum over categories of (O_i - E_i)^2 / E_i Each term measures how far one category strayed from its prediction, squared so that shortfalls and surpluses both count, then divided by the expected count. That division is the part candidates most often cannot explain: count data are noisier where counts are larger (variance grows roughly with the mean), so a gap of 17 is startling when you expected 20 and unremarkable when you expected 5,000. Dividing by E puts every category on a comparable noise scale. For the die: deviations are -8, +8, -5, +17, -11, -1; squares are 64, 64, 25, 289, 121, 1; the sum is 564; dividing by 100 gives **chi-square = 5.64**. ## Degrees of freedom The reference distribution is the chi-square family, indexed by degrees of freedom (df). For goodness of fit, df = (number of categories) - 1 - (number of parameters estimated from the same data) The -1 is because the counts must sum to N: once five face counts are known, the sixth is determined, so only five of the six deviations are free to vary. Nothing was estimated here — 1/6 came from the hypothesis, not the data — so df = 6 - 1 - 0 = **5**. If instead you had fitted, say, a Poisson rate to the data and then tested how well that Poisson fits, you would subtract one more df for the estimated rate. Forgetting this makes the test too conservative or too liberal depending on the case, and it is a favourite interview probe. ## Reading the result Large values of the statistic mean poor fit; the test is inherently **one-tailed in the statistic** even though it detects deviations in any direction, because squaring has already thrown away the signs. On 5 df, the upper 5% point of the chi-square distribution is about 11.07. Our 5.64 sits well below it, so at the 5% level we fail to reject the hypothesis of a fair die. "Fail to reject" is not "the die is fair". With 600 rolls the test simply lacks the power to detect a small bias; a die loaded to give face 4 a probability of 0.18 instead of 0.167 would often slip through. If the question is whether the bias is small enough to ignore, report the observed proportions and their uncertainty rather than leaning on a non-significant p-value. ## Conditions and common mistakes The chi-square reference distribution is an **approximation** that improves with sample size; it is trustworthy when the expected counts are not tiny (a widely used rule of thumb wants all expected counts at least 5). Note the rule is about *expected* counts, not observed ones — a category with zero observations is fine as long as its expectation is healthy. Other recurring errors: - **Feeding proportions instead of counts.** The statistic scales with N; hand it 0.153 instead of 92 and the answer is arithmetic noise. - **Assuming the rolls are independent.** Every category count must come from independent trials, and each observation must land in exactly one category. - **Declaring which face is guilty after a rejection.** The overall test says only that the fit is poor. To localise the problem, inspect standardised residuals `(O - E) / sqrt(E)` per category and treat values beyond roughly ±2 as the cells driving the result — while remembering you are now looking at six comparisons at once. The same machinery works for any fully specified categorical distribution, not just the uniform one — the only thing that changes is how you compute each E_i.

  • Why does each squared deviation get divided by the expected count rather than by N?
    Because the noise in a cell scales with that cell's expectation. Count data behave roughly Poisson-like, so the standard deviation of a cell is about sqrt(E). Dividing the squared deviation by E turns each term into a squared standardised deviation, which is exactly what makes the sum follow a chi-square distribution.
  • How do the degrees of freedom change if you estimate a parameter from the same data before testing the fit?
    You lose one degree of freedom per estimated parameter: df = k - 1 - m for k categories and m parameters fitted from the data. Estimating from the data pulls the expected counts toward the observations, shrinking the statistic, and the df reduction compensates for that. Ignoring it makes the test anti-conservative.
  • The test rejects. How do you find which category is responsible?
    Compute standardised residuals per category, `(O - E) / sqrt(E)`, which are roughly standard normal under the null. Cells beyond about ±2 are the ones driving the rejection. Treat this as exploratory: you are now scanning several cells, so individual residuals are not clean per-cell significance tests.

It is a budget variance report: for each line you compare actual against forecast, and you judge the miss relative to how big that line was supposed to be, not in absolute pounds.

saying these in an interview costs you the question

  • Runs the test on percentages instead of raw counts
  • Says degrees of freedom equal the number of categories
  • Treats a non-significant result as proof the die is fair
  • Applies the test to continuous measurements with no categories
  • Divides squared deviations by the observed rather than expected count

context