What does Wilks' theorem say about the likelihood-ratio statistic for two nested models?
answer
- compare two fits, one a restriction of the other
- difference of maximised log-likelihoods
- there is a factor of two
- degrees of freedom count parameters removed
- interior restrictions only
basics
~20 sTwice the gap in maximised log-likelihoods between a full model and a nested restricted one converges, when the restriction is true, to a chi-square distribution whose degrees of freedom equal the number of parameters the restriction removes.
solid answer
~50 sFor nested models, define `D = 2 * (logL_full - logL_reduced)` using each model's maximised log-likelihood. Wilks' theorem says that if the reduced model is true and the usual regularity conditions hold, `D` converges in distribution to `chi-square` with degrees of freedom equal to the number of free parameters the reduction eliminates. So a model with one extra rate parameter is compared against `chi-square` with 1 degree of freedom, and a restriction of three coefficients against `chi-square` with 3. The practical use is deciding whether extra parameters earn their keep: fit both models on the same data, take the doubled difference, and compare it to the reference distribution. Two conditions bind. The models must be genuinely nested and fitted to identical data. And the restricted parameters must lie in the interior of the parameter space — restrictions that sit on a boundary, such as a variance fixed at zero, break the stated null distribution.
go deeper
Recall that a bigger model always fits at least as well, so improvements have to be judged against what chance alone would give. Know the statistic compares two maximised log-likelihoods.
Be ready to write the statistic with its factor of two and derive the degrees of freedom by counting free parameters in each model and subtracting.
Show you check the preconditions before quoting a result: genuine nesting, identical data in both fits, an interior restriction, and a sample large enough for the asymptotic reference to hold.
Own the model-selection policy: when the team is entitled to lean on an asymptotic reference distribution, when resampling should calibrate it instead, and how to stop nested-test results being used as a licence for repeated searching.
## Nested models Two models are **nested** when the smaller is a special case of the larger, obtained by fixing or constraining some of the larger model's parameters. Fitting a single Poisson rate to a whole dataset is nested inside fitting two separate rates to two segments: the one-rate model is the two-rate model with the constraint `lambda_1 = lambda_2`. Nesting is what makes the comparison meaningful; two models that merely describe the same data are not comparable this way. ## The statistic Maximise the log-likelihood separately under each model, giving `logL_full` and `logL_reduced`. Because the reduced model's parameter space is a subset of the full one, `logL_full >= logL_reduced` always — the extra freedom can never fit worse. The question is whether the improvement is more than the noise you would expect from that extra freedom alone. The **likelihood-ratio statistic** is `D = 2 * (logL_full - logL_reduced) = -2 * log(L_reduced / L_full)` The doubling is not cosmetic: it is exactly what makes the limiting distribution a chi-square rather than a scaled one. ## Wilks' theorem Under the null hypothesis that the restriction holds, and under regularity conditions, `D` converges in distribution as the sample grows to `chi-square` with `k` degrees of freedom, where `k = (number of free parameters in the full model) - (number of free parameters in the reduced model)` That is, `k` is the number of restrictions imposed. Getting `k` right is the part candidates most often fumble: it counts *parameters removed*, not observations, not groups, not categories. ## Why a chi-square appears Expand the log-likelihood quadratically around the full model's maximum. To second order the surface is a paraboloid whose curvature is the observed information. Restricting `k` parameters projects the estimate onto a lower-dimensional subspace, and the drop in the quadratic form is a sum of `k` squared, asymptotically independent standard Normal quantities — the definition of a chi-square with `k` degrees of freedom. The same quadratic expansion is the origin of the asymptotic Normality of the estimate itself, which is why the two results are two faces of one theory. ## A worked comparison Suppose events arrive at a constant rate and you wonder whether the rate differs before and after a change. The reduced model has one rate parameter; the full model has two. Fit both by maximum likelihood on the same events, compute `D = 2*(logL_two_rate - logL_one_rate)`, and compare to `chi-square` with `2 - 1 = 1` degree of freedom. A value near zero says the second rate parameter bought nothing; a large value says the single-rate description is implausibly poor. Note that the comparison is only legitimate if both fits used exactly the same observations — dropping rows in one fit and not the other silently invalidates the statistic. ## The three classical tests Wilks' likelihood-ratio test is one of a trio that are asymptotically equivalent under the null: - **Likelihood-ratio**: fits both models, uses the gap in log-likelihoods. - **Wald**: fits only the full model, asks how far the estimate is from the restricted value relative to its standard error. - **Score (Lagrange multiplier)**: fits only the reduced model, asks how steep the log-likelihood still is in the restricted directions. They agree in the limit but can disagree noticeably in finite samples. The likelihood-ratio version is generally the best behaved of the three and, unlike the Wald statistic, is invariant to how the parameter is parameterised — a Wald test on a coefficient and on its exponential are not the same test, while the likelihood-ratio statistic is unchanged. ## Where it fails - **Boundary restrictions.** If the restricted value sits on the edge of the parameter space — a variance component fixed at zero, a mixing weight at zero — the quadratic expansion is one-sided and the null distribution is not `chi-square_k`. In the single-variance-component case it is a 50:50 mixture of a point mass at zero and `chi-square` with 1 degree of freedom, so using the plain reference distribution is conservative there. - **Non-nested models.** The statistic has no chi-square justification at all; comparison needs a different tool. - **Different data.** Any difference in the rows entering the two fits voids the comparison. - **Small samples.** It is an asymptotic result; with few observations the reference distribution can be a poor approximation, and resampling-based calibration is safer. - **Non-identified extra parameters.** If a parameter present only in the full model is unidentified under the null, the regularity conditions fail and the standard degrees-of-freedom count does not apply. ## What a strong answer contains The formula with the factor of two, the correct degrees-of-freedom count, the requirement that models be nested and fitted to the same data, and at least one named failure mode. Reciting the theorem without the boundary caveat is a middle-tier answer; naming when you would not trust it is the senior one.
- How do you count degrees of freedom for the likelihood-ratio statistic?Count free parameters in each model and subtract: the degrees of freedom equal the number of restrictions the reduced model imposes. Constraining three coefficients to zero gives 3; forcing two group rates to be equal gives 1, because two free parameters collapse to one. It is a count of parameters removed, never of observations or groups.
- Why can the likelihood-ratio and Wald tests disagree in a finite sample?They approximate the same limiting behaviour from different directions. The Wald test relies on a quadratic approximation around the estimate and is sensitive to how the parameter is parameterised; the likelihood-ratio test uses the actual log-likelihood at both fits and is parameterisation-invariant. When the log-likelihood is skewed, the Wald version can be badly calibrated while the likelihood-ratio version stays reasonable.
- The extra parameter you are testing is a variance constrained to be non-negative. What breaks?The null puts the parameter on the boundary of its space, so the two-sided quadratic expansion behind Wilks' theorem does not hold. The null distribution is not the plain chi-square with one degree of freedom; for a single variance component it is a 50:50 mixture of a point mass at zero and chi-square with 1 degree of freedom. Using the plain reference is conservative.
- Can you use this statistic to compare two models that are not nested?No. The chi-square limit depends on one parameter space being a restriction of the other, so with non-nested models the statistic has no known reference distribution and the comparison is meaningless. Model comparison there needs a different framework, and even then the two fits must at minimum use the same observations.
Adding parameters is like adding adjustable screws to a bracket: the fit can only get tighter. The theorem tells you how much tightening pure chance buys you per screw, so you can see whether a screw actually did work.
saying these in an interview costs you the question
- Omits the factor of two in the statistic
- Counts degrees of freedom as observations rather than restricted parameters
- Applies the test to models that are not nested
- Compares fits made on different subsets of the data
- Uses the plain chi-square when the restriction sits on a parameter boundary
- Claims a larger full-model log-likelihood by itself proves the extra parameters help