skip to content

What null hypothesis does the overall F-test in a linear regression output test?

level: middleimportance: should knowfreq 58%

answer

  1. compares against predicting the mean
  2. a joint test, not one per coefficient
  3. all slopes zero simultaneously
  4. two degrees of freedom: k and n minus k minus 1
  5. with one predictor it is t squared

basics

~20 s

It tests whether every slope coefficient is zero at once, meaning the model does no better than an intercept-only model that predicts the outcome's mean. A small p-value says at least one predictor carries signal, without saying which.

solid answer

~50 s

The overall F-test pits the fitted model against the intercept-only model that predicts the outcome's mean for every row. The null is that all k slope coefficients are simultaneously zero; the alternative is that at least one is not. The statistic is `F = (R2/k) / ((1 - R2)/(n - k - 1))`, referred to an F distribution with `k` and `n - k - 1` degrees of freedom. Because it is a single joint test, a significant F alongside no individually significant t-statistic is a normal outcome when predictors are correlated: the block explains variation even though no single column can be credited with it. The reverse also happens - with many predictors, one t-statistic clearing 5 percent by chance need not move the joint test. With exactly one predictor, F equals that predictor's t-statistic squared.

go deeper

for a junior

Know that the F line at the foot of a regression summary is about the whole model at once, not about any single predictor's coefficient.

for a middle

State the null as all slopes zero simultaneously, name both degrees of freedom, and reproduce the F equals t squared identity in the single-predictor case.

for a senior

Show judgment about what significance buys at scale: on large samples the joint test is nearly always significant, so it stops being the question worth asking.

for a principal

Decide what evidence standard a model must meet before it informs a decision, and steer reviewers away from treating a significant F as the model's approval stamp.

## What is being compared The overall F reported at the foot of a regression summary is a comparison of two models on the same rows. The **restricted** model contains only an intercept: it predicts `y_bar` for every observation, and its residual sum of squares is the total sum of squares `TSS`. The **unrestricted** model is the one you fitted, with `k` slopes plus the intercept, and residual sum of squares `RSS`. The null hypothesis is the joint statement `H0: beta_1 = beta_2 = ... = beta_k = 0` with the intercept left free. The alternative is not that all of them are non-zero - it is that **at least one** is. ## The statistic `F = ((TSS - RSS)/k) / (RSS/(n - k - 1))` which, dividing top and bottom by TSS, can be written in terms of the fit statistic as `F = (R2/k) / ((1 - R2)/(n - k - 1))` The numerator is the variation the slopes explain, per slope. The denominator is the leftover variation, per residual degree of freedom. Under the null and the usual linear-model assumptions, this ratio follows an F distribution with `k` numerator and `n - k - 1` denominator degrees of freedom. Large values are evidence against the null; the p-value is the upper tail. ## Joint versus individual tests This is where most interview follow-ups live. Each t-statistic in the coefficient table tests one coefficient **given the others are in the model**. The F tests all of them at once. The two can disagree in both directions: - **Significant F, no significant t.** The typical cause is correlated predictors. Several columns carry overlapping information, so the model as a whole explains real variation, but the shared information inflates every individual standard error and no single coefficient can be separated from zero. The joint test does not suffer from this because it never has to attribute the explained variation to a particular column. - **A significant t, unremarkable F.** With many predictors and weak signal, one t-statistic crossing the 5 percent line is unsurprising by chance alone, and one small contribution among many need not move the joint statistic. With a single predictor there is no room for disagreement: `F = t^2` exactly, and the two p-values coincide. ## What significance here does and does not buy The alternative hypothesis is extremely weak - the model beats the mean by some non-zero amount. The denominator degrees of freedom grow with the sample, so on a large dataset even a tiny R-squared clears the threshold comfortably. With `n = 50000` and `k = 5`, an R-squared of 0.002 already gives a large F. So on modern data volumes a significant overall F is close to guaranteed and carries almost no information; the interesting questions are about effect sizes and specification, not about whether the model beats a flat line. Significance also says nothing about whether the model is **correctly specified**. A curved relationship fitted with a straight line, a model missing its main confounder, and a model contaminated by a predictor derived from the outcome will all produce a hugely significant F. The test is a floor to clear, not a certificate. ## Assumptions worth naming The exact F distribution follows from the usual linear-model conditions: correctly specified mean, independent observations, constant residual variance, and approximately normal errors (or a large enough sample for the approximation to hold). Heavy dependence between observations, such as repeated measurements on the same unit, inflates the statistic because the effective sample size is smaller than `n`. ## How to answer crisply Name the null as all slopes simultaneously zero, name the comparison model as intercept-only, name both degrees of freedom, and finish with the joint-versus-individual point. Adding the `F = t^2` identity for the single-predictor case reliably signals that you understand the two tests are the same machinery at different scopes.

  • The overall F-test is significant but no individual t-statistic is. What is going on?
    Usually correlated predictors. The columns jointly explain variation, but the information they share inflates each coefficient's standard error, so no single one is distinguishable from zero. The joint test is untroubled by this because it never has to allocate the explained variation to a specific column.
  • Why is the overall F-test almost always significant on a very large dataset?
    The denominator degrees of freedom grow with the sample, so even a trivially small explained share produces a large statistic. At n = 50,000 an R-squared of 0.002 is comfortably significant. Significance answers whether the model beats a flat line at all, not whether the improvement is worth acting on.
  • Does a significant overall F-test mean the model is well specified?
    No. It rules out only the intercept-only model. A curved relationship fitted with a straight line, an omitted confounder, or a predictor derived from the outcome can all produce an enormous F. Specification is judged by residual structure and by what the predictors actually are, not by this p-value.

saying these in an interview costs you the question

  • Says the overall F checks whether one specific predictor matters
  • Reads a significant F as proof the model is correctly specified
  • Assumes a significant F guarantees at least one significant t
  • Confuses the F-test's significance with practical usefulness
  • Cannot name the two degrees of freedom

context