How do you test whether a block of three interaction terms improves a regression fit?
answer
- one joint test, not three t-tests
- two nested fits on the same rows
- difference in residual sums of squares
- numerator df equals the number of restrictions
- watch for rows dropped by missing values
basics
~20 sFit the model with and without the three terms on identical rows and run a nested F-test: the drop in residual sum of squares per added parameter, divided by the full model's residual variance, on 3 and n-k-1 degrees of freedom.
solid answer
~50 sTest the block jointly rather than reading three separate t-statistics. Fit the reduced model without the interactions and the full model with them, on exactly the same rows, then compute `F = ((RSS_r - RSS_f)/q) / (RSS_f/(n - k - 1))`, where `q = 3` is the number of restrictions, `k` is the number of predictors in the full model and `n` the sample size. Refer it to an F distribution on 3 and `n - k - 1` degrees of freedom. Two traps decide whether the test means anything. The models must be genuinely nested - the reduced model is the full model with those three coefficients forced to zero. And missing values must not silently drop different rows from the two fits, because then the two residual sums of squares are computed over different data and their difference is not the cost of the restriction.
go deeper
Know that several terms added together are tested as a group, and that the comparison requires both models fitted to exactly the same rows.
Be able to build the F statistic from two residual sums of squares and state both degrees of freedom without looking them up.
Demonstrate the operational care: identical row sets, genuine nesting, a sane residual variance estimate, and a decision that survives the interactions being correlated with their main effects.
Own the standard for how added model complexity is justified on your team, including what happens to reported p-values once terms are dropped after a test.
## Why a joint test Three interaction terms are one hypothesis, not three. The question a stakeholder is really asking is whether the interaction structure as a whole earns the three degrees of freedom it consumes. Reading three individual p-values answers a different question - whether each term is distinguishable from zero given the other two are in the model - and interactions are typically correlated with each other and with their main effects, so all three can look unimpressive while the block explains real variation. Three separate looks also raise the chance that one of them clears 5 percent for no reason. ## The statistic Let the **reduced** model omit the three terms and the **full** model include them, both fitted on the same `n` rows. With `RSS_r` and `RSS_f` the two residual sums of squares, `q = 3` restrictions and `k` predictors in the full model: `F = ((RSS_r - RSS_f) / q) / (RSS_f / (n - k - 1))` The numerator is the extra variation the block explains, per degree of freedom spent. The denominator estimates the residual variance from the full model. Under the null that all three coefficients are zero, this follows an F distribution with `q` and `n - k - 1` degrees of freedom. Note the direction: because the reduced model is nested inside the full one, `RSS_r >= RSS_f` always, so the numerator is never negative. Getting a negative difference means the models were not actually nested or were not fitted on the same rows. ## The two operational traps **Nesting.** The reduced model must be exactly the full model with those coefficients constrained to zero. Swapping in a different variable, changing a transformation, or re-binning a categorical breaks nesting and the F-test no longer has a distribution you can look up. **Identical rows.** This is the failure that silently produces wrong answers in real work. If one of the interaction inputs has missing values, the full model quietly drops those rows while the reduced model keeps them. Then `RSS_r` and `RSS_f` are computed over different observations, `RSS_r - RSS_f` partly measures which rows were included rather than the cost of the restriction, and the difference can even come out negative. The fix is to build the complete-case subset the full model requires and refit both models on it. ## Relationship to the overall F The overall F reported in a regression summary is exactly this test with the intercept-only model as the reduced model and `q = k`. Same formula, same logic; only the restricted model changes. Recognising that unifies two things candidates often learn as separate rules. ## Reading the outcome A significant block says the interaction structure explains more variation than chance would supply for three parameters. It does not say the interactions are large, interpretable, or stable, and it does not tell you which of the three carries the effect - that requires looking at the coefficients with their uncertainty. A non-significant block does not compel deletion. If the interaction encodes the hypothesis the analysis was designed to test, reporting a non-significant block with its confidence intervals is the finding. The more serious error is the other one: dropping terms because they failed a test, then reporting the reduced model's p-values as though no search had happened. Those p-values no longer have their nominal meaning, because the model was chosen using the same data. ## Practical checks around the test Before trusting the number, confirm the residual variance estimate in the denominator is sensible - severe heteroscedasticity or dependence between observations distorts it, and with clustered data the effective sample size is smaller than `n`, which inflates the statistic. Also sanity-check the size of the fitted interaction effects; with a large sample a statistically clear block can still represent an effect too small to change any decision. ## Answering in an interview Say 'one joint test, not three t-tests', write the statistic from the two residual sums of squares, name both degrees of freedom, and volunteer the identical-rows trap before you are asked. That last point is the one that signals having actually run this comparison on real data rather than only having read about it.
- Why not just read the three interaction p-values individually?Because the question is whether the block as a whole earns its three degrees of freedom. Interactions are usually correlated with each other and with their main effects, so each t-statistic can look weak while the set jointly explains real variation. Three separate looks also inflate the chance that one clears 5 percent by accident.
- What happens to the nested F-test if the two models are fitted on different numbers of rows?It becomes invalid. The two residual sums of squares then cover different observations, so their difference reflects row membership as well as the restriction, and it can even come out negative. Refit both models on the complete-case subset that the larger model requires, then compare.
- How does this relate to the overall F-test in a regression summary?The overall F is the special case where the reduced model is the intercept-only model and the number of restrictions equals the number of slopes. Same formula, same reasoning; the only thing that changes is which restricted model you compare against.
- If the block is not significant, must you drop the interactions?Not automatically. If the interactions encode the hypothesis the analysis was designed to test, or a stakeholder needs the estimate with its uncertainty, reporting a non-significant block is itself the finding. Dropping terms after a test and then reporting the reduced model's p-values as if no search occurred is the worse error.
saying these in an interview costs you the question
- Compares two models fitted on different row sets
- Reads three separate p-values instead of one joint test
- Applies the F-test to models that are not nested
- Subtracts the residual sums of squares in the wrong direction
- Reports post-selection p-values as if no search happened