skip to content

A colleague refits the plan-tier model with Enterprise as baseline instead of Free and reports different significant tiers — what happened?

level: seniorimportance: should knowfreq 45%

answer

  1. same model, different parameterisation
  2. predictions and R-squared untouched
  3. each t-test asks a different question
  4. new Pro coefficient is b_Pro minus b_Ent
  5. significance belongs to a comparison, not a tier

basics

~20 s

Nothing about the fit changed — fitted values, residuals and R-squared are identical. Switching the baseline re-expresses each coefficient as a contrast against Enterprise instead of Free, so the individual t-tests now test different comparisons and can flip significance.

solid answer

~40 s

Re-basing is a reparameterisation, not a different model. The two fits span the same column space, so fitted values, residuals, residual sum of squares, R-squared and residual degrees of freedom are identical, and so is the joint F-test for whether plan tier matters at all. What changes is what each printed number means: with Free as baseline the Pro coefficient tests Pro against Free; with Enterprise as baseline it tests Pro against Enterprise. Those are different null hypotheses with different standard errors, so one can be significant and the other not. The arithmetic is mechanical — if `b_Pro` and `b_Ent` were the Free-based estimates, the Enterprise-based ones are `b_Pro - b_Ent` for Pro and `-b_Ent` for Free, with the intercept becoming `b0 + b_Ent`. Neither table is more correct; both need the baseline stated.

go deeper

for a junior

Recall that the choice of reference level is arbitrary as far as the fit is concerned, and that a dummy coefficient always means 'this level minus the reference level'.

for a middle

Be able to do the arithmetic on the spot: derive the new intercept and each re-based coefficient from the old ones, and explain why standard errors do not carry over unchanged.

for a senior

Recognise a reparameterisation instantly rather than chasing a data bug, and show the reporting discipline of naming the baseline and answering omnibus questions with the joint test rather than a scan of p-values.

for a principal

Own how contrasts reach decision-makers: which baseline is standard, how pairwise comparisons are pre-registered rather than harvested, and how model summaries are written so no one reads significance as a property of a tier.

## The situation A plan tier variable has three levels — Free, Pro, Enterprise. One analyst fits the model with Free as the reference; another refits with Enterprise as the reference. The predictions agree exactly, yet the two coefficient tables look different and disagree about which tiers are 'significant'. Both analysts are right, and the disagreement is about the question being asked, not the answer. ## What is invariant Dropping a different level does not change which functions of the data the model can represent. Both specifications span the same column space, so the least-squares projection of the outcome onto it is identical. Concretely, all of the following are unchanged to the last digit: - every fitted value and every residual - the residual sum of squares, R-squared and adjusted R-squared - the residual degrees of freedom and the residual standard error - the estimated difference between any two specific tiers - the joint F-test of the null that all tier effects are equal — the omnibus 'does plan tier matter' question - the coefficients and standard errors on any *other*, non-tier predictors in the model That last point is worth saying out loud in an interview: re-basing a factor leaves the rest of the model untouched. ## What changes The intercept and every tier coefficient change, because they are now measured from a different anchor. Start with Free as the baseline and estimates `b0`, `b_Pro`, `b_Ent`. The expected outcomes are `b0` for Free, `b0 + b_Pro` for Pro, `b0 + b_Ent` for Enterprise (other predictors at zero). Now switch the baseline to Enterprise. The new intercept must be Enterprise's expected outcome, so it is `b0 + b_Ent`. The new Free coefficient is Free minus Enterprise, which is `-b_Ent`. The new Pro coefficient is Pro minus Enterprise, which is `b_Pro - b_Ent`. Nothing was re-estimated in any meaningful sense; the same fitted surface has been rewritten in a different basis. Standard errors change too, and not by any simple rescaling: `SE(b_Pro - b_Ent)` depends on the covariance between the two original estimates, so it can be larger or smaller than either original standard error. Consequently each t-statistic and p-value changes, because each tests a different null hypothesis. ## Why significance can flip Suppose Pro sits close to Free but far from Enterprise. With Free as baseline, the Pro coefficient is small and its p-value large — Pro is 'not significant'. With Enterprise as baseline, the Pro coefficient is the large Pro-minus-Enterprise gap and its p-value is small — Pro is 'significant'. No contradiction: the first result says Pro is indistinguishable from Free, the second says Pro is distinguishable from Enterprise. Both are true statements about the same fitted model. The lesson generalises: **significance of a dummy coefficient is a property of a comparison, not of a level.** A sentence such as 'Pro was significant in the model' is incomplete to the point of being misleading; the sentence must name the baseline. ## How to handle it in practice First, fix the baseline deliberately rather than letting an alphabetical or first-seen default choose it. Prefer a level that is a real status quo or control, and prefer one with a healthy number of observations, since every contrast is measured against it and inherits its estimation noise; a sparse baseline makes every other coefficient noisier. Second, if the substantive question is 'does plan tier matter at all', answer it with the joint test across the tier coefficients, which is baseline-free, rather than by scanning individual p-values. Third, if several specific pairwise comparisons matter, recognise that you are doing multiple comparisons and treat them as such rather than harvesting whichever baseline makes the story cleanest. Refitting under many baselines and reporting the significant ones is a form of selective reporting. Fourth, report the contrasts stakeholders actually asked about, in their units, with the baseline named in the same sentence — 'Enterprise accounts convert 12 points above Free' rather than a bare coefficient table. ## The trap to avoid The failure mode this question is really testing is a colleague who concludes the model 'changed' or that one of the two fits is wrong, and starts hunting for a data bug. Recognising a reparameterisation on sight — same fit, different basis — is the senior signal here, followed by the discipline of never reporting a dummy coefficient without its reference level.

  • Which reported quantities should be byte-identical across the two fits?
    Fitted values, residuals, residual sum of squares, R-squared and adjusted R-squared, residual degrees of freedom and residual standard error, the estimated gap between any two named tiers, the joint F-test that plan tier has no effect, and the coefficients and standard errors of every non-tier predictor in the model.
  • How would you answer 'does plan tier matter at all' without depending on the baseline?
    Use the joint test of the null that all tier coefficients are zero — equivalently, compare the model with the tier dummies against the same model without them. That test is invariant to which level was omitted, unlike the individual t-tests, so it answers the omnibus question directly instead of inferring it from a scan of p-values.
  • Why is a sparsely populated level a poor choice of baseline?
    Every other coefficient is a contrast measured from the baseline, so all of them inherit the baseline's estimation noise. A level with few observations has a large standard error, which inflates the standard error of every contrast and makes the whole table look weaker than the data warrant. Prefer a well-populated, substantively meaningful reference level.
  • What is wrong with refitting under several baselines and reporting whichever contrasts came out significant?
    It is selective reporting across multiple comparisons. Every baseline exposes a different subset of the same set of pairwise contrasts, so cherry-picking the significant ones inflates the false-positive rate well beyond the nominal level. Decide which comparisons matter before fitting, and adjust for the number you intend to test.

saying these in an interview costs you the question

  • Claims the two fits are genuinely different models
  • Starts debugging the data over a reparameterisation
  • Says a tier 'is significant' without naming the baseline
  • Expects R-squared to change when the baseline changes
  • Picks the baseline that maximises significant coefficients

context