skip to content

Why can AIC and BIC select different models from the same 10,000-row dataset?

level: seniorimportance: nice to knowfreq 36%

answer

  1. both are fit plus a complexity penalty
  2. the penalties differ by one factor only
  3. one penalty grows with the sample size
  4. log n overtakes 2 very early
  5. different goals: prediction versus recovering the truth

basics

~20 s

Both add a parameter penalty to the same fit term, but AIC charges 2 per parameter while BIC charges log n - about 9.2 at n = 10,000. BIC is far stricter there, so it favours the smaller model.

solid answer

~50 s

`AIC = -2 log L + 2k` and `BIC = -2 log L + k log n`, where `log L` is the maximised log-likelihood, `k` the number of estimated parameters and `n` the sample size; lower is better for both. They share the fit term and differ only in the price of a parameter: AIC always charges 2, BIC charges `log n`, which passes 2 once n is above roughly 8. At n = 10,000, `log n` is about 9.2, so BIC is more than four times as strict and drops borderline terms that AIC keeps. The disagreement is by design: AIC targets predictive accuracy without assuming the true model is among the candidates, while BIC is consistent, homing in on the true model as n grows if it is in the set. Neither number means anything alone; only differences between models fitted to the same observations are interpretable.

go deeper

for a junior

Recall that AIC and BIC are lower-is-better scores trading fit against the number of parameters, and that only differences between models are meaningful.

for a middle

Write both formulas, point at the single term that differs, and say which criterion tightens as the sample grows and why.

for a senior

Show that you check comparability first - same rows, same response scale, consistent parameter counts - before trusting any ranking the numbers produce.

for a principal

Take an explicit position on what the selection criterion is optimising for on a given project, and stop the team from shopping between criteria until one agrees with a favoured model.

## The two criteria Both are penalised fit scores of the form 'badness of fit plus a price for complexity', and for both, lower is better: `AIC = -2 log L + 2k` `BIC = -2 log L + k log n` Here `log L` is the maximised log-likelihood of the fitted model, `k` counts the estimated parameters, and `n` is the number of observations. The first term rewards fitting the data; the second discourages spending parameters to do it. The fit term is identical in the two formulas - everything that separates them lives in the penalty. ## Where the disagreement comes from AIC charges a flat 2 per parameter, no matter how much data you have. BIC charges `log n`, which grows with the sample. The crossover is where `log n = 2`, i.e. `n` around 7 to 8; above that BIC is the harsher of the two, and the gap widens as data accumulates. At `n = 10,000`, `log n` is roughly 9.21, so each extra parameter costs about 9.2 under BIC against 2 under AIC. A term that improves the fit term by, say, 5 units is worth keeping under AIC and not worth keeping under BIC. That is the whole mechanism behind a disagreement on one dataset: the larger the sample, the more often BIC returns the smaller model. ## Why they were built differently The difference reflects different goals rather than one being wrong. - **AIC** is derived as an approximately unbiased estimate of the expected information loss when the fitted model is used to represent the data-generating process. It does not assume the true model is among your candidates, and it is oriented toward accurate prediction. It is not consistent: as `n` grows, AIC retains a non-vanishing chance of choosing a model slightly larger than the true one. - **BIC** comes from a large-sample approximation to the posterior probability of each model under roughly equal prior weights. It is consistent: if the true model is in the candidate set, the probability that BIC selects it tends to 1 as `n` grows. The price is that when the true model is not in the set - the normal situation with real data - BIC can under-select and discard structure that would have improved predictions. So 'AIC keeps more terms, BIC keeps fewer' is not a preference between right and wrong. It is a choice between leaning toward predictive richness and leaning toward parsimony and recovering a sparse structure. ## Conditions for a valid comparison Both criteria compare candidates only under strict conditions: - **Same observations.** If one model dropped rows with missing values, its log-likelihood is computed over fewer terms and is not commensurable with the other's. This is the most common way a comparison silently breaks. - **Same response variable on the same scale.** A model of `y` and a model of `log y` have log-likelihoods with respect to different densities; their scores cannot be lined up without a change-of-variable correction. - **Consistent parameter counting.** Count everything estimated. For a Gaussian linear model that is the intercept, every slope, and the error variance. Absolute values shift with counting conventions and with dropped additive constants, so the requirement is to count identically for every candidate. Unlike the F-test, neither criterion requires the models to be **nested**, which is precisely why they are reached for when comparing structurally different candidates. ## Reading the numbers The value of an AIC in isolation is meaningless - it carries arbitrary additive constants. Only differences within one candidate set on one dataset are interpretable. Common rules of thumb treat a difference under about 2 as weak evidence between two models and a difference beyond roughly 10 as strong evidence against the higher-scoring one, but these are soft guidance, not tests with error rates. If AIC and BIC disagree, the honest report is that the evidence is not decisive and the choice depends on whether the model is meant to predict well or to identify a parsimonious structure - not a hunt for whichever criterion endorses the model you already liked.

  • Do AIC and BIC require the compared models to be nested?
    No, and that is a large part of their appeal - they can rank structurally different candidates, such as different variable sets or different link functions. What they do require is the same observations, the same response on the same scale, and parameters counted the same way for every candidate.
  • How large a difference in AIC between two models is worth acting on?
    Rules of thumb put differences under about 2 in the 'both supported' zone and treat differences beyond roughly 10 as strong evidence against the higher-scoring model. Treat them as soft guidance: the value is a relative score within one candidate set, not a test with a controlled error rate.
  • How should k be counted for a Gaussian linear model?
    Count every estimated parameter: the intercept, each slope, and the error variance. The absolute score shifts with counting conventions and with dropped constants, so what matters is counting identically across candidates. An inconsistent count is a common way to make one model look better than it is.

AIC charges a flat entry fee for every parameter. BIC charges a fee that rises with how much data you brought, so on large samples it is the tougher doorman.

saying these in an interview costs you the question

  • Reports a single AIC value as if it were meaningful alone
  • Compares criteria across models fitted to different rows
  • Believes AIC and BIC require nested models
  • Assumes BIC always penalises more heavily regardless of n
  • Switches criteria until one endorses the preferred model

context