In a random-intercept model of test scores nested in schools, what does the intercept variance tell you?
answer
- one number for how schools differ
- two variance components, between and within
- feeds the intraclass correlation
- dummies describe only sampled schools
- shrinkage depends on cluster size
basics
~20 sIt estimates how widely school mean scores spread around the overall mean, as one variance on the outcome scale. Divided by the total variance it gives the intraclass correlation, and it treats the sampled schools as draws from a wider population.
solid answer
~50 sA random-intercept model writes a pupil's score as `overall mean + school offset + pupil noise`, with the school offsets drawn from a distribution of variance `var(u)`. That variance component is a one-number summary of how much schools differ once the predictors are accounted for, on the outcome's own scale, and `var(u) / (var(u) + var(e))` is the intraclass correlation. School dummy variables cannot tell you this: they spend a coefficient per school, describe only the schools you sampled, and absorb any school-level predictor you wanted to test. The random intercept treats schools as exchangeable draws, so it generalises beyond them, allows school-level predictors, and partially pools each school's estimate toward the overall mean in proportion to how noisy that school's data are. The price is assuming the school effect is uncorrelated with the predictors; when it is not, dummies are safer.
go deeper
Know the shape of the model: a school-level offset plus pupil-level noise, with the offsets sharing one distribution. Recognising that the intercept variance summarises how much schools differ is enough here.
Explain the two variance components, how their ratio gives the intraclass correlation, and why one variance parameter replaces thirty-nine dummy coefficients.
Be ready to defend the choice in a real analysis: name the no-correlation assumption behind random effects, describe shrinkage and how cluster size drives it, and say when you would switch to dummies or to robust standard errors instead.
Own the reporting standard. Decide whether the team publishes shrunken cluster estimates or raw ones, be able to explain to stakeholders why a small school's headline number was pulled toward the average, and set the bar for when modelling the structure is worth its assumptions.
## The model For pupil `i` in school `j`, a random-intercept model says `y_ij = b0 + b1 * x_ij + u_j + e_ij` with `u_j` drawn from a distribution with mean 0 and variance `var(u)`, and `e_ij` drawn independently with variance `var(e)`. There are two variance components: `var(u)`, the between-school variance, and `var(e)`, the within-school residual variance. The estimated `var(u)` is what the question asks about. ## What the variance component says **It quantifies heterogeneity in one number.** If test scores are on a 0-100 scale and `var(u)` comes out at 25, then school means have a standard deviation of 5 points around the overall mean after adjusting for the pupil-level predictors. That single number answers a real question — how much does the school you attend matter? — in a way a table of forty school coefficients does not. **It gives you the intraclass correlation.** `ICC = var(u) / (var(u) + var(e))` is the share of residual variance sitting between schools and the correlation between two pupils in the same school. This is the quantity that drives how badly an analysis assuming independence would have misled you. **It implies a structure for the errors.** A random intercept says the correlation between any two pupils in the same school is the same constant, whichever two you pick — an exchangeable, or compound-symmetric, structure. That is an assumption, and it is testable in spirit: if pupils in the same class within a school are more alike than pupils in different classes, you need a further level, not just one random intercept. **It defines the population.** The schools are modelled as draws from a population of schools, so `var(u)` describes that population and the model can say something about a school you have not sampled. School dummies describe forty specific schools and stop there. ## Against school dummies A fixed-effects specification with a dummy per school (minus a reference) is the classic alternative, and the comparison is a favourite interview probe. **Dummies cost degrees of freedom.** Forty schools cost thirty-nine parameters; the random intercept costs one variance parameter. With many small clusters that difference is large. **Dummies absorb school-level predictors.** Anything constant within a school — its funding band, urban or rural, a policy it adopted — is collinear with the dummies and cannot be estimated. If a school-level predictor is the point of the study, dummies rule it out. **Dummies give no pooling.** Each school's estimate is its own data alone, so a school with five pupils gets a wildly noisy estimate. The random-intercept model shrinks each school's predicted intercept toward the overall mean by a factor of roughly `var(u) / (var(u) + var(e) / n_j)`: a large school moves barely at all, a school of five moves a long way. That partial pooling is why predicted school effects from a mixed model are more stable and rank schools more sensibly than raw school means. **But dummies are robust where the random intercept is not.** The random-intercept estimator assumes `u_j` is uncorrelated with the predictors. If schools with better intakes also happen to apply the treatment more, that assumption fails and the pupil-level coefficient is biased toward the school-level confounding. Dummies absorb every time-invariant school characteristic, observed or not, and identify the coefficient purely from within-school comparisons. When the identification of a causal effect is the goal and school-level confounding is plausible, dummies are the conservative choice; a common compromise is to include the cluster mean of the predictor as an extra regressor, which recovers the within-school contrast while keeping the random-effects machinery. ## When to use which, in one breath Use a random intercept when the clusters are a sample from a population you want to describe, when cluster-level predictors matter, when clusters are numerous and small, or when the heterogeneity itself is a deliverable. Use dummies when the clusters are the entire population of interest, when cluster-level confounding is the threat, or when you want inference that does not lean on a distributional assumption for the cluster effects. And if all you want is honest uncertainty for a population-average slope, without any claim about the structure, cluster-robust standard errors on the plain fit deliver that with fewer assumptions — they simply give you no variance component and no shrunken school predictions. ## Reading the output A warning sign is a between-school variance estimated at or near zero, which usually means either the schools genuinely do not differ after adjustment or there are too few schools to identify the component; with a handful of clusters the variance component is poorly estimated and its standard error is not to be trusted. Reporting the intraclass correlation alongside the raw variance is good practice, since a variance in squared score units is hard for a stakeholder to interpret while a share of variance is not.
- When would you prefer school dummy variables over a random intercept?When the school effect is plausibly correlated with the predictor of interest. Dummies absorb every time-invariant school characteristic and identify the coefficient from within-school comparisons only, so they are robust to school-level confounding that would bias the random-intercept estimate. The costs are thirty-nine parameters instead of one, no variance component to report, and no ability to estimate any school-level predictor.
- What does partial pooling do to a school with only five pupils?It pulls that school's estimated intercept strongly toward the overall mean, because five pupils give a noisy school mean and the model weights it by roughly var(u) / (var(u) + var(e) / 5). A school with two hundred pupils barely moves. The effect is that extreme rankings driven by tiny samples are damped, which is usually the behaviour you want when reporting school effects.
- When would cluster-robust standard errors be the better tool than a mixed model?When you want a population-average slope with honest uncertainty and no commitment to a correlation structure. Robust errors assume only independence across clusters, so a misspecified within-cluster structure cannot hurt them, and they are simpler to defend. You give up efficiency, the variance component, and the shrunken cluster predictions, so if the heterogeneity itself is the deliverable the mixed model earns its assumptions.
- What does a between-school variance estimated at essentially zero suggest?Either the schools genuinely do not differ once the predictors are adjusted for, or there are too few schools to identify the component. Variance components are bounded below at zero, so estimates pile up on that boundary when information is thin, and their standard errors are unreliable there. With a small number of clusters, treat a zero estimate as uninformative rather than as evidence that clustering can be ignored.
saying these in an interview costs you the question
- Calls the variance component the average school effect
- Thinks random intercepts and school dummies are interchangeable
- Ignores that random effects assume no correlation with predictors
- Expects to estimate a school-level predictor alongside school dummies
- Reads a near-zero variance component as proof clustering is harmless