What do cluster-robust standard errors change in an OLS fit with many observations per user?
answer
- coefficients untouched, variance rebuilt
- sum within cluster, then square
- arbitrary correlation inside, independence across
- asymptotics in clusters, not rows
- precision capped by cluster count
basics
~20 sCluster-robust standard errors leave the coefficients unchanged and only rewiden the uncertainty around them. They permit arbitrary correlation inside each user while assuming independence across users, so reported precision reflects the number of users rather than the number of rows.
solid answer
~50 sThey change the variance estimate, not the fit. The coefficients from a regression on 40,000 page views contributed by 800 users are identical before and after; what changes is that the sandwich formula sums contributions cluster by cluster instead of row by row, so any pattern of correlation within one user is absorbed rather than assumed away. The result is normally a wider interval, sometimes dramatically wider, because the effective information is bounded by the 800 users and not the 40,000 rows. The key assumptions are that clusters are independent of each other and that there are enough of them, since the theory is asymptotic in the number of clusters. They are not a cure-all: they buy no efficiency, they do not fix omitted variables or a wrong functional form, and heteroscedasticity-robust errors that still assume independent rows do not fix clustering.
go deeper
Remember the headline: the fit does not move, only the uncertainty does, and it usually gets wider because rows from one user repeat information. Being able to say when to reach for them is enough here.
Explain the mechanics: the variance is assembled from independent cluster contributions instead of independent rows, so any correlation pattern inside a cluster is allowed and only independence between clusters is assumed.
Demonstrate judgment about the clustering level, including nested structures, and be able to say what the tool leaves unfixed: no efficiency gain, no bias repair, and unreliability once the cluster count is small.
Own the convention across the team: which level is the default for which dataset, when a structured model is preferred to a robust patch, and how results are reported so that a widened interval reads as honesty rather than as a weaker finding.
## What the estimator is doing Ordinary least squares gives you coefficients and, separately, an estimate of how uncertain those coefficients are. The default uncertainty calculation assumes every row's error is an independent draw. A cluster-robust variance estimator keeps the first part untouched and replaces the second: instead of building the variance from `n` independent row contributions, it builds it from `G` independent cluster contributions, where each cluster contributes the sum of its own rows before that sum is squared. Whatever correlation exists inside a cluster is therefore carried along rather than assumed to be zero, and only independence *between* clusters is required. This is why the name is a mouthful and the effect is simple. Point estimates: identical. Fitted values, residuals, R-squared: identical. Standard errors, t-statistics, confidence intervals, p-values: recomputed, usually larger. ## Why it usually widens Take 40,000 page views contributed by 800 users, so 50 views per user on average. Users differ persistently — heavy readers, mobile users, people who arrive from one referrer — and those persistent differences make the errors within a user positively correlated. Write the within-user correlation of the outcome as the intraclass correlation `ICC`. For a mean with equal cluster sizes `m`, the variance is inflated by roughly `1 + (m - 1) * ICC`. With `m = 50` and an `ICC` of 0.2 that factor is `1 + 49 * 0.2 = 10.8`, so the effective sample size is about `40,000 / 10.8`, roughly 3,700 rather than 40,000, and the standard error is about `sqrt(10.8)` — a bit over three times — larger than the naive one. Push the correlation to its extreme, where every view from a user is a perfect copy of that user, and the effective sample size falls to the number of users: 800. That ceiling is the intuition to carry into the interview — clustering can never buy you more independent information than you have clusters. The widening is not guaranteed in every case. If a predictor varies freely within a user and is uncorrelated within the cluster, the clustered and naive standard errors can come out similar. Negative within-cluster correlation, which is rarer, can even shrink them. But the direction of the typical surprise is upward, and a clustered standard error that comes out much *smaller* than the naive one deserves a second look rather than a celebration. ## Choosing the level The rule of thumb is to cluster at the level at which the correlation actually operates, and when levels are nested, at the coarsest level that plausibly matters. If users sit inside cities and there is a city-level shock — a marketing campaign, a holiday, an outage — clustering by user leaves that shock unabsorbed and the standard errors remain too small. Clustering by city absorbs both, because a city cluster contains whole users. The cost is that you now have far fewer clusters, and the estimator's reliability depends on that count. Clustering too finely is the common mistake and it fails silently: clustering by page view when the correlation lives at the user level is barely different from doing nothing. ## What it does not do **No efficiency.** The estimator is still ordinary least squares. A model that captures the structure directly, such as a random intercept per user, can be more efficient. Cluster-robust standard errors only stop you from lying about precision. **No bias repair.** If a user-level variable is omitted and correlates with your predictor, the coefficient is wrong and a wider interval around a wrong number does not help. **No rescue from too few clusters.** The theory is asymptotic in `G`, the number of clusters, not in `n`. With a handful of clusters the estimator is downward-biased and the test over-rejects. **No substitute for the right unit.** If the question is really about users, reporting a per-view effect with clustered errors may answer a question nobody asked. ## How to talk about it A strong answer states the invariance first — coefficients unchanged — then the assumption structure — arbitrary correlation within, independence across — then the practical consequence: your precision is governed by the number of clusters. Mentioning that the choice of clustering level is itself a modelling judgment, and that too fine a level quietly reproduces the original problem, is what separates a candidate who has used the tool from one who has read about it.
- Users sit inside cities — at which level should you cluster?At the coarsest level where correlation plausibly operates, so city if any city-level shock could move many users at once. A city cluster contains whole users, so clustering by city absorbs both layers, while clustering by user leaves the city shock unabsorbed and the errors too narrow. The tension is that the coarser level leaves you with far fewer clusters, and the estimator needs enough of them.
- Do clustered standard errors always come out larger than the default ones?Usually, but not necessarily. The inflation depends on both the within-cluster correlation of the errors and of the predictor, so a predictor that varies freely within a cluster may barely move. Negative within-cluster correlation can even shrink them. Clustered estimates are also noisier, so a clustered error that lands far below the default is a reason to check the clustering level rather than to relax.
- What problems do cluster-robust standard errors not solve?They do not make the estimator efficient, they do not correct omitted-variable bias or a wrong functional form, and they do not fix a mismatch between the unit you modelled and the unit the question is about. They also require a decent number of clusters, since the asymptotics run in the cluster count. They repair inference only, and only the part that comes from within-cluster dependence.
saying these in an interview costs you the question
- Clustered standard errors change the coefficient estimates too
- Clustering at the finest available level is the safe default
- Ordinary heteroscedasticity-robust errors already handle repeated users
- Clustered errors make the estimator efficient again
- They fix bias from an omitted user-level variable