Your standard errors are clustered by store but you only have 8 stores — what goes wrong?
answer
- asymptotics in clusters, not rows
- eight independent contributions only
- downward-biased and noisy variance
- t cutoff on G minus one df
- wild cluster bootstrap imposing the null
basics
~10 sCluster-robust standard errors are justified by the number of clusters growing, not the number of rows. With 8 stores they are downward-biased and noisy, so tests over-reject however many transactions each store contributes.
solid answer
~50 sThe clustered sandwich estimator is consistent as the number of clusters grows, so with 8 stores it is in a small-sample regime whatever the row count. The estimated variance is biased downward and is itself very noisy, which means real rejection rates can run several times the nominal 5%. Common rules of thumb want something like 30 to 50 clusters before the normal approximation is comfortable. The mitigations are, in order: apply the finite-sample correction factor, use a `t` reference distribution with `G - 1 = 7` degrees of freedom rather than the normal, and above all use a wild cluster bootstrap that imposes the null and resamples cluster-level signs. Note that with 8 clusters, two-point Rademacher weights give only `2^8 = 256` distinct draws, so p-values are coarse and six-point weights are preferred. Collecting more transactions per store does not help; only more stores do.
go deeper
Remember the headline rule: clustered standard errors need many clusters, and a handful of stores is a small sample no matter how many rows sit inside them.
Explain the mechanism: the variance is built from one contribution per cluster, so with eight it is both downward-biased and noisy, and the resulting test rejects far more often than its nominal level.
Show the fix ladder you would actually apply: the finite-sample correction, a t cutoff on G minus one degrees of freedom, and a null-imposed wild cluster bootstrap, plus the awareness that more rows per store buy nothing here.
Own the reporting norm: cluster counts published next to every clustered standard error, and a stated policy for what the team does below the threshold, including when a study should be declared underpowered at design time rather than patched at analysis time.
## Why the cluster count is the sample size A cluster-robust variance estimator builds the variance from one contribution per cluster. Those `G` contributions are the independent quantities being averaged, so all the usual large-sample reasoning applies to `G`, not to `n`. Eight stores means eight effective observations for the variance estimation step. Adding transactions inside a store makes each cluster contribution more precisely measured but does not add another independent contribution, which is why a dataset of two million transactions across eight stores is, for this purpose, a sample of size eight. ## The two failures **Downward bias.** Residuals used in the sandwich are fitted residuals, which are systematically too small because the regression has already fitted the cluster's data. With many clusters this bias is negligible; with eight it is severe, so the estimated standard error is too small. **Noise.** Even unbiased, a variance estimated from eight contributions is highly variable. A t-statistic formed by dividing by a noisy denominator has much fatter tails than the normal distribution, so using 1.96 as the cutoff rejects far too often. Reported rejection rates in the neighbourhood of 10-20% at a nominal 5% are not unusual with single-digit cluster counts, and the problem is worse when the clusters are unbalanced or when the predictor of interest is concentrated in a couple of clusters. Both failures point the same way: over-confidence. This is the dangerous direction, because the analysis looks rigorous — it says the word clustered — while being less reliable than it appears. ## What actually helps **Finite-sample correction.** Standard implementations multiply the variance by a factor of roughly `G / (G - 1)` times `(n - 1) / (n - k)`. With eight clusters that factor is small comfort but it is not nothing, and reporting whether it was applied is part of a serious answer. **A t reference distribution with `G - 1` degrees of freedom.** Instead of comparing to 1.96, compare to the two-sided 5% point of `t` with 7 degrees of freedom, which is about 2.365. That is a materially wider interval and it costs nothing. **Wild cluster bootstrap.** This is the accepted remedy. Refit the model with the coefficient of interest restricted to its null value, take the restricted residuals, and then repeatedly regenerate the outcome by flipping the sign of every residual in a cluster together, according to a random weight drawn once per cluster. Refit on each regenerated dataset, collect the t-statistics, and read the p-value off that bootstrap distribution rather than off the normal. Imposing the null is the part people skip and it matters: the null-imposed version performs far better with few clusters. There is a catch specific to tiny cluster counts. Two-point Rademacher weights take the values -1 and +1 with equal probability, one per cluster, so with 8 clusters there are only `2^8 = 256` distinct sign patterns. The bootstrap p-value can therefore only take a few coarse values, and no number of replications fixes that. Six-point weights give `6^8`, about 1.68 million, distinct patterns and are the usual recommendation below roughly a dozen clusters. **Aggregate and test at the cluster level.** With a cluster-level treatment, collapse each store to one number and run a small-sample test on eight values. It is low-powered and honest, and it is easy to explain. **Randomization inference.** If assignment to treatment was random across the eight stores, enumerate the possible assignments and place the observed statistic in that distribution. This gives exact inference without leaning on asymptotics at all. ## What does not help Collecting more transactions per store. Clustering at a finer level such as transaction or day, which simply hides the problem and returns you to the standard errors you should not have trusted in the first place. A pairs bootstrap that resamples rows, which reproduces the same independence error. And declaring victory because `n` is large. ## How to answer in an interview Lead with the sentence that the asymptotics run in the number of clusters, so eight is a small sample regardless of the row count. Then name the direction of the failure — over-rejection, intervals too narrow — and finally give the ladder of fixes ending in the wild cluster bootstrap. Adding that you would report the cluster count prominently alongside every clustered standard error, so a reader can judge for themselves, is the mark of someone who has been burned by this once already.
- Would collecting a year more of transactions from the same 8 stores help?Not for this problem. More rows sharpen each store's own contribution but leave the number of independent contributions at eight, and it is that count the variance estimator's reliability depends on. The interval will not become trustworthy, only more confidently wrong. The only structural fix is more stores; failing that, use a wild cluster bootstrap or randomization inference and be explicit about the limitation.
- Why do Rademacher weights become a problem with only 8 clusters?Because they take just two values, plus and minus one, drawn once per cluster, so eight clusters admit only 2^8 = 256 distinct sign patterns. The bootstrap distribution is therefore coarse and the p-value can only land on a handful of values, however many replications you run. Six-point weights give about 1.68 million patterns and are the usual recommendation when the cluster count is in single digits.
- Is clustering at a finer level, such as by transaction, a reasonable fallback?No, it is the worst option. Clustering below the level at which correlation actually operates leaves the within-store dependence unabsorbed, so the standard errors return to being too small while the write-up still claims clustered inference. That is more misleading than reporting a plain standard error with an honest caveat. If the correlation lives at the store level, the clustering has to be at the store level.
Estimating the spread of a distribution from eight numbers is shaky whichever eight numbers they are. Measuring each of the eight more carefully does not make the spread estimate steadier.
saying these in an interview costs you the question
- Says a large row count compensates for few clusters
- Clusters at a finer level to get more clusters
- Uses 1.96 rather than a t cutoff with G minus one df
- Bootstraps rows instead of whole clusters
- Assumes robust standard errors are always conservative