In a Gaussian mixture, what do spherical, diagonal and full covariance each buy you?
answer
- shape of the constant-density ellipse
- round, axis-aligned, or tilted
- off-diagonal terms encode correlation
- one, d, then d(d+1)/2 numbers
- capacity paid for in data per component
basics
~20 sSpherical forces round components of one width; diagonal allows axis-aligned stretching; full allows tilt and correlation. Each step up fits shape better but costs more parameters, so full covariance needs far more data per component.
solid answer
~50 sThe covariance choice is a capacity dial. **Spherical** gives each component a single variance, so it is a round ball — 1 parameter per component. **Diagonal** gives one variance per feature, so components stretch along the axes but stay axis-aligned — `d` parameters. **Full** gives the whole covariance matrix, so a component can be elongated and tilted at any angle — `d(d+1)/2` parameters. Two elongated clusters lying at 45 degrees are the discriminating case: full covariance captures each with one tilted ellipse, while diagonal must approximate a tilted cigar with axis-aligned boxes and typically spends extra components doing it, and spherical is worse still. The tradeoff is bias versus variance in the parameter estimates: with `d = 50` and only a few hundred points per component, a full covariance has 1275 free numbers and its estimate becomes unstable or singular, so a diagonal fit is often the better model in practice.
go deeper
Recall the three shapes: spherical is a circle, diagonal is an axis-aligned oval, full is an oval that can tilt. Know that more flexibility means more parameters to estimate.
Give the parameter counts — 1, d, and d(d+1)/2 per component — and explain that only full covariance carries off-diagonal terms, which is what encodes correlation and tilt.
Diagnose from symptoms: too many components strung along a diagonal means a restriction problem, and a near-singular full covariance means too little data per component. Standardise or decorrelate features before restricting.
Own the policy: which covariance form is the default for your dimensionality and data volume, how it is validated, and whether the model's interpretability to stakeholders survives the choice.
## What the covariance controls Each Gaussian component in a mixture has a mean, which places it, and a covariance, which sets its **shape**: how wide it is along each direction and whether the directions are correlated. Drawing the component as an ellipse of constant density makes this concrete — the covariance chooses the ellipse's axis lengths and its orientation. Restricting the covariance restricts the family of shapes the component can take. The standard ladder, from most restricted to least: **Spherical** — the covariance is a single variance times the identity. The ellipse is a circle (a ball in higher dimensions). Every feature has the same spread and no feature is correlated with another. Cost: **1** free number per component. **Diagonal** — one variance per feature, zero off-diagonal terms. The ellipse can be stretched independently along each coordinate axis, but its axes must line up with the feature axes; it cannot tilt. Cost: **d** numbers per component. **Full** — the complete symmetric covariance matrix. The ellipse can have any axis lengths and any orientation, so the component can represent correlated features. Cost: **d(d+1)/2** numbers per component (the diagonal plus one triangle of off-diagonals). Some implementations also offer a *tied* variant, where all components share one covariance matrix; that trades per-component flexibility for a single, much better-estimated shape. ## The discriminating example Picture two clusters in two dimensions, both elongated cigars, both tilted at roughly 45 degrees, lying side by side. Feature 1 and feature 2 are strongly correlated within each cluster. - **Full covariance**: two components suffice. Each fits one tilted ellipse straight onto one cigar. The off-diagonal term is what encodes the tilt. - **Diagonal covariance**: no component can tilt. The best a single diagonal component can do with a 45-degree cigar is an axis-aligned blob that covers the cigar's bounding box, which also covers a lot of empty space and bleeds into the neighbouring cluster. EM typically responds by splitting each cigar into a chain of several small axis-aligned components — the fit improves, but you now have six components where two would do, and the component count no longer means what you wanted it to mean. - **Spherical covariance**: worse again, since it cannot even stretch; it needs still more, still rounder blobs. That failure is diagnosable. If a mixture keeps wanting more components than the domain suggests, and the fitted components come out in strings along a diagonal direction, the covariance restriction is the culprit, not the component count. ## Why you would ever restrict, then Because parameter count is not free. With `K` components in `d` dimensions, the full-covariance model has `K * d(d+1)/2` covariance parameters, and covariance estimation is data-hungry: to estimate a `d x d` covariance stably you want substantially more than `d` effective points, and to have it be non-singular you need at least `d + 1` points in general position contributing to that component. At `d = 50`, full covariance means 1275 numbers **per component**. If a component only owns 300 points, that estimate is wildly noisy; if it owns 40, the matrix is singular and the fit degenerates outright. So the choice is a bias-variance decision made on parameters: - **Full** = low bias in shape, high variance in the estimate. Good when `d` is small and you have plenty of data per component. - **Diagonal** = biased (it cannot represent correlation) but the estimate is stable, and it scales linearly in `d`. This is the workhorse in higher dimensions. - **Spherical** = strongly biased toward round, equal-width clusters, but extremely cheap and hard to break. Reasonable when features are on a common scale and clusters really are isotropic. A practical middle path: rotate or decorrelate the feature space first, so that the correlation the diagonal model cannot represent has largely been removed before fitting; a diagonal mixture in a decorrelated space is far less biased than a diagonal mixture in the raw space. ## Scaling matters more as you restrict Spherical assumes one width for all features, so it is only sensible when features share a scale — a mixture of amounts in dollars and counts in single digits will be dominated by the dollar axis. Diagonal tolerates differing scales, since it estimates a per-feature variance. Full covariance is the least sensitive to relative scaling, because it can adapt shape freely. If you standardise features before fitting, spherical and diagonal both become far more defensible. ## How to choose in practice Fit the candidates and compare them on a penalised likelihood or on held-out log-likelihood. The parameter count is exactly what the penalty is built from, so a full-covariance model has to earn its extra flexibility with a substantially better fit before it wins. Do not choose by eye on a two-dimensional projection: a projection can make correlated clusters look axis-aligned and mislead you toward diagonal. And check for degeneracy — if the full-covariance fit produces a component with a near-singular covariance, that is not a better fit, it is a collapse, and it is a reason to step down the ladder or add a regulariser.
- How many free covariance parameters does a full-covariance mixture have with 5 components in 20 dimensions?Each full covariance in 20 dimensions has 20*21/2 = 210 free entries, so five components carry 1050 covariance parameters, plus 100 mean parameters and 4 free weights. The diagonal version would carry 100 covariance parameters instead of 1050. That ten-fold gap is the whole argument: unless you have many thousands of points per component, the diagonal model will generalise better.
- You see a diagonal-covariance mixture fitting six components where the domain expects two — what do you suspect?That the true clusters are correlated and tilted, and the model is tiling each one with a string of axis-aligned pieces because it cannot rotate. Check whether the fitted means lie along a common diagonal direction and whether the within-cluster features are correlated. Refitting with full covariance, or decorrelating the features first, usually collapses the chain back to the expected count.
- Does restricting covariance change the EM algorithm itself?Only the covariance update in the M step. The E step is unchanged, and the weight and mean updates are unchanged. For a diagonal model you keep only the responsibility-weighted per-feature variances; for a spherical model you average those into one number per component. The monotone-likelihood guarantee still holds, because you are maximising over a restricted parameter set.
Fitting a coat to a body: spherical offers one size of round poncho, diagonal lets you pick height and width separately but the seams stay vertical, full lets the tailor cut on the bias and follow any slant.
saying these in an interview costs you the question
- Thinks diagonal covariance means the components cannot be stretched
- Believes full covariance is always the better model
- Says spherical components can represent correlated features
- Ignores that covariance parameters grow quadratically with dimension
- Chooses the covariance type by eye on a 2-D projection