Why can a Gaussian mixture's log-likelihood run to infinity during EM, and how do you stop it?
answer
- the objective has no global maximum
- determinant in the denominator
- a component sitting on duplicate rows
- monotone climbing offers no protection
- put a floor under every eigenvalue
basics
~20 sA component can centre on a few near-identical rows and shrink its covariance toward zero, so their density diverges and the likelihood is unbounded above. Fix it with a covariance floor or a restricted covariance form.
solid answer
~50 sMaximum likelihood for a Gaussian mixture is a badly posed problem: the likelihood surface has singularities. Put one component's mean exactly on a data point and let its variance go to zero — that point's density diverges, so the whole product diverges, and the log-likelihood goes to plus infinity while the fit explains nothing. EM is a hill-climber, so if it wanders into that basin it will happily ride the singularity down. The classic trigger is a component capturing three near-duplicate rows in a dataset with duplicated or low-variance records. Symptoms: a component with a tiny mixing weight, a determinant near zero, and a log-likelihood that jumps by orders of magnitude in a single sweep. Fixes, in order of preference: add a small ridge to every covariance diagonal, restrict to diagonal or spherical covariance, deduplicate the rows, or detect a collapsing component and reinitialise it elsewhere.
go deeper
Know the headline: a Gaussian component can shrink onto a couple of points, its density there becomes enormous, and the fit is meaningless even though the score looks perfect.
Explain the mechanism through the density's covariance determinant, and name the standard remedy of adding a small constant to the covariance diagonal so no direction can collapse to zero width.
Show you would catch it in a pipeline — assertions on effective counts, weights and determinants, suspicion of an unusually high log-likelihood, plus data hygiene on duplicate rows and near-constant features.
Own the position that the unpenalised objective is ill-posed, and decide what your teams fit instead: a regularised or prior-penalised objective, with the floor and the component-count limits written down as defaults rather than rediscovered per project.
## The singularity, precisely For a Gaussian mixture the likelihood of the data is a product over points of `sum_k pi_k * N(x_i | mu_k, Sigma_k)`. Now do the following thought experiment. Pick any single data point `x*`. Set component 1's mean exactly to `x*` and shrink its covariance toward the zero matrix. The Gaussian density has the determinant of the covariance in the denominator under a square root. As the covariance shrinks, the density evaluated *at its own centre* grows without bound. So the term for `x*` in the product contains `pi_1 * N(x* | x*, Sigma_1)`, which diverges. Every other point still contributes a finite positive amount from the remaining components. The product therefore diverges, and the log-likelihood goes to `+infinity`. This is not a numerical artefact. It is a genuine property of the objective: **the maximum-likelihood problem for a Gaussian mixture with unconstrained covariances has no global maximum**. The supremum is infinite and it is attained only in a degenerate limit. Any "solution" you actually use is a well-behaved *local* optimum, which is a slightly uncomfortable but standard state of affairs. A single point is the cleanest illustration, but the practical trigger is usually a small clump: three near-duplicate rows — the same customer recorded three times, a sensor stuck on one reading, a default value repeated — let a component shrink onto them with a covariance that is singular in every direction where the rows agree exactly. ## Why EM in particular falls in EM has a monotone-likelihood guarantee, which sounds protective and is not. It guarantees the objective never decreases; a run marching toward a singularity is *satisfying* that guarantee perfectly. Nothing in the algorithm knows the difference between climbing toward a good fit and climbing into a spike. The mechanics are self-reinforcing. Once a component's covariance is small, its density at nearby rows becomes enormous, so the E step gives it responsibility ~1 for those rows and ~0 for everything else. Its effective count `N_k` collapses to about three. The M step then computes the covariance as the responsibility-weighted scatter of three points about their own mean — which, in `d` dimensions with `d >= 3`, is singular or nearly so. The next E step sharpens further. The loop tightens each sweep. ## Recognising it The symptoms are specific enough to alert on: - A component whose mixing weight is a tiny fraction of a percent, or whose effective count `N_k` is smaller than the dimension. - A covariance determinant approaching zero, or a condition number exploding; equivalently a near-zero eigenvalue. - A log-likelihood that rises smoothly for many sweeps and then leaps by orders of magnitude. - Refits from different initialisations that disagree wildly, with the "best" log-likelihood belonging to an obviously useless partition. - A numerical failure when inverting the covariance — often the first thing you actually see. Be suspicious of a fit that scores far better than every other run. On this objective, a dramatically higher log-likelihood is more often a collapse than a discovery. ## The fixes **Covariance regularisation (the default).** Add a small positive constant to the diagonal of every covariance in the M step. This puts a floor under every eigenvalue, so no direction can shrink to zero, the determinant stays bounded away from zero, and the singularity is removed from the parameter space entirely. The constant should be small relative to the feature variances — which is another argument for standardising features first, since it makes one floor value meaningful across all of them. This is equivalent to a mild prior on the covariance, and it changes the objective slightly: you are maximising a penalised likelihood that actually has a maximum. **Restrict the covariance form.** Diagonal or spherical covariances are much harder to collapse, because a spherical component has a single variance to drive to zero and needs the clump to be degenerate in *every* direction at once. This costs modelling flexibility, so it is a tradeoff rather than a free win. **Fix the data.** Deduplicate exact and near-exact rows, and look hard at constant or near-constant features — a column that is the same value for 90 percent of rows invites a collapse along that axis. Dropping a zero-variance feature removes a whole direction in which any component can degenerate. **Detect and restart.** Watch each component's effective count and covariance determinant during the run; when one crosses a threshold, reinitialise that component at a random point or drop it and continue. Cruder than regularisation but effective, and it keeps you from silently shipping a collapsed fit. **Reduce the component count.** Collapses become common when you ask for more components than the data supports. If the fit is stable at 4 components and collapses at 9, that is information about the component count, not just a numerical nuisance. **Go Bayesian in spirit.** Placing a prior on the covariance and maximising the posterior rather than the likelihood is the principled version of the ridge fix: the prior vanishes as the covariance shrinks, so the degenerate solution is no longer optimal. ## The interview point The answer that lands states three things: the likelihood is genuinely *unbounded above* (not merely ill-conditioned), EM's monotone guarantee offers no protection because a collapse increases the likelihood, and the standard fix is a floor on the covariance. Candidates who describe this as an overfitting problem or a convergence-speed problem have not understood the shape of the objective.
- Doesn't EM's monotone-likelihood guarantee protect against this?No — it is what enables it. The guarantee says each sweep never decreases the likelihood, and a run collapsing onto duplicate rows increases it enormously, so the guarantee is satisfied throughout. EM is a hill-climber on an objective with infinitely tall spikes; the guarantee describes the climbing, not the terrain. That is precisely why the fix has to change the objective or the parameter space.
- How would you detect a collapse automatically in a training pipeline?Assert on each component after every fit: effective count above a floor such as the dimension plus a margin, mixing weight above a minimum, and covariance determinant or smallest eigenvalue above a threshold. Also flag a log-likelihood that jumps by orders of magnitude between sweeps. Fail the run and either reinitialise the offending component or reduce the component count rather than shipping the model.
- Is this the same as overfitting?Related but distinct. Overfitting is a model that generalises worse than it fits; here the fitted object is degenerate even on the training data, and the objective itself has no maximiser. A collapsed component describes three rows and nothing else, so held-out likelihood is usually terrible — but the root cause is an unbounded objective, not merely excess capacity.
It is like scoring a map by how tightly it hugs the coastline: the winning map is one that traces a single pebble at infinite magnification and ignores the continent. The score is real; the map is useless.
saying these in an interview costs you the question
- Calls it a numerical precision bug rather than an unbounded objective
- Believes EM's monotone guarantee prevents degenerate solutions
- Says more EM iterations or a tighter tolerance will fix it
- Treats the highest log-likelihood run as automatically the best
- Confuses it with ordinary overfitting on held-out data