How does a trained weight matrix's singular-value spectrum tell you whether to factorize it?
answer
- look at the decay, not the values
- cumulative squared singular values
- well conditioned is bad news
- stable rank in one number
- weight-space error is not task error
basics
~20 sBy how fast the singular values decay. A steep decay means a few directions carry the map, so a low-rank replacement loses little; a flat, well-conditioned spectrum means every direction matters and factorizing at any useful rank will cost accuracy.
solid answer
~50 sFor each candidate layer, take the singular values of its trained weight matrix and look at the cumulative share of squared singular values - the energy - captured by the top r. Find the smallest r that retains the share you are willing to bet on, then compare it against the break-even rank `m * n / (m + n)`. If the rank you need is comfortably below break-even, that layer is factorizable; if the spectrum is nearly flat, so the matrix is well conditioned and its effective rank is close to full, no cheap factorization exists and you should leave it alone. Two cautions: energy retained in weight space is only a proxy for task loss, because the input distribution decides how much a discarded direction actually mattered, and errors compound across layers. So use the spectrum to *shortlist* layers and ranks, then confirm with a per-layer sensitivity sweep on validation data and a recovery fine-tune.
go deeper
Know that singular values rank how much a matrix stretches along each of its directions, and that many near-zero ones mean the matrix has redundancy a smaller pair of matrices could imitate.
Explain retained energy as the cumulative share of squared singular values, and be able to state why the rank you need must also come in under the break-even point for the factorization to save anything.
Demonstrate the full workflow: size the layers first, screen with spectra, then confirm with a one-layer-at-a-time sensitivity sweep on validation data, allocate rank non-uniformly, and fine-tune to recover.
Own the argument that weight-space error is the wrong objective and that the budget should be allocated by measured accuracy-per-parameter across layers. Be ready to decide when analysis time is better spent than the compression it buys.
## What the spectrum is telling you Every real matrix `W` has a set of non-negative singular values, conventionally sorted from largest to smallest. They describe how much the linear map stretches along each of a set of orthogonal input directions, mapped to a set of orthogonal output directions. A large singular value means the map does a lot along that direction; a near-zero one means the map barely acts there at all. Plotted in order, they form the matrix's spectrum, and the *shape* of that plot is the single most informative artifact when deciding whether a layer will survive being factorized. If the spectrum falls off a cliff - a handful of large values then a long tail of small ones - the layer is doing something that a rank-r matrix can imitate closely. If it decays slowly and every value is of similar magnitude, the matrix is well conditioned in the classical sense (largest over smallest is near 1), its map genuinely uses all its directions, and imposing a low rank throws away a proportional chunk of it. A well-conditioned layer refuses to factorize, and no amount of fine-tuning changes the fact that the starting approximation was poor. ## Turning that into a rank The usual quantitative criterion is retained energy. Squared singular values sum to the squared Frobenius norm of the matrix, so the ratio of the top-r squared values to that total is the fraction of the matrix's "mass" preserved by the best rank-r approximation, which is obtained by truncating the decomposition. Pick a threshold - many teams start at 95% or 99% - and read off the smallest r that clears it. A useful scalar shortcut is the stable rank: Frobenius norm squared divided by spectral norm squared, i.e. the sum of all squared singular values divided by the largest one. It answers "how many directions' worth of energy is here?" in one number, without a plot, and it is bounded above by the true rank. A layer whose stable rank is a small fraction of its dimensions is a candidate; one whose stable rank is close to full is not. Then apply the economic filter. Factorization only saves anything while the rank stays below `m * n / (m + n)`. Both facts must hold at once: the layer must *tolerate* a rank that is also *below break-even*. Plenty of layers tolerate rank 700 out of 1024 - and rank 700 in a 1024-square layer costs more than the dense original. That layer is not factorizable in any useful sense even though 95% of its energy fits. ## Why the spectrum is a shortlist, not a verdict The energy criterion measures error in *weight space*: how far `A B` is from `W`. What you actually care about is error in *task space*: how far the network's outputs and loss move. These differ for two reasons. First, inputs are not isotropic. A direction with a small singular value can still matter enormously if the activations that hit this layer have large variance along it, and a direction with a large singular value can be nearly irrelevant if nothing ever excites it. This is why data-aware factorization exists: instead of minimising weight-space error, you minimise the change in the layer's *outputs* on a calibration batch, which weights directions by how often and how strongly they are actually used. That routinely finds a lower usable rank than the plain spectrum suggests. Second, errors compound. A layer whose output shifts slightly feeds a layer whose behaviour was tuned to the unshifted distribution, and small per-layer errors can accumulate or interact non-linearly by the time they reach the loss. A model with ten layers each at 99% retained energy is not a model at 99% of its accuracy. ## The workflow that actually works 1. Count parameters per layer and rank layers by size. There is no point analysing a layer holding 0.3% of the model. 2. Compute each big layer's spectrum, its retained-energy curve and its break-even rank. Discard layers whose required rank sits above break-even. 3. Run a per-layer sensitivity sweep: factorize exactly one layer at a time at a couple of candidate ranks, measure validation loss with everything else untouched. This produces an empirical accuracy-per-parameter curve per layer, which is the thing you should be allocating against. 4. Allocate the compression budget non-uniformly - cheap layers get aggressive ranks, sensitive ones get left alone or barely trimmed. Uniform ranks across a network are almost always the wrong answer. 5. Fine-tune the whole factorized model to recover, and re-measure. Recovery capacity differs per layer too, so a layer that looked bad before fine-tuning may be fine after it, and the ordering can change. ## Which layers usually win The widest, most over-parameterized matrices - large hidden-to-hidden projections, oversized classifier heads, big input tables - both hold the parameters and tend to have the most redundant spectra. Narrow layers lose twice over: they hold few parameters, and their break-even rank is so small that the usable window barely exists.
- A layer keeps 99% of its spectral energy at rank r, yet accuracy drops. Why?Because retained energy measures distance in weight space, not the effect on the loss. The discarded directions may be exactly the ones the input distribution excites most strongly, so a small weight-space error becomes a large output error. And per-layer errors compound down the network, so several layers each at 99% do not leave the model at 99%. The fixes are data-aware factorization, which minimises output change on a calibration batch, and a measured per-layer sensitivity sweep.
- What single number summarises how many directions of a weight matrix carry real energy?The stable rank: the squared Frobenius norm divided by the squared spectral norm, equivalently the sum of all squared singular values over the largest one. It is bounded above by the true rank and drops toward 1 when one direction dominates. It gives a fast per-layer screen without plotting a full spectrum, though it should still be compared against the break-even rank before you conclude a layer is worth factorizing.
- Which layers of a trained network usually turn out to be the most factorizable?The widest and most over-parameterized ones - large hidden-to-hidden projections, oversized classifier heads and big input tables. They hold most of the parameters and are typically the ones with the most redundant spectra. Narrow layers are doubly unattractive: they contain few parameters to save, and their break-even rank is so low that there is almost no gap between a rank the layer tolerates and a rank that saves nothing.
saying these in an interview costs you the question
- Applies one uniform rank to every layer in the model
- Treats 99% retained energy as proof accuracy is safe
- Reads a flat spectrum as a healthy sign for factorization
- Ignores break-even rank and factorizes tiny layers
- Never runs a recovery fine-tune after truncating
- Analyses spectra before checking where the parameters actually are