Your 80M stored 2048-dimension embeddings show only about 30 large singular values — what happened?
answer
- declared width is only a ceiling
- mean-centre before taking singular values
- C classes need only C-1 directions
- effective rank from the spectrum's entropy
- instance identity was never in the loss
basics
~20 sThe embeddings occupy a roughly 30-dimensional subspace of a 2048-dimensional space — dimensional collapse. The training objective and data diversity, not the layer width, decided that, so the remaining coordinates cost storage and carry almost no information.
solid answer
~50 sStack a large sample of embeddings, mean-centre, and look at the singular values: if about 30 dominate and the rest fall away, the representation lives in a ~30-dimensional subspace. The usual cause is the objective. If the backbone was trained to classify 32 marketplace categories, the head is a 32-row linear map and gradients reach the features only through it, so the features are pressed into separating 32 groups — the between-class structure needs at most `C - 1` directions, and within-class variation gets no incentive to survive. That is fatal for near-duplicate dedup, which is an instance-level task: two different bicycles were never asked to differ. I would quantify it as an effective rank I can track across runs, then fix the training signal — finer or instance-level supervision, or a self-supervised objective with an explicit anti-collapse term — rather than widening the layer.
go deeper
Know that the number of dimensions a vector declares is not the amount of information it holds, and that you can inspect how spread out a set of embeddings really is.
Explain the measurement: mean-centre a large sample, take singular values, and summarise them as an effective rank. Be able to say why a C-class objective needs only about C-1 directions.
Diagnose from the production symptom — duplicates and non-duplicates overlapping on similarity — trace it to the objective rather than the width, and propose a training-signal fix with a measurable before-and-after.
Decide whether low effective rank is even a defect for the portfolio of tasks the encoder serves, and weigh a retraining programme against simply shrinking the stored vector to what it truly carries.
## Measuring it before naming it Take a large random sample of the stored vectors, say a few hundred thousand rows, and stack them into a matrix. **Mean-centre the columns** — skipping this is the single most common mistake, because an uncentred set with a large common offset shows one enormous leading singular value that is just the mean, not structure. Then take the singular values `s_1 >= s_2 >= ...`. Squared and normalised to sum to one, they give the share of variance in each orthogonal direction. A useful single number is the **effective rank**: normalise the squared singular values into a distribution `p_i`, compute its entropy, and report `exp(entropy)`. This is a smooth count of *how many directions actually participate*. Alternatively, report the number of directions needed to reach 99% of the variance — cruder, but easier to explain in a review. Either way, pin the sample size and the centring so the number is comparable between training runs, and track it as a training-time metric rather than discovering it in production. When effective rank is 30 out of a declared 2048, the representation has undergone **dimensional collapse**: not the degenerate case where every vector is identical (rank one, complete collapse), but the partial case where the vectors span a thin subspace of the space you are paying to store. ## Why the objective sets the rank The declared width is a ceiling. What fills it is the training signal. With supervised classification over `C` classes, the head is a `C x d` matrix and the features receive gradient **only** through it. The loss is satisfied as soon as the `C` class groups are linearly separable, and separating `C` groups needs at most `C - 1` directions of between-class structure. Within-class variation earns nothing and, in the late phase of training, actively shrinks — the published *neural collapse* phenomenon describes exactly this: class means converge toward a maximally-separated simplex configuration and within-class scatter goes toward zero. A 32-category label set producing about 30 usable directions is not a coincidence; it is the objective working as specified. Two other causes are worth checking before you blame the head. **Under-diverse data**: if the training set is narrow, the features never needed to spread. And **a collapse-prone unsupervised objective**: an objective that only pulls similar things together, with no repulsion, no negatives and no explicit variance or decorrelation term, has a trivial optimum where the encoder outputs a constant. Methods in that family carry an anti-collapse mechanism precisely because of this. ## Why it matters here and when it does not The key judgment — and the one interviewers are actually testing — is that **low effective rank is not bad in the abstract; it is bad relative to the task**. If the only downstream use were the same 32-way categorisation, 30 directions is exactly right and the collapse is efficient compression. It is bad here because the marketplace task is *instance-level*: detect that this listing photo is a re-post of that one, among 80M items in the same handful of categories. Instance identity is precisely the within-class variation the objective was told to discard. So the symptom in production is a similarity distribution where every pair of items in a category scores high, duplicates and non-duplicates overlap heavily, and no threshold separates them cleanly. You are also paying for 2048 stored coordinates per item to carry 30 directions of information. ## What to do about it **Do not widen the layer.** Going to 4096 doubles storage for the same 30 directions; width was never the binding constraint. **Change the training signal.** In rough order of cost: introduce finer supervision — attribute labels, a deeper taxonomy, or instance-level identity where you have it, which directly asks the features to keep within-category differences. Or train the encoder with a self-supervised objective that treats two augmented views of the same photo as a positive pair, which makes instance identity the training target and typically produces a much fuller spectrum, provided it includes negatives or an explicit variance/decorrelation regulariser. **Consider an earlier layer as a stopgap.** Layers further from the head are less collapsed by the projection and often retain more variance, at the cost of less semantic organisation. **Then, and only then, revisit width.** Once effective rank is measured, shrinking the stored vector toward it is nearly free in quality and a large win in storage — but understand that this is a cost optimisation, not a quality fix. Reducing 2048 collapsed dimensions to 64 does not add information; it merely stops paying for emptiness. ## The honest limit of the diagnostic A healthy spectrum is necessary, not sufficient. High effective rank only says variance is spread across many directions; those directions may encode lighting, compression artefacts, watermarks or background rather than anything about the object. Always pair the spectrum with a task metric — duplicate precision and recall at your operating threshold — and treat effective rank as the cheap early warning, not the verdict.
- Would widening the penultimate layer from 2048 to 4096 fix the collapse?No. Effective rank is set by the objective and the diversity of the data, not by the declared width, so a wider layer yields the same ~30 participating directions at twice the storage and comparison cost. Fix the training signal first — finer or instance-level supervision, or a self-supervised objective with an anti-collapse term — and only then revisit width, with the measured effective rank in hand.
- Why does mean-centring matter before you read the singular values?Without centring, a large common offset shared by every embedding appears as one dominant singular value that reflects the mean, not any structure in the data. That inflates the apparent leading direction and can hide how flat the rest of the spectrum is. Centre over a large fixed sample, and keep sample size and centring identical between runs so the numbers are comparable.
- Can the spectrum look healthy while the embedding is still bad for your task?Yes. Effective rank only says variance is spread across many directions; it says nothing about what those directions encode. They may capture lighting, background or compression artefacts rather than object identity. Treat the spectrum as a cheap early warning and always confirm with a task metric — duplicate precision and recall at the operating threshold you actually ship.
You rented a 2048-shelf warehouse and the goods only ever occupy 30 shelves. Buying a bigger warehouse does not create more goods; changing what you order does.
saying these in an interview costs you the question
- Believing declared dimension equals information content
- Proposing a wider layer as the fix for collapse
- Reading singular values without mean-centring first
- Calling low effective rank bad regardless of the downstream task
- Treating a healthy spectrum as proof the embedding is good