skip to content

How would you validate cutting a legacy 4096-dimension encoder down to 128 dimensions with PCA?

level: seniorimportance: should knowfreq 34%

answer

  1. variance kept is not relevance kept
  2. one fitted transform for docs and queries
  3. recall@10 on labeled queries decides it
  4. sweep widths, look for the cliff
  5. score thresholds do not transfer

basics

~20 s

Judge it by retrieval quality, not by explained variance. Fit PCA on a corpus sample, project queries with the same transform, then sweep widths and compare recall@10 and MRR on labeled queries against the full-dimension ranking. Ship the smallest width before the cliff.

solid answer

~50 s

Explained variance is the wrong acceptance criterion: it measures how much of the vectors' *spread* survives, and retrieval depends on the directions that separate relevant from irrelevant documents, which are not necessarily the highest-variance ones. The correct protocol is: fit PCA on a representative sample of document vectors (centred, and using the mean from that fit), project **queries through the identical transform**, re-normalise afterwards if your pipeline assumes unit vectors, then evaluate recall@10, MRR, and nDCG at 64, 128, 256, and 512 dimensions on a labeled query set. Use the full-4096 ranking as a reference baseline but also score against human labels, so you measure real quality rather than agreement with the old system. Expect flat-then-cliff and pick the smallest width still on the flat part with margin. Report the hard-query slice separately, and re-tune any score thresholds, which do not transfer across widths.

code

python · 15 lines
python
import numpy as np
from sklearn.decomposition import PCA

docs = np.random.randn(50_000, 4096).astype(np.float32)

pca = PCA(n_components=128, random_state=0).fit(docs)  # fit once, on documents

def project(vectors: np.ndarray) -> np.ndarray:
    out = pca.transform(vectors)                        # same transform everywhere
    return out / np.linalg.norm(out, axis=1, keepdims=True)

doc_index = project(docs)
query = project(np.random.randn(1, 4096).astype(np.float32))
scores = doc_index @ query.T
print(doc_index.shape, float(pca.explained_variance_ratio_.sum()))

go deeper

for a junior

Know that PCA projects vectors onto fewer axes, and that the same fitted transform has to be applied to queries and documents alike.

for a middle

Explain why explained variance and reconstruction error do not predict ranking quality, and name the mechanical steps: fit on a sample, store the mean, project queries identically, re-normalise.

for a senior

Design the evaluation: labeled query set with a hard slice, a width sweep, recall@10 and nDCG against labels plus top-k agreement with the baseline, cost metrics per width, and re-tuned thresholds before rollout.

for a principal

Weigh the projection against simply migrating to a narrower modern encoder, and set the policy that a fitted transform is versioned data whose refit is a governed re-index, not a config tweak.

## The setup and why it comes up You inherit a search system built on a 4096-dimension encoder. Re-embedding the corpus with a modern, narrower model is the clean answer, but it may be blocked — a fine-tuned in-house encoder, a frozen contract with downstream consumers, or simply a corpus too large to re-encode this quarter. Cutting the stored vectors to 128 dimensions with a fitted linear projection is the pragmatic move: it shrinks storage and per-comparison work by 32x without touching the model. Unlike a nested-representation model, an ordinary encoder has no privileged prefix, so you cannot slice. You must fit a transform. ## Fitting the projection correctly PCA finds the orthogonal directions of greatest variance in a sample of vectors and projects onto the top k of them. The mechanics that people get wrong: **Fit on a representative sample.** A few hundred thousand document vectors is usually plenty; the sample must match the live distribution across languages, document types, and topics, or the retained directions will be tuned to a subpopulation. **Centring is part of the transform.** PCA subtracts the mean before projecting. That mean is fitted, must be stored, and must be applied identically at query time. **Queries go through the same transform.** This is the single most common bug. Fitting one PCA on documents and another on queries — or projecting documents and leaving queries at full width — produces coordinate systems that are not comparable and silently destroys ranking. One fitted transform, applied everywhere. **Re-normalise after projecting.** Projection does not preserve norms. If downstream code assumes unit vectors so that dot product equals cosine, L2-normalise the projected vectors. **Version the transform with the index.** The projection is now part of your data contract. Refitting it invalidates every stored vector, so treat a refit as a full re-index and give it a version identifier. ## Why explained variance misleads The temptation is to keep enough components to reach, say, 95% explained variance and declare success. Two problems: 1. **Variance is not relevance.** The highest-variance directions capture what varies most across the corpus — often topic or document-type, sometimes language or formatting artefacts. What retrieval needs is whatever separates a relevant document from a plausible-but-wrong one, which can live in modest-variance directions. Discarding a low-variance direction that happens to encode negation or a distinguishing entity costs real recall while barely moving the variance number. 2. **Reconstruction is not ranking.** Low reconstruction error means the projected vectors are close to the originals on average, but ranking is decided by *order*, and small perturbations near a decision boundary flip order. A 98%-variance projection can still reshuffle the top of the result list. So explained variance is a useful diagnostic for choosing candidate widths to test, and useless as an acceptance criterion. ## The evaluation protocol 1. **Build a labeled query set** from real traffic — a few hundred queries minimum, each with judged relevant documents. Stratify it: head queries, tail queries, and a deliberately hard slice with near-duplicate candidates. 2. **Establish the full-dimension baseline.** Score recall@10, MRR, and nDCG@10 at 4096 dimensions. This is the reference the business is currently getting. 3. **Sweep widths.** Fit projections at 64, 128, 256, 512 and score each on the same query set with the same index configuration. 4. **Report two things per width.** Absolute quality against the labels, and agreement with the 4096 ranking (for example, overlap of the top ten). Agreement alone is not enough — the old ranking is not ground truth — but a sharp drop in agreement flags behaviour change users will notice. 5. **Read the curve.** You will typically see quality nearly flat down to some width, then a cliff. Choose the smallest width still on the flat part *with margin*, because the cliff moves as the corpus drifts. 6. **Look at the hard slice separately.** Aggregate means hide targeted damage; reduction hurts fine discrimination first, and that is exactly the hard slice. 7. **Re-tune thresholds.** Any similarity cutoff, minimum-score filter, or router threshold was calibrated on 4096-dimension score distributions and does not transfer. Also record what you came for: bytes per vector, index size, p95 latency, and index build time at each width. If 256 dimensions costs 0.5 points of recall against 128 and the memory saving between them is immaterial at your scale, take the 256. ## Alternatives worth putting on the table **Whitening** — rescaling components to equal variance — sometimes helps and sometimes hurts, because it amplifies low-variance directions that may be noise. Treat it as another arm of the sweep, never a default. **Non-linear reduction such as UMAP** is the wrong tool here. It is optimised to preserve neighbourhood structure for inspection, its out-of-sample transform is approximate and comparatively expensive per query, and it gives no clean guarantee that metric relationships needed for ranking survive. Linear projections are cheap, deterministic, and applicable to a query in microseconds. **Just re-embed.** Run the numbers on a modern narrower encoder before committing to the projection. If a 512-dimension model beats the PCA-128 pipeline on your labeled set, the migration cost may be the cheaper long-term answer, and it removes the fitted transform from your data contract entirely. ## The two-pass escape hatch If the reduced index loses top-of-list precision but keeps recall, you do not have to accept the reduced ranking. Retrieve a few hundred candidates on the 128-dimension index and rescore them with the retained 4096-dimension vectors. Full-width cost is then paid per query on a few hundred documents rather than across the corpus.

  • Why is agreement with the original 4096-dimension ranking not a sufficient acceptance test?
    Because the old ranking is a baseline, not ground truth — it has its own errors, and a reduction that diverges from it may be diverging toward better results. Score both systems against human labels so you measure quality directly. Keep top-k overlap as a secondary signal, since a large drop predicts visible behaviour change for users even when label-based metrics hold steady.
  • What breaks if you refit the PCA transform after the index is already built?
    Everything already stored becomes incomparable, because the new transform defines different axes and a different mean. Queries projected with the new fit will not align with documents projected under the old one, and ranking degrades in ways that look like a quality regression rather than a config error. Treat the transform as versioned data: a refit is a full re-index, gated the same way.
  • When would you skip PCA entirely and re-embed the corpus instead?
    When a modern narrower encoder beats the projected pipeline on your labeled set, and the re-embedding cost is affordable. Re-embedding removes a fitted transform from the data contract, usually improves quality outright rather than trading it away, and avoids the ongoing burden of keeping the projection versioned and consistent across every producer of embeddings.

saying these in an interview costs you the question

  • Accepting a reduction because it retains 95% of explained variance
  • Fitting separate PCA transforms for documents and for queries
  • Comparing reduced document vectors against full-width query vectors
  • Reporting only mean recall and never the hard-query slice
  • Reusing similarity thresholds calibrated at the original dimension

context