How do you choose 64, 512 or 2048 as the embedding width for 80M stored listings?
answer
- capacity ceiling, not a quality guarantee
- cost is linear in width at this scale
- measure the curve, take the knee
- check effective rank before comparing widths
- changing width later means a full re-embed
basics
~20 sWidth sets a capacity ceiling, not quality, while storage and per-comparison cost grow linearly with it. Train at the widest width you can afford, plot the downstream metric against width, and ship the knee of that curve rather than the widest option.
solid answer
~50 sThree budgets are coupled. Capacity: width is an upper bound on how much the representation can encode, and it is only a ceiling — if effective rank sits far below it, extra coordinates buy nothing. Serving cost: at 80M items, storage and per-comparison work scale linearly in width, so 2048 versus 64 is a 32x factor on both. Downstream fragility: wider vectors need more data or regularisation for whatever consumes them. The method is empirical — train wide, then measure the metric that pays the bills (duplicate precision and recall at the shipping threshold) as a function of width, and take the knee plus a margin. Then weigh the fact that width is expensive to change later: re-embedding 80M items is a backfill, a dual-read window and a threshold recalibration, so bias slightly wide if the roadmap adds consumers.
go deeper
Know that a wider embedding costs proportionally more to store and compare, and that more dimensions does not automatically mean a better representation.
Explain how to get the evidence: fix the encoder, produce candidate widths, and score each against the downstream metric at the operating threshold rather than against a training loss.
Show that you would check effective rank before comparing widths at all, and that you would price the migration — backfill, dual-read window, threshold recalibration — as part of the decision.
Own the irreversibility argument and the multi-consumer contract: who sets the width, what is published about it, and what investment makes a future re-embed a routine batch job rather than a blocked project.
## Why this is a judgment call and not a lookup Embedding width is one of the few architecture decisions that leaks straight into infrastructure, product quality and organisational commitments at once. It is worth structuring the answer around the three budgets it spans. ### Budget one: representational capacity Width is a **ceiling**. It bounds how much a vector *can* carry, and guarantees nothing about how much it *does*. A 2048-wide layer whose effective rank is 30 has the quality of a 30-dimensional representation and the cost of a 2048-dimensional one. So the first question is never *how wide* but *how much distinct information does the task require*: a 32-way categorisation needs a handful of directions; instance-level near-duplicate detection over 80M listings needs enough to separate individual objects within a category, which is far more. A reasonable prior: coarse classification tasks saturate early, retrieval and near-duplicate tasks keep improving for longer, and multi-consumer embeddings need more than any single consumer would. ### Budget two: serving cost At 80M stored items, everything downstream is linear in width. Storage per item grows in proportion. The work in a single similarity comparison grows in proportion. Memory pressure on whatever holds the vectors grows in proportion. Rebuild and backfill times grow in proportion. Going from 512 to 2048 does not add a rounding error to the bill — it multiplies that whole column by four. This is also where you decide the numeric precision you store, which trades against width for the same budget. The point for the interview is to show you know the cost is a *product* of item count, width and per-value size, and that at this item count the width term is a first-class infrastructure decision, not a hyperparameter. ### Budget three: downstream fragility Everything that consumes the vector inherits its width. A downstream model fitted on 2048 inputs has more parameters and needs more labelled data or stronger regularisation than one fitted on 64. If your labelled downstream set is small, a narrower embedding is often *better*, not merely cheaper — the classic bias-variance trade appears here as an architecture choice. ## The method Do not argue the three budgets in the abstract; measure. 1. **Train once at the widest width you can afford.** This is the reference and it bounds what any narrower option could achieve. 2. **Measure effective rank.** If it is far below the declared width, stop — the binding constraint is the objective, and comparing widths is measuring the wrong thing. 3. **Get the quality-versus-width curve.** The cheap version is a linear projection of the wide vector down to each candidate width, fitted on a sample; the expensive and more faithful version is retraining the encoder at each width, which lets the network allocate its capacity rather than merely discarding directions. Run the cheap version first to find the region, retrain at two or three widths around it. 4. **Score with the metric that pays the bills**, at the operating threshold you would actually ship — duplicate precision and recall for dedup, not a proxy loss. 5. **Take the knee, plus margin.** These curves are typically steep then flat; the interesting choice is where flat begins. ## The part that makes it a lead's decision Quality curves do not settle it, because width is **hard to change later**. The stored vectors, every consumer's model, every calibrated threshold and every cached result are all downstream of the choice. Changing it means a full re-embed of 80M items, a dual-write or dual-read window while old and new coexist, and a recalibration of every threshold — similarity thresholds do not transfer between encoders, so nothing tuned against the old vectors survives the swap. That asymmetry shapes the decision: - **Bias wider** when the roadmap adds consumers, when the task will get finer over time, or when the team's capacity for large migrations is thin. Paying 4x storage is often cheaper than a migration you cannot schedule. - **Bias narrower** when serving cost dominates the unit economics, when the task is stable and well-understood, or when downstream consumers have little labelled data. - **Reduce reversibility cost** as an explicit investment: pin and version preprocessing, keep the extraction job reproducible end to end, and store encoder provenance with every vector. A re-embed you can run as a routine batch job makes width a decision you can revisit; one that requires archaeology makes it permanent. ## What a weak answer looks like Naming a number. There is no correct width in the abstract — the same encoder that is over-provisioned for 32-way categorisation is under-provisioned for instance-level dedup, and the deciding evidence is a curve you measured plus a migration cost you priced.
- The 2048-wide vector beats 512 by one percent on duplicate recall. Which ships?Convert the percent into absolute outcomes at your volume before deciding. One point of recall on a rare positive class may be a handful of missed duplicates a day, against a fourfold increase in storage and comparison cost across 80M items. If the misses are cheap to catch by another route, 512 wins; if each missed duplicate is a fraud or trust incident, the wide vector is cheap. State the trade in incidents and currency, not in percentage points.
- Can you just project the 2048 vector down to 512 instead of retraining a narrower encoder?Yes, and it is the right first experiment because it is fast and bounds the answer. The difference is that a post-hoc projection can only discard directions, whereas retraining at 512 lets the network allocate its capacity to what matters and often lands slightly better. Use the projection to find the region of interest, then retrain at two or three candidate widths to make the final call.
- Who should own the width decision when several teams consume the embedding?The team that owns the encoder should own the number, but only under an explicit contract: published width, precision, preprocessing version and encoder provenance, plus a stated deprecation and re-embed process. Consumers get a say through their measured quality-versus-width curves, not through preference. Without that contract, width gets set by whoever shipped first and becomes unchangeable by default.
saying these in an interview costs you the question
- Naming a default width without reference to the task
- Assuming a wider embedding is always a better embedding
- Ignoring that cost scales linearly with 80M stored items
- Forgetting that thresholds must be recalibrated after a width change
- Comparing widths before checking whether the current width is even used