skip to content

How do you choose the embedding width for a 2-million-item catalog id that dominates the model?

level: principalimportance: should knowfreq 44%

answer

  1. rows times width is the parameter count
  2. one layer can outweigh the whole network
  3. data per id, not number of ids
  4. budget the memory before sweeping

basics

~10 s

Size it from observations per id and a stated memory budget, not from cardinality. Two million rows at width 64 is 128 million parameters, so width is the model's main size decision.

solid answer

~50 s

A table's parameter count is rows times width, so 2 million items at width 64 is 128 million parameters -- next to a 3-million-parameter network, the table effectively is the model, and optimizer state multiplies its training memory several times over. Width is therefore a budget decision, not a hyperparameter to sweep blindly. The input that matters is observations per id, not cardinality: a wide row is only supportable if enough examples carry that id, so I read the median count per id and the traffic share of the head. Rules of thumb like the fourth root of cardinality are a first point, not an answer -- they disagree with each other by an order of magnitude. I set the serving memory budget first, sweep width with everything else fixed, and take the smallest width the validation metric supports.

go deeper

for a junior

Recall that an embedding table's parameter count is the number of ids times the width, and that this single layer can hold far more parameters than every dense layer combined. Be able to do that multiplication on the spot.

for a middle

Explain where the memory actually goes -- the table itself, the optimizer's per-parameter state, the checkpoint, the serving process -- and why a lookup dominates memory without adding meaningful compute.

for a senior

Show the diagnosis: read the distribution of examples per id, sweep width against the validation metric, and recognise the too-wide signature of falling training loss with a flat validation metric on the tail.

for a principal

This is your tier. Set the memory budget before the sweep, state what metric gain would justify raising it, and say what you sacrifice first -- tail rows, per-id resolution, then head width -- when the budget binds.

## The arithmetic first An embedding table's parameter count is simply `rows * width` -- cardinality times embedding dimension. A 2-million-item retail catalog at width 64 is `2,000,000 * 64 = 128,000,000` parameters. If the rest of the network -- the layers that combine features and produce the prediction -- holds 3 million, then **97% of the model is one lookup table**, and essentially every decision you make about model size is a decision about this one number. That imbalance is normal for recommendation and tabular deep learning, and it changes what "model size" means: - **Optimizer state multiplies it.** An optimizer that keeps one or two moving averages per parameter needs two or three times the table's memory during training, on top of the table itself. - **Checkpoints and transfer inherit it.** Every save, copy and rollback moves those hundreds of megabytes. - **Serving RAM is the hard wall.** Dense layers are cheap to hold; the table has to be resident (or sharded, or partly offloaded) for every request that touches an arbitrary id. - **Compute does not scale with it.** A lookup is a row copy. The table dominates *memory*, not FLOPs -- which is why "our model is huge but fast" is the normal state here, and why cutting width buys memory rather than latency. ## Cardinality is the wrong input to the decision The instinct is to scale width with cardinality: more ids, wider rows. But width sets how many parameters each *row* has, and a row is trained only by the examples carrying that id. The quantity that decides whether a width is supportable is **observations per id**, not the number of ids. Two million items with 5,000 interactions each can support a wide table; two million items with a median of three interactions each cannot, no matter how large the dataset looks in total. So the questions to ask before picking a number: 1. What is the **distribution** of examples per id -- median, not mean? What share of traffic do the top 1% of ids carry? 2. What is the **serving memory budget**, in numbers, before any tuning starts? 3. How often is the model retrained, and how quickly do new ids need rows? 4. Would the same capacity spent on side features (category, price band, brand, recency) do more than spending it on per-id rows for a thin tail? ## The heuristics, and why they are only a starting point Two rules of thumb circulate: width around the fourth root of the cardinality, and width around `min(50, cardinality / 2)`. For 300,000 merchants the first suggests roughly 23 and the second caps at 50, while production recommender teams routinely run 128 or more for a field like that. They disagree by an order of magnitude, which tells you what they are: a defensible first point in a sweep, not an answer. The honest method is to fix everything else and sweep width -- 16, 32, 64, 128 -- and read the validation metric against the memory cost, then take the smallest width whose metric you can defend. The signal that a table is too wide is not the parameter count; it is **training loss continuing to fall while the validation metric flattens or worsens**, with the gap concentrated on tail ids. The extra columns are being spent memorising rows that have almost no data behind them. Conversely, if halving the width costs nothing on validation, the width was never doing work. ## Small fields are not worth arguing about A day-of-week field has 7 values. At width 12 the whole table is 84 parameters; at width 64 it is 448. Seven rows already have all the freedom they need at a modest width -- the following layers are free to use them however they like -- and either choice is invisible in the model's memory. The engineering attention belongs on the 300,000-value merchant field, where width 128 is a 38.4-million-parameter commitment. A good answer says this out loud: **the width decision only matters where cardinality is large**, and spending review time on the small fields is misallocated effort. ## Ways to buy capacity where the data is, not everywhere - **Mixed widths by frequency.** Give frequent ids wide rows and tail ids narrow ones, projecting the narrow rows up to the common width. Parameters follow the data instead of being handed out uniformly. - **A shared row for the tail.** Fold ids below a frequency threshold into one bucket. You lose per-id resolution you never estimated, and you delete most of the table. - **Hashing into a fixed bucket count.** Decouples table size from cardinality entirely, at the cost of collisions between ids. - **Prune and re-fit at retrain.** Vocabularies grow monotonically if nobody prunes them; ids with no recent traffic are pure memory. ## The judgment call to own Embedding capacity is the cheapest capacity in the model to *add* -- one number in a config -- and among the most expensive to *serve*, which is exactly the shape of decision that gets made by accident. Set the memory budget before the sweep, state what metric improvement would justify raising it, and make the width a reviewed choice with a number attached to both sides. Then say what you would give up first when the budget is hit: tail rows before head width, per-id resolution before side features, and resolution on the fields whose ids the business cannot even act on individually.

  • Is there a rule of thumb you would start from?
    Two circulate: width near the fourth root of the cardinality, and width near min(50, half the cardinality). For 300,000 merchants the first suggests about 23 and the second caps at 50, while production recommenders often run 128 or more -- they disagree by an order of magnitude, which tells you they are a starting point in a sweep rather than a decision.
  • Would you give every id in a field the same width?
    Not necessarily. Mixed widths by frequency give head ids wide rows and tail ids narrow ones projected up to a common width, so parameters follow the data instead of being handed out uniformly. The tail can also be collapsed into a shared row or hashed into buckets, which deletes most of the table while losing resolution that was never estimated.
  • What tells you the table is wider than the data supports?
    Training loss keeps falling while the validation metric flattens or worsens, with the gap concentrated on rare ids -- the extra columns are memorising rows with almost no data behind them. The cheap check runs the other way: if halving the width costs nothing on validation, the width was never doing work.
  • Does a 7-value day-of-week field deserve the same analysis?
    No, and saying so is part of the answer. At width 12 that table is 84 parameters and at width 64 it is 448 -- invisible either way, and seven rows already have all the freedom they need. Width only matters where cardinality is large; spending review time on the small fields is misallocated effort.

Embedding width is shelf space per product. Cardinality tells you how many products exist; sales per product tells you which ones deserve a wide shelf. Giving every item the same frontage fills the warehouse with stock nobody buys.

saying these in an interview costs you the question

  • Chooses width from cardinality alone, ignoring data per id
  • Assumes a wider table always learns more
  • Forgets optimizer state multiplies the table's training memory
  • Counts only the dense layers when sizing the model
  • Tunes width with no serving memory budget stated
  • Argues at length about a seven-value field's width

context