Why do a recommender's 500-million-parameter embedding tables barely affect its per-request FLOPs?
answer
- Two metrics, two different constraints
- Reuse per request is the bridge
- A lookup is not a multiplication
- Convolution kernels invert the relationship
- Size constrains storage, FLOPs constrains compute
basics
~20 sAn embedding table is read, not multiplied. A request looks up a handful of rows, so the table dominates model size and memory while contributing almost no arithmetic; the small dense layers on top do nearly all the compute.
solid answer
~50 sParameter count answers "how much space does this model occupy"; FLOPs answers "how much arithmetic does one request cost". An embedding table decouples them completely: it holds hundreds of millions of weights, but serving one request retrieves a few rows by index, which is memory traffic rather than multiplication. The dense layers stacked on those embeddings might hold a few million parameters and still account for essentially all the FLOPs, because every one of their weights participates in a multiply-accumulate for every request. The reverse case exists too: a small convolution kernel reused across a high-resolution feature map is tiny in parameters and enormous in FLOPs. So a parameter count constrains your storage, download size and weight memory; it tells you very little about per-request compute, and the two constraints are optimized by different changes.
go deeper
Be able to say that parameter count measures how much the model stores and FLOPs measures how much arithmetic a request performs, and that a lookup table adds a lot of the first and almost none of the second.
Explain reuse as the bridge: FLOPs is roughly parameters used times uses per request. Work the convolution case out loud — kernel weights are tiny and reused at every spatial position.
Demonstrate that you diagnose before you optimize: identify which resource is actually binding, then pick the lever that moves it, rather than shrinking whichever number is easiest to shrink.
Own how efficiency targets are written for your teams. Insist that a target name its resource — bundle size, weight memory, per-request compute — so that engineering effort lands where the constraint is.
## Two different questions Parameter count and FLOPs answer different deployment questions, and confusing them is one of the most common efficiency mistakes. - **Parameter count** is how many learned values the model stores. Multiply by bytes per value to get **model size**: 500 million parameters in 32-bit floats is 2 GB; in 8-bit form, 500 MB. This is the number that governs download size, on-disk footprint, and how much memory the weights occupy while the model is resident. - **FLOPs** is how much arithmetic one forward pass performs. This is the number that scales with request cost. The link between them is **reuse**: FLOPs is roughly "parameters used, times how many times each is used per request". Whenever that reuse factor swings wildly between layers, the two metrics come apart. ## The embedding-table case: many parameters, no arithmetic A large recommender keeps an embedding table with one learned vector per user, item, or categorical feature value. With tens of millions of items and a modest vector width, the tables can easily reach hundreds of millions of parameters and dominate everything else in the model. Serving one request touches almost none of it. The model takes the indices present in that request — this user, these few dozen candidate items, a handful of categorical fields — and reads the corresponding rows. Mathematically it is a multiplication by a one-hot vector, but it is implemented as a lookup, and the arithmetic performed is essentially zero. Each retrieved parameter is used **once**, and the vast majority are not touched at all. What is left is a stack of dense layers over the concatenated vectors, perhaps a few million parameters. Every one of those weights participates in a multiply-accumulate on every request. That small tail carries nearly all the FLOPs. The operational consequence: this model is **not** compute-constrained, it is **memory- and capacity-constrained**. Shrinking the dense layers saves you nothing you care about. What matters is table size — hashing the vocabulary, narrowing the vectors, pruning cold rows, storing the table in lower precision, or keeping it in a separate store. And because a lookup is a scattered read, the real serving cost of the tables is memory: how they are held, and how quickly random rows can be fetched. ## The mirror case: few parameters, enormous arithmetic A convolution inverts the relationship. A 3x3 kernel with 64 input and 64 output channels holds `3 * 3 * 64 * 64 = 36,864` weights — negligible storage. Applied across a 224x224 output map, it performs `36,864 * 224 * 224` multiply-accumulates, about 1.85 billion. Each weight is reused once per spatial position. A convolutional feature extractor can be a few megabytes on disk and still be the most expensive thing in the request. This is why input resolution is a compute lever that leaves parameter count untouched: halving each spatial dimension of the input cuts convolutional FLOPs by roughly four, and changes the model size not at all. ## Choosing which number to quote Start from the constraint someone actually wrote down: - "The app bundle must stay under 100 MB", "the model must fit in the device's weight memory", "we cannot push a 2 GB update to fleet devices" → that is a **parameter-count-times-precision** problem. - "Each request must cost under X" or "we must serve N requests per accelerator" → that is a **compute** problem, and FLOPs is only the paper proxy for it. Quoting the wrong one produces work that does not move the constraint: teams that shave FLOPs off a model whose real problem is a 2 GB embedding table, or that quantize weights to shrink a model whose real problem is milliseconds per request. ## What a strong answer sounds like State the definitions apart, give the reuse factor as the bridge between them, name one case in each direction — a huge table with a per-request lookup, a small kernel applied everywhere — and then say which deployment constraint each metric answers. The interviewer is checking whether you optimize the metric your constraint names, or the metric you happen to know how to reduce.
- Give the opposite case — a layer with few parameters and a large FLOP cost.A convolution kernel. A 3x3 kernel over 64 input and 64 output channels holds under 37 thousand weights, but applied across a 224x224 output map it performs close to two billion multiply-accumulates, because every weight is reused at every spatial position. Model size stays trivial while compute dominates the request, which is why lowering input resolution is a compute lever that leaves parameter count unchanged.
- If the embedding tables are the problem, which levers actually help?The ones that shrink the table: narrowing the vector width, hashing or trimming the vocabulary, dropping cold rows that see almost no traffic, and storing the table at lower precision. Shrinking the dense layers on top saves FLOPs the system was not short of. The right first step is confirming which resource is binding — weight memory or per-request compute — before choosing.
A dictionary has enormous bulk but you consult two entries per sentence; a short multiplication table is tiny yet you use every cell constantly. Bulk and use are separate costs.
saying these in an interview costs you the question
- Treats parameter count as a proxy for inference cost
- Calls a bigger model automatically a slower model
- Forgets that convolution weights are reused across positions
- Quotes model size when the constraint is per-request latency
- Assumes an embedding lookup performs a full matrix multiply