skip to content

Across eighty million players, what does one 512-dimension vector of 32-bit floats per player cost to store, and what does halving its width save?

level: middleimportance: should knowfreq 40%

answer

  1. width times bytes times population
  2. 512 coordinates is 2,048 bytes
  3. payload is a floor, not a budget
  4. the same bytes cross the wire per request
  5. any width change is a new encoder version

basics

~20 s

512 coordinates at four bytes each is 2,048 bytes per player, so eighty million players hold roughly 164 GB of vector payload before keys, stamps or extra copies. Halving the width to 256 coordinates halves that to about 82 GB.

solid answer

~40 s

The arithmetic is `width x bytes-per-scalar x players`: 512 x 4 x 80,000,000 is about 164 GB, and that is payload only - entity keys, the encoder-version stamp, timestamps and whatever copies the online tier keeps all sit on top, and a two-namespace cutover doubles it for the duration. It also shows up on the read side: 2,048 bytes leave the online store on every scoring request, so 20,000 scoring requests per second is roughly 41 MB/s of vector payload. Halving the width, or halving the bytes per coordinate, halves both numbers. Neither is free: a narrower vector carries less information, and any change to what the stored vector *is* makes a new encoder version that every consumer must retrain against.

code

pseudocode · 10 lines
pseudocode
bytesPerVector  = width * bytesPerScalar        # 512 * 4 = 2048
storeBytes      = players * bytesPerVector      # 80_000_000 * 2048 = 163_840_000_000
readBytesPerSec = scoringQps * bytesPerVector   # 20_000 * 2048 = 40_960_000

# same population, half the width, same precision
bytesPerVector  = 256 * 4                       # 1024
storeBytes      = 80_000_000 * 1024             # 81_920_000_000

# during a two-namespace cutover both copies are live
peakStoreBytes  = 163_840_000_000 * 2           # 327_680_000_000

go deeper

for a junior

Be able to do the multiplication out loud: coordinates times bytes per coordinate times the number of entities. Most sizing questions start exactly there.

for a middle

Add what the payload figure excludes - row overhead, the copies the tier keeps, the overlap during a cutover - and give the read-bandwidth number alongside the storage one.

for a senior

Show that shrinking a vector is not a storage change: it makes a new encoder version, so it costs an encode pass and a retrain for every consuming model.

for a principal

Decide whether the saving is worth the coordination. At this scale, the argument that usually wins is per-request bytes, not the monthly storage bill.

## The arithmetic One stored embedding costs `width x bytes-per-scalar` bytes, and the store costs that times the population. At eighty million players: | width | scalar | bytes per vector | eighty million players | |---|---|---|---| | 512 | 32-bit float | 2,048 | ~164 GB | | 512 | 16-bit float | 1,024 | ~82 GB | | 256 | 32-bit float | 1,024 | ~82 GB | | 128 | 16-bit float | 256 | ~20.5 GB | These are the numbers to quote in a design conversation, and they are worth doing out loud rather than asserting: 512 x 4 = 2,048; 2,048 x 80,000,000 = 163,840,000,000 bytes. ## What the number leaves out The payload figure is a floor, not a budget. On top of it: - **Row overhead** - the entity key, the `encoderVersion` stamp, the `computedAt` timestamp and the store's own per-row bookkeeping, typically tens of bytes per row, which is small against 2,048 but not against 256. - **Copies** - whatever the online tier keeps for durability and availability multiplies the whole figure. - **The offline history** - the columnar store keeps rows for training, and if it keeps one row per player per encoder version it grows with every adoption rather than staying flat. - **The cutover overlap** - two live namespaces means two full copies, 328 GB at 512 coordinates, until the superseded one is deleted. ## The read side The same 2,048 bytes cross the wire on every scoring request that fetches a player's vector. At 20,000 scoring requests per second that is 20,000 x 2,048 = 40,960,000 bytes per second, roughly **41 MB/s** of vector payload out of the online store - before any other feature in the request. Halving the width halves that too, which is often the argument that actually wins, because bandwidth and deserialisation are felt on every request while storage is a monthly line item. ## Three ways to make it smaller, and what each costs 1. **Fewer coordinates.** Train the encoder narrower, or fit a projection from 512 down. A random projection has a distance-preservation guarantee - the **Johnson-Lindenstrauss lemma** bounds how much pairwise distances can change for a given target width - whereas simply keeping the first 256 coordinates of a learned encoder has no such guarantee: nothing orders a learned encoder's coordinates by importance unless it was trained to produce that ordering. 2. **Fewer bits per coordinate.** Halving the scalar size halves the bytes and perturbs each value slightly. Note the asymmetry: a *width* change is caught loudly by a shape check, while a *precision* change is not, so it needs the same version discipline as any other silent change. 3. **Quantise.** **Product quantization** stores short codebook indices in place of raw floats and cuts bytes far more aggressively, at the price of a reconstruction error and an extra decode step on read. ## Why every one of them is a new encoder version All three change what the stored vector *is*. That makes a new encoder version, which means a new namespace, a re-encode of the population, and a retrain of every consuming model - so "let's make the vectors smaller" is never a storage ticket. Its real cost is one encode pass plus one retrain per consumer, and the storage saving has to be worth that before anyone starts. The honest framing in a design round is comparative: at eighty million players, 164 GB is not a large number for a columnar offline store and is a noticeable one for a low-latency key-value tier priced on memory. Say which tier you are sizing before you argue about whether the number is big.

  • Is 164 GB a lot?
    It depends which tier you are sizing, which is why the question is worth asking back. In a columnar offline store it is unremarkable. In a low-latency key-value tier priced on memory, and multiplied by the copies that tier keeps and by a second namespace during a cutover, it becomes a real line item and the first place a width reduction pays for itself.
  • Can the team just keep the first 256 coordinates of each stored vector?
    Mechanically yes, and it does halve the bytes, but a learned encoder's coordinates carry no importance ordering unless it was trained to give one, so slicing discards an arbitrary half of the signal. A fitted projection down to 256 keeps far more of the geometry for the same stored size, and either way the consuming models have to be retrained.

saying these in an interview costs you the question

  • Quotes vector payload as the whole storage budget, ignoring keys and copies
  • Assumes the leading coordinates of a learned encoder carry the most signal
  • Treats a precision change as invisible because the width did not move
  • Forgets that both namespaces are live during a cutover
  • Counts storage only and ignores the bytes read on every scoring request