How much memory do 50 million 1536-dimension float32 embeddings require?
answer
- vectors times dimensions times bytes
- float32 is four bytes each
- raw figure is a floor, not a bill
- index, metadata and replicas ride on top
- the binding constraint is RAM
basics
~10 sRoughly 300 GB: 50,000,000 x 1536 x 4 bytes is about 307 GB of raw vector data, before any index structure, metadata, or replicas. Dimension count multiplies every byte you store, ship, and compare.
solid answer
~50 sThe arithmetic is `vectors x dimensions x bytes-per-value`: 50M x 1536 x 4 bytes for float32 is about 307 GB (roughly 286 GiB) of **raw vectors alone**. That number is the floor, not the bill. On top of it you pay for the index structure that makes search fast, for stored document ids and filter metadata, and for however many replicas you run — so planning for two to three times the raw figure is normal. Because that working set usually has to sit in RAM to hit low-latency search, dimension is really a *memory-tier* decision, not a disk one. Dimension also scales the per-comparison work — a similarity computation is one pass over d values — plus index build time and the payload size of every embedding you push over the network. Halving d roughly halves all of those at once.
go deeper
Be able to do the multiplication on the spot: vectors times dimensions times four bytes for float32, and say the result in gigabytes without hesitating.
Explain that the raw figure is a floor and name what rides on top — index structure, metadata, replicas — and that per-comparison compute scales linearly with dimension too.
Turn the number into a capacity plan: which tier the working set must live in, what a rebuild costs, and which lever you would pull first when the bill is too high.
Frame dimension as one point on a cost-quality frontier alongside vector count, chunking policy, and index choice, and insist the cliff point be established empirically before anyone commits a budget.
## The arithmetic An embedding is a fixed-length array of numbers. Its storage size is entirely mechanical: ``` bytes = number_of_vectors x dimensions x bytes_per_value ``` With 50 million documents, 1536 dimensions, and 32-bit floats (4 bytes each), that is 50,000,000 x 1536 x 4 = 307,200,000,000 bytes — about 307 GB in decimal units, or roughly 286 GiB. A single 1536-dimension float32 vector is 6,144 bytes; a 768-dimension one is 3,072 bytes. There is nothing subtle here, and interviewers ask it precisely because a candidate who cannot do this arithmetic will size a system wrong by an order of magnitude. ## Why the raw number is only the floor Three things sit on top of the raw vectors: **Index overhead.** A brute-force scan over 307 GB is not going to answer in tens of milliseconds, so you build an approximate-nearest-neighbour structure over the vectors. Every such structure adds its own bytes per vector — graph links, partition assignments, or centroid tables depending on the family. Treat this as a real multiplier, sized from the specific index you pick. **Payload and metadata.** Document ids, tenant ids, timestamps, and any attributes you filter on all live alongside the vectors, and filterable metadata often has to be resident too. **Replication and headroom.** One replica for availability doubles the figure. Rebuilds, compactions, and background merges want spare room on top. So a 307 GB raw corpus is realistically a 600 GB to 1 TB memory-tier commitment. At typical cloud memory prices that is a five-figure monthly line item, which is why dimension shows up in budget conversations and not just in model-selection ones. ## Where else the cost lands Engineers tend to think of dimension as a storage question. It touches four budgets at once: 1. **RAM.** Latency-sensitive vector search wants the searchable structure resident. This is usually the binding constraint, and RAM is the expensive tier. 2. **CPU per comparison.** Computing a dot product or Euclidean distance is one pass over d values. Doubling d roughly doubles the work of every candidate comparison the search performs. With an approximate index you do far fewer comparisons than a full scan, but each one still scales linearly in d. 3. **Index build and rebuild time.** Every distance computation during construction carries the same d factor, so a re-embedding or re-index job takes proportionally longer. 4. **Network and serialization.** Embeddings crossing a service boundary as JSON are dramatically worse than the binary figure — a float serialized as text can run 10 or more bytes per value. Batch ingestion pipelines and cross-region replication feel this immediately. ## Dimension as a dial Because all four scale together, dimension is the cheapest single lever you have on infrastructure cost, and unlike most levers it is continuous. Going from 1536 to 768 dimensions cuts the raw footprint of that 50M corpus from ~307 GB to ~154 GB and roughly halves comparison work. The question is always what quality you give up, and the honest answer is: measure it. Retrieval quality against dimension is usually **flat then cliff** — quality barely moves as you shrink, then falls off sharply past some corpus-specific point. Nobody can tell you where your cliff is from first principles; you sweep dimensions and plot recall@10 on your own labeled queries. ## Practical routes to a smaller footprint - **Pick a lower-dimension model.** Often a smaller modern encoder matches a larger older one at a fraction of the width. - **Shorten the vectors you already produce.** Models trained with nested representations let you keep a prefix of the dimensions; a general encoder can be projected down with a fitted linear reduction. Both need validation on your data. - **Two-pass retrieval.** Shortlist on short vectors, rescore the top candidates on full-width vectors. You pay full width only for a few hundred documents per query. - **Store fewer vectors.** Chunking strategy multiplies vector count as surely as dimension multiplies vector size; 50M vectors from 5M documents is a chunking decision, and it is often the larger lever. There are also compression schemes applied at the index layer that reduce bytes per value rather than the number of values — a separate axis from the dimension question, and one you evaluate on its own terms. ## What a good answer sounds like Do the multiplication out loud, state the ~300 GB figure, then immediately say it is the floor and name index overhead and replicas. Finish by pointing at RAM as the binding constraint. That progression — arithmetic, overhead, tier — is what the question is testing.
- Why is the JSON payload of the same embedding so much larger than the binary figure?Because each float is serialized as decimal text plus a separator — commonly 10-20 bytes per value instead of 4. A 1536-dimension vector that is 6 KB in binary can exceed 20 KB as JSON. On high-volume ingestion or cross-service hops that inflation dominates bandwidth and parsing cost, which is why bulk embedding transfer usually moves to a binary or compact encoding.
- If 50 million vectors is too many, is reducing dimension always the right lever to reach first?Not necessarily. Vector count and vector width multiply, and count is often driven by chunking policy. If aggressive chunking turned 5 million documents into 50 million chunks, revisiting chunk size or de-duplicating near-identical chunks can cut the footprint more than halving dimension would, and without touching retrieval quality per chunk.
- Does halving the dimension halve query latency in an indexed search?It reduces it but rarely halves it. Distance computation scales linearly in d, so each comparison gets about twice as cheap, but query time also includes index traversal, candidate bookkeeping, filtering, and network overhead that do not scale with d. Expect a meaningful improvement in the compute-bound portion, and measure rather than assume the full factor.
saying these in an interview costs you the question
- Quoting only raw vector bytes and ignoring index and replica overhead
- Assuming a float takes one byte per dimension
- Treating the corpus as a disk-sizing problem when RAM is the constraint
- Claiming dimension affects storage but not query compute
- Assuming halving dimension always halves retrieval quality too