How would you fit 50M 1536-dim vectors in Qdrant on a fixed RAM budget?
answer
- start with a multiplication, not a feature
- dimensions times four bytes per vector
- which copy is the resident one
- compressed in RAM, originals on disk
- recall floor is a product decision
basics
~20 sStart from arithmetic: 50M x 1536 x 4 bytes is roughly 300GB of raw vectors plus graph overhead. Push originals to disk with on_disk=True, pin a quantized copy in RAM with always_ram=True, and buy back recall with oversampling and rescoring — then validate against exact search.
solid answer
~40 sFirst size the problem honestly. Raw float32 storage is dimensions x 4 bytes per vector — about 6KB here, so ~300GB for 50M points, before HNSW links. That will not sit in RAM at sane cost. The standard shape is: `models.VectorParams(size=1536, distance=models.Distance.COSINE, on_disk=True)` so originals are memory-mapped, plus a quantization config with `always_ram=True` so only the compressed copy is resident. Scalar int8 leaves ~75GB; product at X16 leaves ~19GB; binary leaves ~10GB, and 1536 dimensions is squarely in the range where binary works. Then buy recall back with `models.QuantizationSearchParams(oversampling=3.0, rescore=True)` — the wide pass is in RAM, the exact pass hits disk for a few dozen vectors, so NVMe is effectively mandatory. Finally decide sharding, measure recall against `exact=True`, and treat the recall floor as a product decision, not an infrastructure one.
code
python · 15 linesfrom qdrant_client import QdrantClient, models
client = QdrantClient(url="http://localhost:6333")
client.create_collection(
collection_name="corpus",
vectors_config=models.VectorParams(
size=1536,
distance=models.Distance.COSINE,
on_disk=True,
),
quantization_config=models.BinaryQuantization(
binary=models.BinaryQuantizationConfig(always_ram=True)
),
hnsw_config=models.HnswConfigDiff(m=16, ef_construct=200),
)go deeper
Know that vector memory is roughly dimensions times four bytes per point, and that Qdrant can keep a compressed copy in RAM while the full vectors live on disk.
Explain the placement flags — on_disk for vectors, always_ram for the quantized copy, on_disk for the graph — and compute what each quantization mode leaves resident.
Show the measurement discipline: build a table of configuration, resident memory, p95 latency and recall against exact ground truth, and call out storage class as a functional requirement because rescoring does random reads.
Own the whole trade — recall floor negotiated with the product owner, sharding aligned to the query pattern, replication factored into the budget, and enough headroom to re-index a collection of this size while it still serves traffic.
## Do the arithmetic before choosing anything Capacity planning for a vector store starts with a multiplication, not a product feature. Raw vector bytes are `points x dimensions x 4` for float32: 50,000,000 x 1536 x 4 ≈ 307GB. On top of that sits the HNSW graph, whose size scales with `m` and the point count, and the payload store. A common planning rule of thumb is to budget roughly 1.5x the raw vector size to account for index overhead. So the honest starting statement is: this collection does not fit in RAM on any machine you want to pay for, and every subsequent decision is about which representation is resident. ## The layered storage decision Qdrant lets you place three things independently: - **Original vectors** — `models.VectorParams(..., on_disk=True)` memory-maps them instead of holding them in RAM. - **Quantized vectors** — `always_ram=True` in the quantization config pins the compressed copy. - **The HNSW graph** — `models.HnswConfigDiff(on_disk=True)` memory-maps the graph itself. The high-value configuration for this scale is originals on disk plus quantized copy in RAM. Traversal — which touches enormous numbers of vectors — runs entirely on the resident compressed data; disk is read only when rescoring the final candidate set. Putting the graph on disk too is the next lever if RAM is still short, at the cost of page faults during traversal, which hurt far more because they happen on every hop rather than once at the end. ## Sizing each quantization option At 1536 dimensions and 50M points: - **Scalar int8** — one byte per component, ~1.5KB per vector, ~77GB resident. Accuracy loss is usually under a percent. Still a large machine, but a plausible one. - **Product at X16** — ~384 bytes per vector, ~19GB. Fits comfortably, but distances are meaningfully approximate and index building is much slower. - **Binary** — one bit per component, 192 bytes per vector, ~10GB. Cheapest and fastest, and 1536-dimensional embeddings are exactly the regime where the sign pattern retains enough signal. Add the graph on top of whichever you choose, and leave headroom — an instance running at 95% memory has no room for the optimizer to merge segments. ## Recall is the currency you are spending Every step above trades accuracy for cost, so the plan is incomplete without a measured recall number. Sample real query vectors, compute ground truth once with `models.SearchParams(exact=True)`, and evaluate each candidate configuration at your intended `hnsw_ef` and `models.QuantizationSearchParams(oversampling=..., rescore=True)`. The deliverable of this exercise is a small table: configuration, resident memory, p95 latency, recall@10. That table is what makes the decision defensible, and it is what a principal-level answer produces rather than a preference for one mode. Set the recall floor with the product owner, not the infrastructure team. "96% recall@10" and "99.5% recall@10" can differ by a factor of several in machine cost, and only someone who understands the downstream use — RAG grounding, recommendations, dedup — can say which is acceptable. ## Storage class is part of the design With originals on disk, every query performs a bounded number of random reads during rescoring. On local NVMe that is a few milliseconds and the design holds. On network-attached or throttled storage, those reads dominate the request and the architecture quietly inverts into a latency disaster. Specify the storage class explicitly; it is not an infrastructure detail here, it is a functional requirement. ## Horizontal shape At this size, also decide the horizontal layout: how many shards, how many replicas, and whether the workload is naturally partitionable — by tenant, language, or time — so that most queries touch a subset rather than everything. Partitioning that maps onto the query pattern reduces per-query work far more effectively than any parameter, and it caps blast radius when a node is lost. Replicas multiply resident memory by the replication factor, so factor that into the budget rather than discovering it after provisioning. ## Operational headroom Finally, plan for change. Altering `m`, the quantization mode, or the storage flags triggers background re-indexing across the whole collection, which is hours of sustained CPU and disk at this size. Provision so that a re-index can run while the collection still serves traffic, and rehearse it on a copy before doing it on the production collection. A configuration you cannot change safely is a configuration you have to get right the first time — which is exactly the situation measurement is meant to avoid.
- Why not simply put the HNSW graph on disk as well and keep everything cheap?Because graph page faults happen on every hop of every query, not once at the end like rescoring reads. Memory-mapping the graph is a legitimate lever when RAM is genuinely short and latency targets are loose, and it works acceptably on fast local NVMe with a warm page cache. But it degrades tail latency much more sharply than putting original vectors on disk, so it should be the second lever, chosen with measurements rather than as a default.
- How would you justify the recall floor to a product owner?Translate recall into user-visible outcomes rather than percentages. For RAG, a missed neighbour means an answer generated without the relevant passage — so show sampled queries where the quantized configuration lost the true top result and what the downstream answer became. Then present the cost delta between configurations. The decision is a business trade between machine spend and answer quality, and framing it that way gets a real number instead of 'as accurate as possible'.
- What would change your plan if the workload were naturally partitioned by tenant?A great deal. If almost every query is scoped to one tenant, the working set per query is a fraction of the corpus, so you can shard along that boundary and keep only hot partitions fully resident while cold ones rely more heavily on disk. That often beats any quantization choice for both cost and latency, and it caps blast radius on node loss. The trade is skew — a few very large tenants can undo the balance and need special handling.
saying these in an interview costs you the question
- Picks a quantization mode before computing the raw footprint
- Forgets HNSW graph overhead in the memory budget
- Assumes rescoring is cheap on network-attached storage
- Ignores replication factor when sizing resident memory
- Treats the recall floor as an infrastructure preference