skip to content

Your 30M-chunk vector index no longer fits RAM — what do you change?

level: principalimportance: should knowfreq 34%

answer

  1. do the memory arithmetic first
  2. shorter vectors before cleverer indexes
  3. compress to navigate, rescore to rank
  4. partition on a key queries never cross
  5. consider not storing vectors at all

basics

~20 s

Start with the arithmetic — dimensions times bytes times chunks — then work the levers in order: fewer dimensions, quantized vectors with a full-precision rescoring pass, a disk-resident index, sharding along a natural tenancy boundary, and finally the option of not storing vectors at all.

solid answer

~60 s

First quantify. Thirty million chunks at 1024 dimensions in float32 is about 120 GB of vectors before graph overhead, which is why the question is architectural rather than a tuning knob. The levers, cheapest first: **reduce dimensions**, since many modern models support truncating an embedding with modest quality loss; **quantize**, with int8 giving roughly 4x and binary roughly 32x reduction, paired with an oversampled candidate set rescored against full-precision vectors kept on disk so most of the recall comes back; **move the index to disk**, using designs that keep compressed vectors in memory and the graph and full vectors on SSD, which trades a few random reads for a fraction of the RAM; **shard along a boundary queries never cross** — per matter, per tenant, per year — so no single index must be whole and cold shards can live in object storage; and **question the premise**, since for some workloads iterative keyword search over the source performs close to vector retrieval with no index to host at all. Choose using measured recall, the p95 latency budget, corpus churn, and how much operational surface the team can carry.

go deeper

for a junior

Know that a vector index is largely held in memory and that its size is roughly the number of chunks times the vector dimensions times four bytes, so scale is a hardware question, not just a config one.

for a middle

Explain the concrete reduction levers — fewer dimensions, int8 or binary quantization, disk-resident index layouts — and why quantized search is paired with an oversampled candidate set rescored against full-precision vectors.

for a senior

Show the measurement discipline: state the recall and p95 requirement, shadow-test a candidate configuration against exact ground truth on real traffic, and watch the tail rather than the mean before cutting over.

for a principal

Own the architecture and its reversibility — order the levers by cost to undo, pick a partition key queries never cross, decide between the database you already run and a dedicated engine on operational grounds, and be willing to conclude the workload needs no vector store at all.

## Do the arithmetic before choosing anything The memory bill is not mysterious. Vectors cost `chunks * dimensions * bytes_per_component`. Thirty million chunks of 1024-dimensional float32 is roughly 120 GB. A graph index adds roughly `8 * m` bytes per vector for its links — about 8 GB at m=32 — plus payloads and identifiers. Any answer that starts with an index parameter rather than this calculation is answering the wrong question, because no HNSW tuning changes the leading term. The second number to establish is the requirement: what recall does this workload actually need, at what p95 latency, at what query rate, and how often does the corpus change? A legal e-discovery matter under a production deadline has a defensibility obligation that makes recall close to non-negotiable and tolerates a slower query; a consumer type-ahead is the reverse. These two numbers — the bill and the requirement — frame every lever below. ## Lever 1: fewer dimensions The cheapest reduction is to store shorter vectors. Several current embedding models are trained so that a prefix of the vector is itself a usable embedding, which means you can truncate — say 1024 down to 512 or 256 — and lose less quality than the size reduction would suggest. Halving dimensions halves the leading term and speeds every distance computation. It requires re-embedding or re-slicing and a recall re-measurement, so validate on a labelled set before committing. ## Lever 2: quantization plus rescoring Quantization stores each component in fewer bits. Scalar quantization to int8 is about 4x smaller; binary quantization, one bit per dimension, is about 32x smaller and turns distance computation into extremely fast bitwise operations. Used naively, aggressive quantization loses meaningful recall. Used properly, it is paired with **rescoring**: search the quantized index for an oversampled candidate set — several times k — then recompute exact distances for those candidates against full-precision vectors held on disk or in a cheaper tier, and take the true top-k. The compressed index does the navigation; the full vectors do the final ordering. This pattern is the workhorse of large deployments because it keeps the hot structure small while preserving most of the accuracy, at the cost of one extra read-and-rescore step per query. ## Lever 3: disk-resident index designs A graph index whose working set exceeds RAM does not degrade gracefully — every hop becomes a random page fault. Indexes designed for disk residency, in the DiskANN family and its Postgres implementations such as pgvectorscale, invert the layout deliberately: compressed vectors stay in memory to guide the search, while the graph and full-precision vectors live on SSD, and the graph is built so that a query needs only a small number of reads. The result is a dramatically smaller memory footprint at a latency cost measured in milliseconds rather than the catastrophic cliff of swapping. This is the standard answer when the corpus genuinely exceeds affordable RAM and recall must stay high. ## Lever 4: shard along a boundary queries never cross In many domains there is a partition key that no legitimate query spans — one legal matter, one tenant, one fiscal year. Sharding on it means no single index must be whole: hot shards stay resident, cold shards live in object storage and are loaded on demand. Object-storage-native and multi-tenant-first designs make this the primary architecture rather than an afterthought, trading a cold-start penalty on a first query for very cheap storage of a long tail. It also simplifies filtering, since the partition key stops being a predicate the index must handle, and it makes deletion of an entire tenant or matter a file operation. The cost is many indexes to operate and a routing layer that must be correct. ## Lever 5: question whether you need vectors at all As of mid-2026 this is a live option rather than a provocation. When an agent can issue iterative keyword searches against the source — grepping a repository, querying an existing full-text index — it can refine its own queries across turns in a way a single embedding lookup cannot. Reported results have put agentic keyword search close to vector retrieval on answer faithfulness for some corpora, and at least one prominent coding tool dropped its local vector index in favour of it. The trade is more model calls and higher per-query latency against no embedding pipeline, no index to host, no reindexing on every document change, and no stale-embedding problem. For corpora that change constantly, or where the queries are naturally lexical, that trade can be strongly favourable. For large heterogeneous prose corpora where paraphrase matching matters, it is usually not. ## Deciding Order the levers by how much they cost you to reverse. Truncation and quantization are configuration and a re-index. A disk-resident index is a deployment change. Sharding is an architectural commitment that reaches into routing and lifecycle. Dropping vector search entirely is a product decision. Take them in that order, and after each one re-measure recall against exact search on the same labelled query set, because that is the quantity you are spending. The supporting choice — whether this lives in the Postgres you already operate or in a dedicated vector engine — should follow the numbers, not enthusiasm. Keeping vectors alongside the relational data you already back up, secure and monitor is worth a great deal, and Postgres extensions now cover both graph and disk-resident designs. A dedicated engine earns its operational cost when query volume, filtering sophistication or multi-tenant scale exceed what the general-purpose database is happy doing. Agent-driven workloads push in that direction, since agents issue far more queries per task than a human does. ## Where teams go wrong Buying more RAM as a reflex, which postpones the decision at increasing cost; quantizing without a rescoring pass and discovering the recall loss in production; sharding on a key that queries do turn out to cross; and — most common — never establishing the recall requirement, so no one can say whether any of these trades was acceptable.

  • Why is binary quantization usually paired with a rescoring pass?
    Because one bit per dimension discards enough information that the quantized ranking is only approximately right, even though it is excellent at identifying a neighbourhood. The pattern is to retrieve several times more candidates than you need from the compact index, then recompute exact distances for just those candidates against full-precision vectors kept on cheaper storage, and re-sort. You keep most of the memory saving and recover most of the recall, paying one extra read-and-rescore step per query.
  • When does agentic keyword search over the source beat maintaining a vector index?
    When the corpus changes faster than you can re-embed it, when queries are naturally lexical — identifiers, symbols, file paths — and when an existing full-text search is already operated and trusted. The agent refines its own queries across turns, which recovers much of what a single embedding lookup provides. You pay more model calls and higher per-query latency, and give up strong paraphrase matching over large heterogeneous prose, which is exactly where vector retrieval still wins.
  • What argues for keeping vectors in the relational database you already run rather than a dedicated engine?
    Operational gravity. The data is already backed up, access-controlled, monitored and joinable against the rows it describes, and the team already knows how to run it, which removes an entire system from the failure surface. Modern Postgres extensions cover both in-memory graph and disk-resident designs, so the capability gap has narrowed. A dedicated engine earns its keep when query volume, filtering sophistication or multi-tenant scale exceed what the general-purpose database serves comfortably.
  • How would you validate a move to a quantized or disk-resident index before cutting over?
    Keep a labelled query set and exact brute-force ground truth for it, then run the candidate configuration in shadow against production traffic. Compare recall@k, p95 and p99 latency, and memory footprint against the current index on the same queries. Watch the tail rather than the mean, since disk-resident designs move variance more than averages, and only cut over once the tail fits the budget at your real query concurrency.

saying these in an interview costs you the question

  • Just add RAM; the index is supposed to be in memory
  • Quantization is lossless if you pick the right number of bits
  • Sharding is only about throughput, not about memory
  • A dedicated vector engine is always better than vectors in Postgres
  • Vector search is the only viable retrieval architecture at scale

context