skip to content

How would you plan Elasticsearch capacity for 100 million 1024-dimension vectors on a fixed RAM budget?

level: principalimportance: should knowfreq 38%

answer

  1. Start with dimensions times bytes times documents
  2. Vector data lives off-heap in page cache
  3. Compression ratios come from index_options
  4. Full-precision vectors stay on disk for rescoring
  5. A smaller embedding model beats any tuning

basics

~20 s

Start from the arithmetic: 100M by 1024 float32 dimensions is roughly 410 GB before the graph. Quantized index_options cut that several-fold, and the target is keeping the searchable vector data and HNSW graph resident in off-heap page cache.

solid answer

~50 s

Do the arithmetic first. 100M × 1024 × 4 bytes ≈ **410 GB** of raw float32 vectors, plus an HNSW graph of roughly `num_vectors × m × 2 × 4` bytes — about 13 GB at the default `m` of 16. Vector data is memory-mapped off-heap, so the design target is that the searchable representation fits in filesystem page cache across the data nodes; once it spills to disk, kNN latency falls off a cliff. That is what quantization is for: `int8_hnsw` stores one byte per dimension (~4×smaller), `int4_hnsw` ~8×, and `bbq_hnsw` around a bit per dimension (~32×), with the full-precision vectors kept on disk so a rescoring pass can recover the recall the compression cost. Then size shards so each node's graphs fit, keep JVM heap modest to leave room for page cache, and validate the chosen quantization by measuring recall@k against an exact baseline before committing.

code

json · 18 lines
json
PUT /vectors
{
  "mappings": {
    "properties": {
      "embedding": {
        "type": "dense_vector",
        "dims": 1024,
        "index": true,
        "similarity": "cosine",
        "index_options": {
          "type": "bbq_hnsw",
          "m": 16,
          "ef_construction": 100
        }
      }
    }
  }
}

go deeper

for a junior

Be able to do the multiplication — documents times dimensions times bytes per dimension — and know that quantized index_options exist to shrink it.

for a middle

Explain that vector data is memory-mapped off-heap, what each quantization type costs in fidelity, and why full-precision vectors are retained on disk.

for a senior

Show the operating judgment: keep heap small so page cache holds the graphs, avoid over-sharding, plan for graph rebuilds during merges, and validate quantization with a recall measurement.

for a principal

Own the whole cost curve — embedding dimensionality, quantization, replica count, node class and hot/cold tiering — and the evaluation programme that proves the chosen point meets the product's recall requirement.

## Step 1 — the arithmetic Capacity planning for vectors starts with multiplication, not intuition. - Raw float32 vectors: `100,000,000 × 1024 × 4 bytes ≈ 410 GB`. - HNSW graph connections: roughly `num_vectors × m × 2 × 4 bytes`. At the default `m` of 16 that is about 13 GB — small next to the vectors, but it grows linearly with `m`. - Everything else in the index: `_source`, other fields, doc values, and transient space for merges. Vector fields dominate, but they are not the whole index. That 410 GB figure is the number to react to. It does not fit on one node, and it will not sit in page cache anywhere cheap. ## Step 2 — understand where the memory lives Vector data and HNSW graphs are memory-mapped and live **off-heap**, in the operating system's page cache. This inverts the familiar Elasticsearch reflex of giving the JVM lots of heap: for a vector workload you want a *modest* heap and as much free RAM as possible for the OS to cache the vector files. When the working set exceeds page cache, each graph traversal turns random reads into disk I/O, and because HNSW hops are random by nature, this is the worst possible access pattern for anything but fast local NVMe. So the real design constraint is: **which representation must be resident**, and does it fit? ## Step 3 — quantization as the primary lever Elasticsearch's `index_options` types trade fidelity for size: - **hnsw** — raw float32. 410 GB here. Highest fidelity, highest cost. - **int8_hnsw** — scalar quantization to one byte per dimension, roughly a 4× reduction (~102 GB). This is the default posture for float vectors in recent 8.x precisely because the recall loss is usually small. - **int4_hnsw** — roughly 8× (~51 GB), with more visible recall loss. - **bbq_hnsw** — better binary quantization, around one bit per dimension, roughly 32× (~13 GB). Suddenly the working set fits comfortably on a handful of nodes. Crucially, the quantized form is what the graph search uses, while the original float vectors remain on disk. That enables **rescoring**: search approximately over the compact representation, then re-score an oversampled set of top hits against the full-precision vectors. In 8.18+/9.x this is exposed on the kNN search as `rescore_vector` with an `oversample` factor, and it recovers most of the recall that aggressive quantization gives up — you fetch, say, three times the candidates from the compact index and re-rank them exactly. The cost is extra disk reads on the top slice only, which is a far better trade than keeping 410 GB resident. Binary quantization also generally wants higher-dimensional vectors to work well; very low-dimensional vectors have too little information per vector to survive one bit per dimension. ## Step 4 — shard and node layout With a target representation size in hand: - Size shards so that the vector data on any one node fits its available page cache with headroom. Larger shards than the log-workload habit suggests, because every shard pays a full `num_candidates` traversal per query — over-sharding multiplies query cost without improving relevance. - Consider dedicated nodes for the vector indices so a heavy aggregation workload elsewhere cannot evict vector pages from cache. - Replicas double the resident footprint but also double query throughput; that is a throughput-versus-RAM decision, not a durability one alone. - If the corpus is time-partitioned and old data is rarely searched semantically, keep the graph only on hot indices and let cold ones fall back to lexical search or exact scoring. ## Step 5 — build cost and merge behaviour HNSW graphs are built at index time and **rebuilt when segments merge**, because a merged segment needs a single coherent graph. That makes vector indices expensive to merge and makes `force_merge` on a large vector index a genuinely heavy operation, not the routine housekeeping it is for a text index. Plan bulk loads with that in mind: index with a relaxed refresh interval, expect merges to be CPU-hungry, and treat a full reindex as a multi-hour capacity event. `ef_construction` and `m` govern build effort and graph quality. Raising them improves recall and raises both build time and graph size. They are index-time decisions; changing them means rebuilding. ## Step 6 — validate before committing None of the above is worth anything without a recall measurement: 1. Sample a realistic query set. 2. Compute exact ground truth for those queries. 3. Measure recall@k for each candidate configuration — quantization type, `m`, `ef_construction`, `num_candidates`, oversampling factor. 4. Plot recall against p99 latency and against RAM cost, and pick the configuration that meets the product's recall requirement most cheaply. The strategic point a lead is expected to make: the cheapest capacity win is often upstream of Elasticsearch. A 1024-dimension model where a 384-dimension one performs comparably on your evaluation set cuts the footprint by nearly two-thirds before any quantization at all — and unlike quantization, it also cuts inference cost, ingest time and network transfer.

  • Why is force_merge unusually expensive on an index with dense_vector HNSW fields?
    A merged segment needs a single coherent HNSW graph, so merging does not just concatenate postings — it rebuilds the graph over the combined vectors, which is CPU-heavy and grows with `m` and `ef_construction`. On a large vector index that turns routine housekeeping into a multi-hour operation. Plan it as a capacity event, run it off-peak, and do not treat it as a free optimisation.
  • How does rescoring recover recall that quantization gave away?
    The graph search runs over the compact quantized vectors and returns an oversampled candidate list — several times the requested k. Those candidates are then re-scored against the full-precision vectors kept on disk and re-ranked. Compression errors that mis-order near neighbours get corrected on the top slice only, so you pay a few extra disk reads per query rather than keeping the uncompressed data resident.
  • What is the argument for changing the embedding model instead of tuning the index?
    Footprint scales linearly with dimensions, so moving from 1024 to 384 dimensions removes roughly two-thirds of the data before any quantization, and it compounds with quantization rather than competing with it. It also reduces inference cost, ingest throughput requirements and network transfer. The prerequisite is evidence: measure the smaller model's nDCG on your judged query set before assuming the quality is comparable.

It is warehouse planning: measure the pallets first, then decide whether to compress the goods, rent more floor space, or ship a smaller product. Compressing everything and hoping is not a plan.

saying these in an interview costs you the question

  • Sizes JVM heap for vectors instead of leaving page cache
  • Assumes vector data is stored on the JVM heap
  • Picks a quantization level without measuring recall
  • Over-shards the vector index for parallelism
  • Treats force_merge on a vector index as routine housekeeping

context