skip to content

kNN & Hybrid Retrieval

dense_vector fields plus an HNSW index give Elasticsearch approximate nearest-neighbour search, and retrievers fuse that list with a BM25 list using reciprocal rank fusion. Interviewers want the trade-offs: index build cost, num_candidates versus recall, and how filtered kNN behaves.

part ofElasticsearchoverview, primer and where to startread it →
on this pageshow

questions

6

What must an Elasticsearch dense_vector mapping declare before a knn search can use it?

level: juniorimportance: must knowfreq 68%

answer

  1. Three things are pinned at mapping time
  2. Length, graph, and distance function
  3. index: false still allows script scoring
  4. The metric must match the embedding model
  5. dims and similarity need a reindex to change

basics

~10 s

The field must be typed dense_vector with a fixed dimension count, left indexed so an HNSW graph is built, and given a similarity metric (cosine, dot_product, l2_norm or max_inner_product) that matches the embedding model.

solid answer

~40 s

A `dense_vector` field pins three things. **dims** fixes the vector length — every document must supply exactly that many numbers, and in recent 8.x the value can be inferred from the first indexed vector instead of declared. **index** decides whether Elasticsearch builds an HNSW graph for approximate nearest-neighbour search; when it is off the field is still stored and usable from `script_score` vector functions, but a `knn` search cannot touch it. **similarity** chooses the distance function used for scoring — `cosine`, `dot_product`, `l2_norm` or `max_inner_product` — and it must match how the embedding model was trained. `element_type` (`float`, `byte`, `bit`) and `index_options` (graph type, `m`, `ef_construction`) round it out. dims and similarity are not editable in place: getting them wrong means reindexing into a new field.

code

json · 19 lines
json
PUT /articles
{
  "mappings": {
    "properties": {
      "title": { "type": "text" },
      "title_embedding": {
        "type": "dense_vector",
        "dims": 768,
        "index": true,
        "similarity": "cosine",
        "index_options": {
          "type": "int8_hnsw",
          "m": 16,
          "ef_construction": 100
        }
      }
    }
  }
}

go deeper

for a junior

Be ready to write the mapping from memory: type dense_vector, the dimension count, indexed, and a similarity. Know that the query vector must come from the same model as the indexed vectors.

for a middle

Explain what index: true actually builds, why cosine normalizes at index time while dot_product demands unit vectors, and what element_type and index_options change about storage.

for a senior

Show that dims and similarity are immutable and that the real cost of getting them wrong is a full reindex behind an alias. Discuss quantized index_options as the default posture rather than raw hnsw.

for a principal

Own the model-versioning story: how embedding-model upgrades are rolled out without downtime, how two vector fields coexist during a cutover, and what recall evidence justifies the switch.

## What a dense_vector field is `dense_vector` is the Elasticsearch field type that holds a fixed-length array of numbers — typically the output of an embedding model, where semantically similar text lands close together in the vector space. Unlike `text` or `keyword`, it is not analysed into terms; it is stored as a numeric array, and optionally indexed into a graph structure that supports approximate nearest-neighbour (ANN) lookup. ## dims — the length contract Every vector in the field must have the same number of dimensions. In older versions you declared `dims` explicitly; in recent 8.x you may omit it and Elasticsearch fixes it from the first document indexed. Either way it becomes immutable for that field. Indexing a 768-dimension vector into a field fixed at 384 is rejected. Because embedding models have a fixed output width, switching models almost always means a new field and a reindex. ## index — whether ANN is available `index: true` (the default in recent 8.x) tells Elasticsearch to build an HNSW graph over the field's vectors, which is what the `knn` search option traverses. With `index: false` the vectors are stored but there is no graph: you can still read them and score them with `script_score` vector functions such as `cosineSimilarity` or `dotProduct`, but that is an exhaustive scan over every document the surrounding query matches. Exhaustive scoring is exact and perfectly reasonable for a few thousand candidates; it is hopeless for millions. ## similarity — the distance function The `similarity` parameter selects how closeness is measured and how the raw distance is turned into a positive `_score`: - **cosine** — angle between vectors. Elasticsearch normalizes vectors to unit length at index time, so internally it degenerates to a dot product. Score is `(1 + cosine) / 2`, so it lands in [0, 1]. - **dot_product** — raw inner product, and Elasticsearch requires the indexed and query vectors to already be unit length; a vector with a different magnitude is rejected at index time. If your model already emits normalized vectors, this is the cheapest choice. - **l2_norm** — Euclidean distance, turned into a score that decreases as distance grows. - **max_inner_product** — inner product without the unit-length requirement, for models where magnitude carries meaning. The metric must match the model. Embeddings trained under cosine similarity ranked with `l2_norm` will not be catastrophically wrong, but they will be subtly worse, and that is the kind of relevance bug nobody notices for months. ## element_type and index_options `element_type` defaults to `float` (32-bit). `byte` stores each dimension in one signed byte, and `bit` stores one bit per dimension for models that emit binary codes — both shrink the field dramatically at some cost in fidelity. `index_options` picks the graph flavour and its build parameters: `hnsw` for raw float graphs, `int8_hnsw` / `int4_hnsw` / `bbq_hnsw` for quantized graphs, and `flat` variants for brute-force-only storage. Recent 8.x defaults float vectors to a scalar-quantized `int8_hnsw` graph rather than raw `hnsw`, because the memory saving usually outweighs the small recall loss. `m` and `ef_construction` inside `index_options` control graph connectivity and build effort. ## What you cannot change later Dimensions, similarity and element type are baked into the field. There is no in-place mapping update for them; the fix is a new field (or a new index) plus a reindex, usually behind an alias so search traffic never sees the switch. This is why the mapping conversation happens before the first bulk load, not after. ## Common failure modes A `knn` search that errors with a message about the field not being indexed for search means `index: false`. A search that returns nothing sensible often means the query vector came from a different model than the indexed ones — same dimension count, completely different space, and nothing in Elasticsearch can detect that for you. An index-time rejection about vector magnitude means `dot_product` with unnormalized vectors. And a field that quietly eats a lot of memory usually means raw `hnsw` on float32 vectors where a quantized graph would have done.

  • If a dense_vector field is mapped with index set to false, how would you still rank documents by vector similarity?
    Run a normal query to narrow the candidate set, then wrap it in a `script_score` query using a vector function such as `cosineSimilarity(params.qv, 'my_vector')`. That is an exact, exhaustive comparison over every matching document, so it only stays affordable when the surrounding query already cuts the candidate set down to something small — thousands, not millions.
  • Your team switches from a 768-dimension model to a 1024-dimension one. What is the migration?
    You cannot widen the existing field. Add a new `dense_vector` field (or build a new index) with the new dims and similarity, re-embed and reindex every document, then flip an alias so searches move over atomically. Run both fields side by side long enough to compare recall and latency before deleting the old one.
  • Why does Elasticsearch reject some vectors when similarity is dot_product but accept them under cosine?
    `dot_product` is only meaningful for unit-length vectors, so Elasticsearch validates the magnitude at index time and rejects vectors that are not normalized. `cosine` accepts any magnitude because it normalizes vectors itself at index time. If your model already emits unit vectors, `dot_product` skips that normalization work.

Think of it like declaring a fixed-width numeric column and choosing its collation up front: the width decides what rows fit, and the collation decides what 'close' means. Neither can be renegotiated once the data is in.

saying these in an interview costs you the question

  • Thinks dims can be changed later with a mapping update
  • Says knn works on a field mapped with index false
  • Picks a similarity at random instead of matching the model
  • Assumes cosine and dot_product are interchangeable without normalization
  • Believes dense_vector fields are analysed into terms like text

context

open as a page

In an Elasticsearch knn search, how do k and num_candidates differ?

level: middleimportance: must knowfreq 76%

basics

~20 s

k is how many nearest neighbours the search returns; num_candidates is how many candidates each shard explores in the HNSW graph before picking its best k. Raising num_candidates buys recall at the cost of latency.

open as a page

How does an Elasticsearch sparse_vector field indexed with ELSER differ from dense_vector kNN?

level: middleimportance: should knowfreq 48%

basics

~20 s

ELSER expands text into weighted vocabulary tokens stored in a sparse_vector field and searched through the ordinary inverted index, so there is no HNSW graph, no num_candidates and no approximation — but query cost grows with the number of expanded tokens.

open as a page

How does the filter inside an Elasticsearch knn search differ from a post_filter?

level: seniorimportance: should knowfreq 58%

basics

~20 s

The knn filter is applied during the HNSW graph traversal, so all k neighbours already match it. A post_filter runs after the neighbours are chosen and simply discards some, often leaving far fewer than k hits.

open as a page

How does Elasticsearch's rrf retriever combine a BM25 query with a kNN search?

level: seniorimportance: should knowfreq 60%

basics

~20 s

The rrf retriever runs each sub-retriever separately, then fuses their ranked lists by summing 1 divided by (rank_constant plus each list's rank). It uses positions, not scores, so incomparable BM25 and vector scores never have to be normalized.

open as a page

How would you plan Elasticsearch capacity for 100 million 1024-dimension vectors on a fixed RAM budget?

level: principalimportance: should knowfreq 38%

basics

~20 s

Start from the arithmetic: 100M by 1024 float32 dimensions is roughly 410 GB before the graph. Quantized index_options cut that several-fold, and the target is keeping the searchable vector data and HNSW graph resident in off-heap page cache.

open as a page