How does an Elasticsearch sparse_vector field indexed with ELSER differ from dense_vector kNN?
answer
- One expands into words, one into coordinates
- It rides the inverted index, not a graph
- No candidate queue, no recall dial
- Query cost grows with the expansion width
- English-only, and inference runs on ML nodes
basics
~20 sELSER expands text into weighted vocabulary tokens stored in a sparse_vector field and searched through the ordinary inverted index, so there is no HNSW graph, no num_candidates and no approximation — but query cost grows with the number of expanded tokens.
solid answer
~50 sELSER is a **learned sparse** model: instead of a dense array of floats it emits a set of token-to-weight pairs drawn from a fixed vocabulary, including terms that never appeared in the source text. Those go into a `sparse_vector` field and are indexed as ordinary Lucene postings, so retrieval is a weighted disjunction of term lookups over the inverted index — exact, not approximate, with no graph to build, no `num_candidates` to tune and no recall dial. The trade-offs run the other way: each expanded token is effectively a term clause, so queries touch many postings lists, and ELSER is English-only. You query it with the `sparse_vector` query, which can run inference on the query text through an `inference_id`. A `semantic_text` field wraps the whole pipeline — chunking long text, running inference at ingest and query time — and is queried with the `semantic` query.
code
json · 9 linesPUT /docs
{
"mappings": {
"properties": {
"body": { "type": "text" },
"body_expanded": { "type": "sparse_vector" }
}
}
}go deeper
Know the shape: ELSER turns text into weighted vocabulary tokens stored in a sparse_vector field, and you search it with the sparse_vector query rather than with knn.
Explain that sparse retrieval runs on the ordinary inverted index — exact, filterable, no candidate queue — and that its cost scales with how many tokens the expansion produces.
Bring the operational view: ML node capacity for per-query inference, token pruning as the latency lever, English-only coverage, and where semantic_text's managed chunking helps or hides too much.
Own the retrieval-strategy choice across the corpus — sparse, dense, lexical or a fused combination — including model lifecycle, licensing and the evaluation evidence that justifies the mix.
## Two ways to beat vocabulary mismatch Lexical search fails when the user and the document choose different words for the same idea. There are two families of fix. **Dense retrieval** embeds text into a few hundred or thousand continuous dimensions and finds neighbours geometrically. **Learned sparse retrieval** keeps the vocabulary-shaped representation of classical search but lets a model decide which terms belong and how much each one weighs — including terms the document never contained. ELSER (Elastic Learned Sparse EncodeR) is the second kind. Given a passage, it produces something like `{"laptop": 1.9, "notebook": 1.4, "computer": 1.1, "portable": 0.7, ...}` — an expansion into weighted vocabulary entries. A query is expanded the same way, and matching is the overlap of the two expansions, weighted. ## How it is stored and searched The output goes into a `sparse_vector` field, which stores token-weight pairs. Lucene indexes them as terms with weights, so a search is a big weighted disjunction: for each query token, look up its postings list and accumulate the product of query weight and document weight. That has real architectural consequences relative to `dense_vector`: - **No approximation.** There is no HNSW graph, so there is no `k`/`num_candidates` trade-off and no recall measurement against a brute-force baseline. What you get is what matches. - **Filters compose naturally.** A sparse retrieval clause is a query clause like any other; putting it in a `bool` with `filter` clauses behaves the way every Elasticsearch engineer already expects. There is no pre-filter/post-filter subtlety. - **No graph build cost.** Indexing is ordinary indexing, and segment merges do not rebuild a graph. - **Cost scales with expansion width.** Each expanded token is another postings list to walk. A query expanded into a hundred tokens is a hundred-clause disjunction, and common tokens have long postings lists. This is why token **pruning** exists: the `sparse_vector` query can drop tokens that are very frequent and carry low weight, trading a sliver of relevance for a large latency win. - **Interpretability.** You can look at the expansion. "Why did this match?" has an answer in terms you can read, which is not true of a dense embedding. ## The query surface The `sparse_vector` query is the current way to search these fields. You can hand it a pre-computed set of token weights, or give it query text plus an `inference_id` so Elasticsearch runs the model at query time. An older `text_expansion` query did the same job and was deprecated in 8.15 in favour of `sparse_vector`; on a current cluster, write `sparse_vector`. ## semantic_text — the batteries-included wrapper Wiring inference by hand means an ingest pipeline with an inference processor, a `sparse_vector` field, your own chunking for documents longer than the model's input window, and a query that names the right model. `semantic_text` (introduced in 8.15) collapses that: you map a field with an `inference_id` pointing at an inference endpoint, and Elasticsearch runs inference at ingest, chunks long text automatically, stores the per-chunk representations, and runs query-time inference for the `semantic` query. It works with a sparse model like ELSER and with dense embedding endpoints, so the field type is the abstraction and the endpoint decides which kind of vector is produced. The cost of the abstraction is control: chunking strategy, model choice per field and the exact representation are handled for you, which is exactly right until you need to tune one of them. ## Operational realities ELSER runs as a model deployment on machine-learning nodes. Inference happens on every ingested document and on every query, so ML node capacity becomes part of your search capacity planning, and cold or under-provisioned deployments show up as ingest backpressure and query latency. ELSER v2 is English-only — for other languages you want a multilingual dense model instead. Running the model also requires an appropriate Elastic subscription level rather than being available on every deployment. ## Choosing between them Sparse retrieval tends to be strong out of the box on general English text without domain fine-tuning, and it inherits the operational simplicity of the inverted index. Dense retrieval handles non-English corpora, lets you swap in a domain-tuned model, and supports genuinely non-textual similarity (images, audio, behavioural embeddings). Many production systems run both, plus BM25, and fuse the ranked lists — which is precisely the case reciprocal rank fusion was designed for.
- Why does an ELSER query get slower as the model expands the query into more tokens?Each expanded token becomes a term lookup, so the query is a weighted disjunction over that many postings lists, and frequent tokens have long lists. Doubling the expansion roughly doubles the postings work. Token pruning in the `sparse_vector` query addresses this by discarding tokens that are both very common and low-weight, which removes most of the cost while barely moving relevance.
- What does semantic_text handle that a hand-built sparse_vector pipeline does not?It runs inference automatically at both ingest and query time via an `inference_id`, and it chunks text that exceeds the model's input window, storing and searching the chunks for you. With a hand-built pipeline you own the ingest inference processor, the chunking strategy and the query-time model reference. The trade is convenience against control over chunking and representation.
- When would you choose a dense embedding model over ELSER for semantic retrieval?When the corpus is not English, since ELSER v2 is English-only; when you want to fine-tune or swap the model for a specialised domain; or when the similarity is not textual at all — images, audio, or behavioural embeddings have no vocabulary to expand into. Dense retrieval also gives you a fixed, predictable per-query cost rather than one that grows with expansion width.
A dense embedding is a coordinate on a map of meaning; an ELSER expansion is a generously rewritten set of index-card keywords, each with a weight — still filed in the same card catalogue everything else uses.
saying these in an interview costs you the question
- Calls ELSER output a dense embedding with fewer dimensions
- Thinks sparse retrieval needs num_candidates or an HNSW graph
- Assumes ELSER works well on non-English text
- Ignores that inference runs per query on ML nodes
- Believes semantic_text is just an alias for sparse_vector