skip to content

Elasticsearch

3 roadmaps144 questionsupdated

The default answer when a product needs full-text search, log analytics, or aggregations over billions of documents: Lucene indexing wrapped in a distributed cluster with a JSON query DSL. Interviews walk the whole stack — how a field is mapped, how a query scores, and how shards keep it running.

on this pageshow

guide

overview

~1 min

Elasticsearch is a search and analytics engine built on Apache Lucene: it turns JSON documents into inverted indexes, spreads them across shards on a cluster of nodes, and answers full-text queries, exact filters and aggregations through one JSON query language. Interviewers use it to test whether you understand what a search engine does at write time, because nearly every surprise at read time — a query that finds nothing, a score that looks wrong, a dashboard that times out — was settled when the document was indexed. The hub follows that write-then-read path. [Indices and mappings](/topics/db-elasticsearch-indices-mappings) decide what each field is and what may be done with it, and why changing that later means building a new index. [Analyzers and tokenization](/topics/db-elasticsearch-analysis) decide which terms a text field actually holds. The [Query DSL](/topics/db-elasticsearch-query-dsl) is how you ask for documents, and [relevance scoring](/topics/db-elasticsearch-relevance-scoring) is why they come back in the order they do, including vector and hybrid retrieval. [Aggregations](/topics/db-elasticsearch-aggregations) are the analytics half, and the part most likely to exhaust a node's memory. [Sharding and replication](/topics/db-elasticsearch-sharding-replication) is the cluster underneath all of it: where a document lives, when it becomes visible, and what survives a lost node. Junior rounds stay on the model — `text` against `keyword`, `match` against `term`, what yellow health means. Senior and staff rounds turn into design and incident conversations: sizing shards for a log pipeline, evolving a mapping with no downtime, explaining a ranking to a product manager, keeping aggregations inside the heap. Start with mappings and analysis. Queries, scoring and aggregations only make sense once you know what the index holds, and the cluster sections assume all four.

primer

### Search reads terms, not documents Elasticsearch keeps the JSON you sent in `_source`, but ordinary queries do not search it. Each field is converted at index time into structures built for one kind of access: an **inverted index** from terms to documents for search, and columnar **doc values** for sorting and aggregations. A question about why something matches, sorts or aggregates is almost always a question about which of those structures the field has. ### The mapping is a one-way decision The mapping fixes each field's type and indexing options before data arrives, and Lucene segments are immutable, so an existing field cannot be retyped in place. Adding fields is cheap; changing one means building a fresh index and copying the data across. That is why aliases, templates and data streams matter so much in practice: they are the tools that make rebuilding an index routine instead of an outage. ### Analysis decides what can match A `text` value is broken into terms by an analyzer, and a full-text query analyzes its input too. Two sides produce terms, and a hit needs them to meet. Most "search returns nothing" and "search returns everything" reports come down to the two sides disagreeing, or to a query that skipped analysis on a field that had it. ### Some clauses score, some only select Every clause either contributes to relevance or simply includes and excludes documents. Scoring clauses are ranked with **BM25** by default; non-scoring filters are cheaper and cacheable. Interviewers probe whether each condition sits in the right place, and whether you read the scoring breakdown before reaching for boosts. ### Aggregations are per-shard and often approximate Aggregations read doc values, run on every shard in parallel, and are merged on the coordinating node. Merged partial results mean counts can be estimates, and some metrics use fixed-size sketches by design. A strong answer says which numbers are exact and what exactness would cost. ### Shards are the unit of scale — and the fixed choice An index is split into primary shards, each a complete Lucene index, with replicas copying them. The primary count is part of how documents are placed, so it is set at creation; replicas can change at any time. Too few shards cap growth, and too many burn heap and coordination on overhead, which is why sizing is a standing design question. ### Near-real-time, and rarely the system of record A write is acknowledged before it is searchable; visibility waits for the next refresh. Durability comes from the transaction log and replicas, and concurrent updates are guarded optimistically. Most architectures still feed Elasticsearch from a primary database and treat the index as rebuildable.

Mapping
The schema of an index: each field's type and how it is indexed and stored. Existing fields cannot be retyped in place.
Inverted index
The Lucene structure that maps each term to the documents containing it, which is what full-text and exact queries look up.
Doc values
An on-disk, column-oriented copy of a field's values per document, used for sorting, aggregations and scripts. Enabled by default for most non-text types.
text field
A string field passed through an analyzer into individual terms, built for full-text search rather than exact matching.
keyword field
A string field indexed as one unanalyzed term, used for exact filters, sorting and aggregations.
Analyzer
The pipeline of character filters, one tokenizer and token filters that turns a string into the terms stored in or searched against the index.
Query context
Where a clause answers how well a document matches and contributes to its _score. Filter context only includes or excludes, and can be cached.
BM25
The default relevance function: it weighs how often a term appears in a field, how rare it is across documents, and how long the field is.
Primary shard
One partition of an index and a full Lucene index in its own right. Every document belongs to exactly one, and the count is set at creation.
Replica shard
A copy of a primary kept on a different node for failover and extra read capacity. The replica count can change on a live index.
Segment
An immutable piece of a Lucene index. New documents create new segments, background merges combine them, and deletes are only marked until a merge.
Refresh
The periodic operation that makes recently indexed documents visible to search by opening a new segment.
Translog
A per-shard write-ahead log of operations not yet committed to Lucene, replayed after a crash to recover acknowledged writes.
Index alias
A secondary name pointing at one or more indices, so clients keep one stable name while the index behind it is replaced.
Data stream
An append-only name over a sequence of hidden, time-based backing indices, used for logs, metrics and other timestamped data.
Index lifecycle management (ILM)
Policies that move time-based indices through phases, rolling over, shrinking, moving to cheaper nodes and finally deleting them.
Coordinating node
The node that receives a request, fans it out to the relevant shards and merges their partial results. Any node plays this role.

Follow one product document from write to result. The client sends it to any node, which routes it to one primary shard computed from its routing value. The primary applies the mapping — creating fields dynamically if the mapping allows — runs each `text` value through its analyzer, and writes terms, doc values and `_source` into an in-memory buffer plus the translog. It then forwards the operation to the replicas and acknowledges. Nothing is searchable yet; the next refresh publishes a new segment, and background merges later fold small segments into larger ones. A search travels the other way. The coordinating node sends it to one copy of every shard involved. Each shard analyzes the query text, finds candidates in the inverted index, scores the scoring clauses, applies filters and computes its share of every aggregation. The coordinating node merges the per-shard top hits and aggregation partials, then fetches `_source` only for the documents that made the final page. One request shows how the sections meet: ```json { "query": { "bool": { "must": [ { "match": { "name": "trail running shoes" } } ], "filter": [ { "term": { "brand": "acme" } }, { "range": { "price": { "lte": 120 } } } ] } }, "aggs": { "by_brand": { "terms": { "field": "brand" } } }, "size": 20 } ``` - The `match` works only because `name` is mapped as `text`, and it matches only the terms its analyzer produced. It is the one clause that shapes the ranking. - The `term` and `range` clauses sit under `filter`, so they add nothing to the score and can be cached across requests; the `term` matches the stored value exactly because `brand` is a `keyword`. - The terms aggregation reads doc values for `brand`, computed per shard and merged, which is where count accuracy and memory questions start. - `size` and the shard count together decide how much each shard returns and the coordinating node sorts, which is what makes deep pagination expensive. Operations attach to the same objects: templates and ILM create and retire indices around this path, and cluster health and shard allocation decide whether every shard it needs has a live copy.

  1. Indices and Mappings →

    Field types, dynamic mapping and why mapping changes need a reindex — every other section assumes you know what the index holds.

  2. Analyzers and Tokenization →

    How text becomes terms on both the index and query side, which explains most missing or unexpected hits.

  3. Query DSL →

    Full-text versus term-level queries, bool composition and filter context, then nested data and pagination.

  4. Relevance Scoring →

    Once queries are clear: how BM25 ranks hits, how boosts and function_score bend it, and where vector retrieval fits.

  5. Aggregations →

    The analytics half, read after the query model because aggregations run over whatever the query matched.

  6. Sharding and Replication →

    The cluster underneath: routing, the search phases, refresh and durability, and diagnosing unassigned shards.

  • Running a term query against a text field and concluding the data is missing, when the stored terms were analyzed and the query string was not.

  • Putting pure yes/no conditions under must instead of filter, which spends scoring work on them and gives up caching.

  • Treating the mapping as editable: a field's type or analyzer cannot be changed in place, so the plan must include a new index, a reindex and an alias switch.

  • Leaving dynamic mapping open on user-controlled keys, then hitting the field limit or a mapping so large it slows the whole cluster.

  • Aggregating or sorting on a text field by enabling fielddata instead of adding a keyword sub-field that already has doc values.

  • Quoting terms aggregation counts or distinct counts as exact on a multi-shard index without saying they can be estimates.

  • Choosing the primary shard count as if it were adjustable later, or creating thousands of tiny shards that cost more in overhead than they hold.

  • Expecting a document to be searchable the moment it is acknowledged, and forcing a refresh on every write to hide it.

  • Answering deep pagination with ever larger from values instead of search_after with a point in time.

This guide assumes Elasticsearch 8.x or later. Several interview questions still turn on what changed across the 6.x, 7.x and 8.x lines, because older clusters and client code stay in production for years: - **Mapping types** were narrowed to one per index in 6.x, deprecated in requests in 7.0 and removed in 8.0. Answers that mention a type name besides `_doc` describe an old cluster. - **7.0** lowered the default primary shard count from five to one, replaced the old discovery setting that required configuring a master quorum by hand, and stopped counting total hits exactly beyond a default threshold unless asked to. - **The 7.x line** introduced composable index and component templates, data streams, and point-in-time searches for consistent deep paging. - **8.0** turned security on by default, so a fresh install configures TLS and authentication out of the box. - **The 8.x line** added approximate kNN search on HNSW graphs, and later retrievers and reciprocal rank fusion for hybrid search. When an answer depends on one of these, name the version you are describing.

Elasticsearch is the storage and query engine of the Elastic Stack, from Elastic: Kibana sits on top for dashboards and exploration, while Logstash, Beats and Elastic Agent collect and ship data into it. Interviewers expect you to place it in an architecture, not only query it. In most product systems it is a secondary store: a relational or document database owns the data, and the index is fed from it by change events or application writes and can be rebuilt. Its closest neighbours share its roots. Apache Solr is the other long-standing search server built on Lucene. OpenSearch is the fork created after Elastic changed its license in 2021, starting from the 7.10 code base and diverging since, so version-specific answers do not always carry across. Against a relational database's built-in full-text search, the trade is operational weight for richer relevance, analysis and scale. For log analytics it competes with columnar analytical databases, which favour aggregate speed and cost over full-text search; for semantic search it competes with dedicated vector databases, where Elasticsearch's argument is combining vectors with text relevance and filters in one query.

explore

report an issue with this guide →

questions

144 · 6 sections

How does an Elasticsearch data stream differ from writing to a single regular index?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A data stream is one name in front of an ordered set of hidden backing indices. Writes are append-only and always land in the newest backing index; searches and rollover target the stream name, not any single index.

open as a page

In an Elasticsearch mapping, what do the dynamic values true, runtime, false and strict each do?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Elasticsearch's dynamic setting decides what happens to an unmapped field: true indexes it and adds it to the mapping, runtime adds it as a query-time field only, false ignores it but keeps it in _source, strict rejects the document.

open as a page

In an Elasticsearch mapping, how do the text and keyword field types differ?

level: juniorimportance: must knowfreq 88%
basics
~10 s

text is analyzed into tokens for full-text search and cannot be sorted or aggregated by default. keyword stores the value verbatim as one term, which is what filters, sorts and aggregations need.

open as a page

In Elasticsearch, what does an index template configure and when is it applied?

level: juniorimportance: must knowfreq 68%
basics
~20 s

An index template attaches settings, mappings and aliases to any index whose name matches its index_patterns. It is applied once, at index creation time; indices that already exist keep whatever configuration they were created with.

open as a page

Why do Elasticsearch applications query an index alias instead of the concrete index name?

level: juniorimportance: must knowfreq 72%
basics
~20 s

An alias is a second name that points at one or more indices. Applications talk to the alias so the physical index behind it can be rebuilt, versioned and swapped atomically, with no client change and no downtime.

open as a page

How do you use the _analyze API to see the exact tokens Elasticsearch produced for a field?

level: juniorimportance: must knowfreq 66%
basics
~20 s

Call POST /_analyze with an analyzer name and text to test any built-in analyzer, or POST /<index>/_analyze with a field and text to run that field's configured analyzer. Add explain: true to see each pipeline stage's output.

open as a page

How do Elasticsearch's standard, simple, whitespace, and keyword analyzers differ on the same input?

level: juniorimportance: must knowfreq 74%
basics
~20 s

The standard analyzer splits on Unicode word boundaries and lowercases; simple splits at any non-letter and lowercases, so digits disappear; whitespace splits only on spaces and keeps case and punctuation; keyword emits the whole input as one token.

open as a page

How is a custom analyzer composed in Elasticsearch, and in what order do its parts run?

level: juniorimportance: must knowfreq 78%
basics
~20 s

A custom analyzer is zero or more character filters, exactly one tokenizer, then zero or more token filters. Character filters rewrite the raw string, the tokenizer splits it into tokens, and token filters transform that token stream in listed order.

open as a page

When is a text field analyzed in Elasticsearch — at index time, at search time, or both?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Both. Elasticsearch analyzes a text field's value when the document is indexed and stores the resulting terms; it analyzes the query string of a full-text query at search time. A hit requires the two term sets to overlap.

open as a page

What tokens do the ngram and edge_ngram tokenizers emit, and how does that drive index size?

level: middleimportance: must knowfreq 64%
basics
~20 s

edge_ngram emits only prefixes anchored at the start, so Quick with grams 2 to 4 gives Qu, Qui, Quic. ngram emits substrings at every position, producing far more terms and a much larger index for infix matching.

open as a page

In an Elasticsearch bool query, what is the difference between a must clause and a filter clause?

level: juniorimportance: must knowfreq 85%
basics
~20 s

Both are mandatory and select exactly the same documents. A must clause runs in query context and adds a relevance contribution to _score; a filter clause runs in filter context, computes no score, and its result set can be reused from the node query cache.

open as a page

In Elasticsearch, what does a match query do with the query string before it looks up terms?

level: juniorimportance: must knowfreq 82%
basics
~10 s

A match query runs the string through the field's search analyzer, turning it into terms, then builds a boolean query over those terms. By default a document matching any single term is a hit.

open as a page

Why does Elasticsearch reject a search with from=10000 and size=10?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Because from + size exceeds index.max_result_window, which defaults to 10,000. Deep paging makes every shard build a top-(from+size) list that the coordinating node merges and then throws almost all of away, so Elasticsearch caps the depth instead.

open as a page

Why does an Elasticsearch term query on a text field usually return no hits?

level: juniorimportance: must knowfreq 88%
basics
~20 s

A term query looks up the exact bytes you supply with no analysis, while a text field stores analyzed tokens that are split and lowercased. Searching for New York finds no such token. Use match, or a keyword sub-field.

open as a page

When do should clauses in an Elasticsearch bool query affect which documents match?

level: middleimportance: must knowfreq 70%
basics
~20 s

Only when minimum_should_match requires it. If the bool query has at least one should clause and no must or filter clauses, minimum_should_match defaults to 1; otherwise it defaults to 0, so should clauses only add to _score.

open as a page

In Elasticsearch, what does a terms aggregation return and what does its size parameter control?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A terms aggregation groups matching documents by the distinct values of a field and returns one bucket per value with a doc_count. size caps how many buckets come back, defaulting to 10, ordered by descending count.

open as a page

Why can an Elasticsearch terms aggregation report wrong doc_count values on a multi-shard index?

level: middleimportance: must knowfreq 72%
basics
~20 s

Each shard independently returns only its own top candidate terms, and the coordinating node sums those partial lists. A term ranked low on one shard contributes nothing from it, so counts can be undercounted or the term missed entirely.

open as a page

Why is Elasticsearch's cardinality aggregation approximate, and what does precision_threshold trade off?

level: middleimportance: must knowfreq 70%
basics
~20 s

The cardinality aggregation estimates distinct counts with a HyperLogLog++ sketch of fixed size instead of tracking every value, so memory stays bounded as cardinality grows. precision_threshold trades memory for accuracy: higher means near-exact counts up to a larger unique count.

open as a page

Why can Elasticsearch's percentiles aggregation report a p99 that differs from the exact value?

level: middleimportance: must knowfreq 55%
basics
~20 s

The percentiles aggregation summarizes the value distribution in a bounded TDigest sketch rather than sorting every value, so results are approximate. TDigest is most accurate at extreme percentiles and least accurate near the median, and per-shard sketches are merged.

open as a page

Why does aggregating on an Elasticsearch text field require fielddata, and why is it off by default?

level: middleimportance: must knowfreq 72%
basics
~20 s

Analyzed text fields have no doc_values, so aggregating on them makes Elasticsearch uninvert the index into heap-resident fielddata, whose size grows with token count. That risks the heap, so it is disabled by default; aggregate on a keyword sub-field instead.

open as a page

In Elasticsearch, how do the ^ field-boost suffix and the boost parameter affect _score?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A boost is a multiplier on one query clause's score contribution: title^3 inside a multi_match triples what a title match adds, and any query's boost parameter does the same. Boosts are relative weights, and filter context ignores them.

open as a page

What must an Elasticsearch dense_vector mapping declare before a knn search can use it?

level: juniorimportance: must knowfreq 68%
basics
~10 s

The field must be typed dense_vector with a fixed dimension count, left indexed so an HNSW graph is built, and given a similarity metric (cosine, dot_product, l2_norm or max_inner_product) that matches the embedding model.

open as a page

How do you use Elasticsearch's _explain API to see why a document got the BM25 score it did?

level: middleimportance: must knowfreq 70%
basics
~20 s

Call GET /<index>/_explain/<id> with the query body, or add "explain": true to a search. Elasticsearch returns a nested tree breaking the score into per-clause weights with the idf and tf formulas and the actual N, n, freq, dl and avgdl values used.

open as a page

In Elasticsearch's function_score query, what is the difference between score_mode and boost_mode?

level: middleimportance: must knowfreq 74%
basics
~20 s

score_mode combines the results of the entries in the functions array with each other; boost_mode combines that single combined function score with the score of the inner query. Both default to multiply, and boost_mode replace discards the query score entirely.

open as a page

In an Elasticsearch knn search, how do k and num_candidates differ?

level: middleimportance: must knowfreq 76%
basics
~20 s

k is how many nearest neighbours the search returns; num_candidates is how many candidates each shard explores in the HNSW graph before picking its best k. Raising num_candidates buys recall at the cost of latency.

open as a page

What do green, yellow, and red cluster health mean in an Elasticsearch cluster?

level: juniorimportance: must knowfreq 80%
basics
~20 s

Green: every primary and replica shard is assigned. Yellow: all primaries are assigned but at least one replica is not. Red: at least one primary is unassigned, so part of the data is missing from search results and writes to that shard fail.

open as a page

Why is a document Elasticsearch just acknowledged as indexed sometimes missing from an immediate search?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Elasticsearch search is near-real-time, not real-time. Indexing puts the document in an in-memory buffer and the translog; it only becomes searchable when a refresh turns that buffer into a new searchable Lucene segment, by default about once a second.

open as a page

How does Elasticsearch decide which shard a document lands on when you index it without a routing value?

level: middleimportance: must knowfreq 70%
basics
~20 s

Elasticsearch hashes the document's routing value, which defaults to its _id, with Murmur3 and reduces that hash modulo the index's routing-shard count to pick exactly one primary shard. Placement is pure arithmetic, computed on any node, never looked up.

open as a page

What happens during the query phase and the fetch phase of an Elasticsearch search?

level: middleimportance: must knowfreq 65%
basics
~20 s

In the query phase each shard runs the search locally and returns only document ids plus sort values or scores; the coordinating node merges these into one globally sorted list. The fetch phase then retrieves the _source of just the winning documents.

open as a page

Why does an Elasticsearch cluster with thousands of tiny shards perform worse than one with fewer large shards?

level: middleimportance: must knowfreq 70%
basics
~20 s

Every shard is a separate Lucene index with fixed overhead in heap, file handles, merge work and cluster-state metadata, and every search becomes a task per shard that must be scheduled and merged. Those per-shard costs dominate once shards are small, which is why guidance targets tens of gigabytes per shard.

open as a page