skip to content

Index-time vs Search-time Analysis

The analyzer that built the index and the analyzer applied to the query must produce compatible terms, and changing the index-time one means reindexing everything. Interviewers use edge n-gram autocomplete as the canonical trap: n-grams at index time, plain analysis at search time.

part ofElasticsearchoverview, primer and where to startread it →
on this pageshow

questions

6

When is a text field analyzed in Elasticsearch — at index time, at search time, or both?

level: juniorimportance: must knowfreq 62%

answer

  1. The raw string is not what gets matched
  2. Two moments in a document's life
  3. One side is baked in, one runs per query
  4. Both document text and query text are tokenized
  5. term queries skip one of the two

basics

~20 s

Both. Elasticsearch analyzes a text field's value when the document is indexed and stores the resulting terms; it analyzes the query string of a full-text query at search time. A hit requires the two term sets to overlap.

solid answer

~40 s

A `text` field is analyzed twice, at two different moments. At **index time** the field value runs through the field's analyzer — character filters, tokenizer, token filters — and the resulting tokens are what actually land in the inverted index; the original string is only kept in `_source`. At **search time**, full-text queries such as `match` or `match_phrase` run the query string through an analyzer for that field as well, and the produced terms are looked up in the index. Matching therefore never compares raw strings: it compares terms to terms. This is why the two analyzers have to be compatible — if index time lowercases and search time does not, `"Quick"` will never find the indexed term `quick`. Term-level queries such as `term` skip search-time analysis entirely and look the string up as-is.

code

bash · 7 lines
bash
# terms the field actually holds (index-time analyzer)
GET /articles/_analyze
{ "field": "title", "text": "The Quick Brown Foxes" }

# terms a query would produce with a different analyzer
GET /articles/_analyze
{ "analyzer": "whitespace", "text": "Quick Foxes" }

go deeper

for a junior

Be ready to say plainly that both the document value and the full-text query string are run through an analyzer, and that matching compares the resulting terms rather than the original text.

for a middle

Explain the pipeline stages (character filters, tokenizer, token filters) and show with the _analyze API why a lowercasing mismatch between the two sides produces zero hits.

for a senior

Demonstrate that you use the asymmetry deliberately: index-time analysis is frozen into segments and needs a reindex to change, while the query side can be swapped per request.

for a principal

Own the policy: which analysis decisions get baked into the index and which stay changeable at query time determines how expensive future relevance iterations are.

## The two moments Elasticsearch never matches raw strings. It matches **terms**, and terms are produced by an analyzer. Analysis happens at two distinct points in a document's life, and understanding which one you are looking at explains most "why does this query return nothing" bugs. **Index time.** When a document is indexed, every `text` field is passed through that field's analyzer. The analyzer is a three-stage pipeline: zero or more character filters (which rewrite the raw character stream, e.g. stripping HTML), exactly one tokenizer (which splits the stream into tokens and records their positions and offsets), and zero or more token filters (lowercasing, stemming, stopword removal, synonym expansion, ASCII folding). The tokens that come out the far end are written into the inverted index as terms, together with document ids, positions and frequencies. The original text is *not* stored in the index structure used for matching — it survives only in `_source` (for returning the document) and, if you enabled them, in `doc_values` or `stored` fields. **Search time.** When a full-text query such as `match`, `match_phrase`, `multi_match` or `query_string` targets that field, the query's input string is analyzed too, using the field's search analyzer. Each produced term becomes a lookup into the inverted index. `match` combines those lookups with a boolean OR by default; `match_phrase` additionally requires the terms to be adjacent according to the positions recorded at index time. ## Why compatibility matters Because matching happens purely between two sets of terms, the pipelines on the two sides must agree on the *shape* of a term. If the index-time analyzer lowercases and folds accents while the search-time one does not, the query term `Café` never finds the indexed term `cafe`. If the index-time analyzer stems (`running` → `run`) and the search side does not, searching for `running` misses documents that only ever produced `run`. The simplest way to guarantee agreement is what Elasticsearch does by default: unless you say otherwise, the same analyzer is used on both sides. You override the query side only with a deliberate purpose — the classic one being edge n-gram autocomplete, where you want many prefix terms in the index but only the whole typed word as the query term. ## Term-level queries are the exception `term`, `terms`, `prefix`, `wildcard` and `range` do **not** analyze their input. They take the string you supply and look it up verbatim. This is precisely correct for `keyword` fields, which are also not analyzed at index time (the whole value becomes one term), and it is the number-one beginner trap on `text` fields: `{"term": {"title": "Quick Brown"}}` on a standard-analyzed field finds nothing, because the index contains `quick` and `brown`, never `Quick Brown`. ## Seeing it for yourself The `_analyze` API is the diagnostic tool. Called with a `field`, it shows how that field's index-time analyzer tokenizes the text; called with an explicit `analyzer` name, it shows any pipeline you like. Comparing the two token lists side by side — the one your document produced and the one your query produces — resolves nearly every empty-result mystery in a minute. ## Practical consequences Two consequences follow, and interviewers usually push toward them: 1. **The index side is frozen into the data.** Terms were computed once, when the document was indexed. Changing the index-time analyzer does not rewrite existing segments; only newly indexed documents get the new terms. Making an analyzer change apply to old data means reindexing them. 2. **The search side is cheap to change.** The query analyzer runs per request, so swapping it — or reloading its synonym rules — takes effect immediately for all documents, old and new, with no reindex. That asymmetry is the whole reason the two sides can be configured separately, and it is why decisions like "expand synonyms at index time or at query time" have very different operational costs even though they can produce identical result sets.

  • If both sides are analyzed, why does a term query on a text field so often return nothing?
    Term-level queries do not analyze their input. The string is looked up verbatim, while the field's index-time analyzer stored lowercased, tokenized, possibly stemmed terms. `{"term": {"title": "Quick Brown"}}` looks for the single term `Quick Brown`, which the standard analyzer never produced. Use a `match` query, or target a `keyword` sub-field where the whole value is one untouched term.
  • How would you prove that a mismatch is an analysis problem and not a query problem?
    Run the `_analyze` API twice: once with the `field` parameter and the document's text, to see the terms actually in the index, and once with the analyzer used at query time and the user's input. If the two token lists share no term, the query cannot match, and the fix belongs in the mapping, not in the query DSL.
  • Does the analyzer affect what is returned in _source?
    No. `_source` stores the original JSON exactly as you sent it, so a hit always returns the untouched text. Analysis only shapes the terms in the inverted index used for matching, scoring and highlighting.

saying these in an interview costs you the question

  • Says Elasticsearch matches the query against the raw stored string
  • Thinks only documents are analyzed and queries are not
  • Believes term queries are analyzed like match queries
  • Assumes changing an analyzer retroactively re-tokenizes existing documents
  • Confuses _source contents with what is searchable

context

open as a page

Why should an edge_ngram autocomplete field in Elasticsearch use a different analyzer at search time?

level: middleimportance: must knowfreq 60%

basics

~20 s

Because n-gramming the query too turns each typed word into many short prefixes, which match almost everything and destroy precision and scoring. Prefixes belong in the index; the query should contribute the whole typed token as one term.

open as a page

In an Elasticsearch mapping, what is the difference between the analyzer and search_analyzer parameters?

level: middleimportance: should knowfreq 70%

basics

~20 s

The analyzer parameter defines the pipeline that turns a field's value into indexed terms, and is also used on queries unless overridden. search_analyzer overrides only the query side, so the terms a search produces can differ from the terms stored.

open as a page

What does it take to change the index-time analyzer of a field on an existing Elasticsearch index?

level: seniorimportance: should knowfreq 55%

basics

~20 s

You cannot update a field's analyzer parameter in place, and existing documents are never re-analyzed. Create a new index with the new analysis settings and mapping, reindex into it, and swap an alias — or add the analyzer to a closed index and reindex anyway.

open as a page

Should synonym expansion in Elasticsearch run at index time or at search time, and why?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Search time is the usual choice: rules can change without reindexing, the index stays clean, and synonym_graph handles multi-word synonyms correctly. Index-time expansion avoids per-query cost but bakes rules into segments and distorts term statistics.

open as a page

What does the search_quote_analyzer mapping parameter do in Elasticsearch?

level: seniorimportance: nice to knowfreq 20%

basics

~10 s

It sets a separate analyzer for text inside quotes in query_string and simple_query_string queries. The usual purpose is to keep stopwords in quoted phrases while the ordinary search analyzer strips them.

open as a page