skip to content

In an Elasticsearch mapping, what is the difference between the analyzer and search_analyzer parameters?

level: middleimportance: should knowfreq 70%

answer

  1. One shapes stored terms, one shapes query terms
  2. Default is to use the same on both sides
  3. Autocomplete is the reason asymmetry exists
  4. Only one of the two is updatable on a live field
  5. Query-body analyzer beats the mapping

basics

~20 s

The analyzer parameter defines the pipeline that turns a field's value into indexed terms, and is also used on queries unless overridden. search_analyzer overrides only the query side, so the terms a search produces can differ from the terms stored.

solid answer

~50 s

On a `text` field, `analyzer` is the index-time pipeline: it decides which terms are written into the inverted index. If nothing else is set, the same analyzer is also applied to full-text query input, which is the safe default because both sides then produce identically shaped terms. `search_analyzer` overrides just the query side. You set it when you deliberately want asymmetry — the canonical case is autocomplete, where the index side emits every prefix via `edge_ngram` while the search side emits only the whole typed token. Resolution order at query time is: an `analyzer` given in the query itself, then the field's `search_analyzer`, then the index-level `analysis.analyzer.default_search`, then the field's `analyzer`. Operationally they differ too: `search_analyzer` can be changed on an existing field with the update mapping API, while `analyzer` cannot — changing it means reindexing.

code

json · 11 lines
json
{
  "mappings": {
    "properties": {
      "title": {
        "type": "text",
        "analyzer": "autocomplete_index",
        "search_analyzer": "standard"
      }
    }
  }
}

go deeper

for a junior

Know that the mapping can name two analyzers for one field, and that only the first one decides what is actually stored in the index.

for a middle

Explain the query-time resolution order and give the autocomplete example as the reason the two parameters exist at all.

for a senior

Show the operational split: search_analyzer is an update-mapping call that takes effect at once, while an index-analyzer change is a reindex project.

for a principal

Frame it as where relevance iteration should live — keeping changeable logic on the query side is what lets a team tune search weekly without rebuilding indices.

## What each parameter controls Both parameters live on a `text` field in an Elasticsearch mapping, and both name an analyzer — a chain of character filters, one tokenizer and token filters. They differ in *when* that chain runs. `analyzer` runs at **index time**. Its output is the set of terms persisted in the inverted index for that field, together with positions and frequencies. Everything downstream — matching, phrase proximity, scoring statistics, highlighting — is computed from those terms. `search_analyzer` runs at **search time**, on the input of full-text queries targeting the field (`match`, `match_phrase`, `multi_match`, the free-text parts of `query_string`). It produces the terms that will be looked up. It has no effect whatsoever on what is stored. If you set only `analyzer`, Elasticsearch uses it on both sides. That default is deliberate and usually correct: symmetric analysis guarantees that a document's own text, fed back as a query, finds the document. ## Resolution order at query time When a full-text query hits a field, Elasticsearch picks the query-side analyzer in this order: 1. an `analyzer` explicitly named inside the query clause, 2. the field's `search_analyzer`, 3. the index-level default named `analysis.analyzer.default_search`, if defined, 4. the field's `analyzer`, 5. the index-level `default` analyzer, otherwise `standard`. Knowing this order matters when debugging: a `search_analyzer` on the field silently loses to an `analyzer` written into the query body, which is a common source of "the mapping says one thing, the results say another". ## When asymmetry is right — and when it is a bug Deliberate asymmetry has a handful of legitimate uses: - **Autocomplete.** The index side uses `edge_ngram` so `"star"` produces `s`, `st`, `sta`, `star`; the search side uses a plain analyzer so the typed prefix `"sta"` becomes exactly one term and matches the stored prefix. Applying n-gramming to the query too would flood it with tiny terms and wreck precision. - **Synonyms.** Keeping the index clean and expanding synonyms only in the search analyzer means rule changes need no reindex. - **Stopwords in phrases.** Index everything, but let the ordinary search analyzer discard stopwords while quoted phrases go through a stopword-preserving pipeline. Accidental asymmetry, on the other hand, is a bug: an index-time chain that stems or folds while the search chain does not means whole classes of queries silently return nothing. The rule of thumb is that the search analyzer should produce terms *in the same normalized form* as the index analyzer — the same casing, the same folding, the same stemming — and may legitimately produce *fewer or coarser* tokens, but never differently shaped ones. ## Mutability: the operational difference This is the part interviewers push on. `search_analyzer` is a query-side concern, so Elasticsearch lets you update it on an existing field with the update mapping API; the next search simply uses the new pipeline, for old and new documents alike. `analyzer` is baked into the terms already written to disk, so it cannot be updated on an existing field — the mapping update is rejected, and applying a new index-time analysis means creating a new index with the new mapping and reindexing into it. A related mechanic: analyzer *definitions* themselves live in `index.analysis`, which is a static index setting. Adding a brand-new custom analyzer to an existing index therefore requires closing the index, updating the settings, and reopening it — even if you only intend to use the new analyzer as a `search_analyzer`. The exception is a search-time synonym filter marked `updateable: true`, whose rules can be refreshed on a live index with the `_reload_search_analyzers` API. ## Debugging checklist When results look wrong, dump both sides. `GET /index/_analyze` with `"field": "title"` shows the index-time tokens for a sample text; calling it with `"analyzer": "<your search analyzer>"` shows what the query becomes. If those token lists cannot intersect, no query rewrite will save you — the fix is in the mapping. The `_validate/query?rewrite=true` and `explain` APIs then confirm which terms the query actually built.

  • Which of the two can you change on an existing field, and why only that one?
    `search_analyzer`, via the update mapping API, because it only affects how queries are tokenized at request time — nothing on disk depends on it. `analyzer` determined the terms already written into segments, so changing it would leave old documents indexed one way and new ones another; Elasticsearch rejects the update and expects a reindex into a new index instead.
  • If a query names its own analyzer, does the field's search_analyzer still apply?
    No. An `analyzer` given inside the query clause wins over the field's `search_analyzer`, which in turn wins over the index-level `default_search`, which wins over the field's `analyzer`. This precedence is a frequent debugging surprise: the mapping looks correct while a client library quietly injects an analyzer into the query body.
  • What is a safe rule for choosing a search analyzer that differs from the index analyzer?
    Keep the normalization identical — same lowercasing, folding and stemming — and let the search side differ only by producing fewer or coarser tokens than the index side. Index-time n-grams with whole-token search is safe; index-time stemming with unstemmed search is not, because the term forms no longer line up.

saying these in an interview costs you the question

  • Thinks search_analyzer changes what is stored in the index
  • Sets a different search analyzer with no shared normalization
  • Believes the index-time analyzer can be updated in place
  • Assumes the field's search_analyzer always wins over the query's
  • Cannot name a real reason to make the two sides differ

context