skip to content

How does Elasticsearch decide which analyzer applies to a text field at index time?

level: middleimportance: should knowfreq 46%

answer

  1. Most specific declaration wins
  2. The fallback lives in settings, not in mappings
  3. It is a specially-named analyzer
  4. Ultimate fallback is one built-in analyzer
  5. Some field types ignore all of this

basics

~10 s

Elasticsearch uses the field mapping's analyzer parameter first; if the field does not declare one it falls back to the index's analysis.analyzer.default setting, and if that is not defined it uses the standard analyzer.

solid answer

~50 s

Resolution is a three-step fallback. The most specific wins: an `analyzer` declared on the field in the mapping. If the field declares none, Elasticsearch looks for an analyzer named `default` in the index's `analysis.analyzer` settings — that is how you change an index-wide default, not by any top-level `index.analyzer` key. If neither exists, the `standard` analyzer is used. A `default_search` analyzer can be defined the same way to change query-time behaviour. Two limits matter in practice: this only concerns analyzed field types, since `keyword` fields skip analysis and take a `normalizer` instead; and changing the resolution changes only what is indexed *from now on* — documents already written keep their existing terms until they are reindexed. To check the effective answer for a field, run `_analyze` against the index with the `field` parameter rather than reading the mapping.

code

json · 14 lines
json
PUT /articles
{
  "settings": {
    "analysis": { "analyzer": { "default": { "type": "english" } } }
  },
  "mappings": {
    "properties": {
      "body":  { "type": "text" },
      "sku":   { "type": "text", "analyzer": "whitespace" },
      "state": { "type": "keyword" }
    }
  }
}
// body -> english (index default), sku -> whitespace, state -> not analyzed

go deeper

for a junior

Know that a field can name its own analyzer and that Elasticsearch falls back to the standard analyzer when nothing is specified. Being able to say which one wins when both exist is the expected answer.

for a middle

Explain the full three-step resolution and where the index default actually lives — a named default analyzer in analysis settings, not a mapping section — and note that keyword fields are outside this mechanism entirely.

for a senior

Demonstrate the operational consequences: analysis changes are not retroactive, the index ends up mixed until reindexed, and verification means running _analyze with the field parameter rather than trusting the mapping.

for a principal

Own the convention: whether index defaults or explicit per-field analyzers are the house style, how that is enforced through index templates, and what the estate-wide reindex cost is if the default turns out to be wrong.

## Three levels, most specific first When Elasticsearch indexes a value into a `text` field, it needs one analyzer. It finds it in this order: 1. **The field's own `analyzer` in the mapping.** `"title": { "type": "text", "analyzer": "english" }` is unambiguous and always wins. 2. **The index's default analyzer**, defined by creating an analyzer literally named `default` under `settings.analysis.analyzer`. Every analyzed field that does not declare its own analyzer picks this up. 3. **`standard`**, the built-in fallback when neither of the above exists. The second step is the one people get wrong. There is no `index.analyzer` setting and no `_default_` mapping section in current versions — the mechanism is a *named* analyzer: ```json PUT /articles { "settings": { "analysis": { "analyzer": { "default": { "type": "english" } } } } } ``` Every text field in `articles` now analyzes with `english` unless it says otherwise. Adding `"sku": { "type": "text", "analyzer": "whitespace" }` overrides it for that one field. ## The search-time counterpart The same idea applies to query-time analysis: a field-level `search_analyzer` beats an index-level analyzer named `default_search`, which beats the field's index-time analyzer. Elasticsearch's default is symmetry — the same analyzer on both sides — because asymmetry silently breaks matching unless it is deliberate. ## Which fields this affects Only analyzed types. A `keyword` field is not analyzed at all: it is indexed verbatim and takes a `normalizer` rather than an analyzer, so an index default analyzer has no effect on it whatsoever. The same is true of numerics, dates, booleans and IPs. This matters because a common bug report — "I set an index default analyzer and my `status` field still matches case-sensitively" — is explained entirely by the field being a keyword. Multi-fields resolve independently: a `text` field with a `.keyword` sub-field has an analyzer on the parent and no analysis on the sub-field, and each is queried by its own name. ## Dynamic mapping interacts with this If you index a document into an index with no mapping for that field, dynamic mapping creates one — typically a `text` field with a `keyword` sub-field for a string. The new `text` field declares no analyzer, so it inherits the index default, and thereafter behaves like every other field. That is convenient, and it is also why an index default is a heavier decision than it looks: it silently governs every field nobody explicitly configured. ## Changes are not retroactive Analysis happens at write time. Once a document is indexed, its terms are fixed in immutable segments. Changing the mapping's analyzer or the index default therefore affects only documents indexed afterwards, leaving the index in a mixed state where old and new documents behave differently for the same query. The remedy is to rebuild the data into a new index and switch traffic over. Analysis settings are also static rather than freely updatable on a live index, which reinforces the same workflow: create a new index with the settings you want and move the data. ## Verifying rather than assuming Reading a mapping tells you what was declared, not what is *effective* — the field may be silently inheriting a default defined in the settings, or the fallback to `standard`. The reliable check is: ```json POST /articles/_analyze { "field": "title", "text": "The Quick Brown Foxes" } ``` That resolves the analyzer exactly as indexing would and prints the tokens, so it answers the question directly. Comparing the same call on two indices is also the fastest way to explain why a query behaves differently across them. ## Practical guidance Setting an index default is appropriate when an entire index has one nature — a single-language document corpus, for example. When an index mixes prose, identifiers and structured strings, per-field analyzers are clearer, because the mapping then documents itself and nobody has to remember that a setting elsewhere is quietly in force. Whichever you choose, express it once in an index template so every index created from it is consistent rather than depending on whoever issued the create call.

  • Does an index default analyzer affect keyword fields?
    No. Keyword fields are not analyzed at all — they are indexed verbatim and accept a `normalizer` instead. An index-wide default analyzer applies only to analyzed types such as `text`, so a case-sensitivity complaint about a keyword field is never fixed by changing the default.
  • How do you confirm which analyzer a field is actually using?
    Run `POST /<index>/_analyze` with the `field` parameter and some sample text. That performs the same resolution indexing would — field mapping, then the index `default`, then `standard` — and prints the resulting tokens. Reading the mapping alone is unreliable, because a field inheriting a default declares no analyzer at all.
  • You change an index's default analyzer — what happens to documents already indexed?
    Nothing. Their terms were written into immutable segments at index time and stay as they were, so the index ends up mixed: old documents answer queries under the old analysis, new ones under the new. The fix is to rebuild the data into a new index with the desired settings and cut traffic over, typically behind an alias.

saying these in an interview costs you the question

  • Thinks the index default analyzer also applies to keyword fields
  • Believes changing the default re-analyzes existing documents
  • Looks for a _default_ section in the mappings
  • Assumes standard cannot be overridden per index
  • Reads the mapping instead of testing with _analyze

context