skip to content

What does it take to change the index-time analyzer of a field on an existing Elasticsearch index?

level: seniorimportance: should knowfreq 55%

answer

  1. Terms on disk are not recomputed
  2. One of the two analyzer parameters is updatable
  3. Analysis settings are a static index setting
  4. Build new, copy, then point the name at it
  5. An alias makes the cutover invisible

basics

~20 s

You cannot update a field's analyzer parameter in place, and existing documents are never re-analyzed. Create a new index with the new analysis settings and mapping, reindex into it, and swap an alias — or add the analyzer to a closed index and reindex anyway.

solid answer

~50 s

Two separate immutability rules bite here. First, analyzer *definitions* live in `index.analysis`, a static index setting, so adding or changing one on an existing index means closing it, updating settings and reopening it. Second, the `analyzer` mapping parameter on an existing field cannot be updated at all — only `search_analyzer` can, via the update mapping API. Even if you get past both, the decisive fact is that terms already written into segments are never recomputed: old documents keep their old terms and new ones get the new ones, leaving an index that matches inconsistently. The clean procedure is therefore to create a new index carrying the new `analysis` settings and mapping, run `_reindex` from the old one, verify with `_analyze` and a query set, then point the alias at the new index and delete the old.

code

bash · 11 lines
bash
PUT /articles-v3
{ "settings": { "analysis": { "...": "new analyzers" } },
  "mappings": { "properties": { "body": { "type": "text", "analyzer": "body_v3" } } } }

POST /_reindex?wait_for_completion=false
{ "source": { "index": "articles-v2" }, "dest": { "index": "articles-v3" } }

POST /_aliases
{ "actions": [
  { "remove": { "index": "articles-v2", "alias": "articles" } },
  { "add":    { "index": "articles-v3", "alias": "articles" } } ] }

go deeper

for a junior

Remember the headline rule: existing documents keep the terms they were indexed with, so an index-time analyzer change means the data has to be rebuilt.

for a middle

Distinguish the two immutability rules — static analysis settings needing a closed index, and the analyzer mapping parameter being unchangeable — and describe the reindex path.

for a senior

Walk the migration end to end: new index, async reindex, verification against a query set, atomic alias swap, and a plan for writes arriving mid-copy.

for a principal

Speak to the programme: making alias-based indexing the default so analysis changes are routine, and budgeting reindex capacity as a normal cost of relevance work.

## Three facts that decide the procedure **1. Analysis settings are static.** Custom analyzers, tokenizers and token filters are declared under `index.analysis` in the index settings. Static settings can only be changed while the index is closed: `POST /idx/_close`, `PUT /idx/_settings` with the new analysis block, `POST /idx/_open`. Closing an index makes it unavailable for reads and writes for the duration, which on a production cluster is already a reason to prefer building a new index instead. **2. The analyzer mapping parameter is immutable on an existing field.** The update mapping API accepts new fields and a small set of parameter changes; the field's `analyzer` is not among them, because the terms it produced are already on disk. `search_analyzer` *is* updatable, precisely because it only affects query-time tokenization. So a change confined to the query side is easy and immediate; a change to the index side is not a mapping edit at all. **3. Existing documents are never re-analyzed.** This is the one candidates most often miss. Nothing in Elasticsearch walks the corpus and recomputes terms — not a close/open cycle, not a `_refresh`, not a `_flush`, not a force merge. Segments hold the terms produced when each document was indexed. If you somehow applied a new index-time analyzer to a live index, documents written before the change and after the change would tokenize differently, and searches would return an arbitrary mixture depending on which half of the corpus a document fell into. That silent inconsistency is worse than an outright error. ## The safe procedure The accepted pattern is build-new-and-swap: 1. **Create a new index** with the new `analysis` settings and the new mapping. Give it a versioned physical name (`articles-v3`) rather than the name your application uses. 2. **Reindex** with the `_reindex` API from the old index into the new one. For a large corpus, run it asynchronously (`wait_for_completion=false`) and watch it through the tasks API; slicing parallelizes it. Consider raising `refresh_interval` and dropping replicas on the target while it loads, then restoring both. 3. **Verify before cutting over.** Run `_analyze` against the new field to confirm the tokens are what you intended, compare document counts, and run a fixed set of queries against both indices to see how results shift — an analysis change is a relevance change, and it deserves the same scrutiny as one. 4. **Swap the alias atomically.** The application should already read and write through an alias; a single `_aliases` call that removes the old index and adds the new one makes the cutover invisible to clients. 5. **Handle writes during the reindex.** `_reindex` copies a point-in-time snapshot, so anything written afterwards must be replayed — dual-write to both indices during the migration, re-run the reindex for a bounded time window using a range query on an updated-at timestamp, or take a short write pause. ## When you can avoid reindexing Before committing to a reindex, check whether the change really belongs on the index side: - A change to synonyms, stopwords or query-side normalization can often live in the `search_analyzer` instead, which is an update-mapping call and takes effect immediately for old and new documents alike. - A search-time synonym filter marked `updateable: true` can have its rules refreshed on a live index with `POST /<index>/_reload_search_analyzers` — no close, no reindex. - Adding a *new* field or multi-field with the new analyzer is a legal mapping update. It does not populate the new field for existing documents on its own, but combined with an update-by-query it can be a cheaper migration path than rebuilding the whole index when only one field changes. ## Why the design is like this The rule is not arbitrary. An inverted index is a precomputed structure: terms, posting lists, positions, and the per-field statistics used for scoring were all derived from a particular analysis chain. Retroactively changing the chain would invalidate all of it. Systems that let you "change the analyzer" transparently are simply doing this same rebuild under the covers. Elasticsearch makes the cost visible so you can schedule it, run it against a copy, and validate relevance before your users see the difference. ## The interview answer in one line Say it as a decision tree: if the change can be expressed on the query side, update `search_analyzer` and you are done; if it must change stored terms, build a new index, reindex, verify, and flip the alias — and never expect existing documents to re-tokenize themselves.

  • Why does adding a new custom analyzer require closing the index?
    Analyzer definitions live under `index.analysis`, which is a static index setting — Elasticsearch only accepts changes to static settings while the index is closed, so the sequence is close, `PUT _settings`, open. On production this downtime is usually reason enough to build a new index instead, where the settings are supplied at creation time.
  • How do you keep the index consistent for writes that arrive during the reindex?
    `_reindex` copies a snapshot, so later writes are missed. Options are dual-writing to both indices for the migration window, replaying a bounded catch-up reindex filtered on an updated-at timestamp, or pausing writes briefly at the end. Whichever you pick, cut over with an atomic `_aliases` update so clients never see a half-migrated index.
  • Which analyzer-related change does NOT require a reindex?
    Changing `search_analyzer` on an existing field via the update mapping API, and refreshing the rules of a search-time filter marked `updateable: true` with `_reload_search_analyzers`. Both affect only how query input is tokenized, so they apply immediately to documents indexed at any time.
  • How do you validate that the analysis change actually improved things?
    Treat it as a relevance change: run `_analyze` on the new field to confirm the tokens, then run a fixed query set against old and new indices and compare the top results against judgements before flipping the alias. Document-count parity confirms the copy; only a query comparison confirms the change was an improvement.

saying these in an interview costs you the question

  • Believes a close/open cycle re-analyzes existing documents
  • Tries to update the analyzer parameter with the update mapping API
  • Forgets writes arriving during the reindex
  • Cuts over by renaming rather than using an alias
  • Thinks a force merge or refresh rebuilds terms

context