Why can't you change an existing field's type in an Elasticsearch mapping?
answer
- Segments on disk are written once
- The encoding depends on the type at write time
- Additive changes only: new fields, sub-fields
- Sub-field added late is empty for old docs
- New index plus reindex is the real fix
basics
~20 sA field's type and analyzer are baked into immutable Lucene segments when each document is indexed, so already-written data cannot be reinterpreted under a new type. Adding new fields is fine; changing an existing one requires a new index and a reindex.
solid answer
~50 sElasticsearch mappings are **additive**. Documents are stored in immutable Lucene segments, and the way a value was tokenised, encoded into postings and written into doc values was decided by the field's type and analyzer at index time. Nothing on disk records the original input in a re-interpretable form, so flipping `price` from `text` to `float` would leave every existing document encoded the old way — Elasticsearch rejects the update rather than silently producing an index where half the documents are unqueryable. What you *can* do with `PUT /<index>/_mapping` is add a brand-new field, add a sub-field under `fields` on an existing field, and adjust a few parameters such as `ignore_above`, `dynamic`, `meta` and `search_analyzer`. Everything else — type changes, index-time analyzer changes, `object` to `nested` — means creating a new index with the corrected mapping and running `_reindex`.
code
json · 10 linesPUT /products/_mapping
{
"properties": {
"sku": { "type": "keyword" },
"title": {
"type": "text",
"fields": { "keyword": { "type": "keyword", "ignore_above": 256 } }
}
}
}go deeper
Recall that mappings are additive: you can add a field but not retype an existing one, and the way out is a new index plus reindex. Knowing that much already puts you ahead at this level.
Explain the Lucene reason — values were tokenised and encoded on write into immutable segments — and list what the update mapping API does accept: new fields, sub-fields, ignore_above, dynamic, search_analyzer.
Show you know the silent failure: a sub-field added to a populated index returns nothing for older documents until _update_by_query rewrites them, and that this is full-index I/O, not a metadata change.
Own the consequence upstream: mappings should be explicit and template-versioned so dynamic guesses never force an unplanned rebuild, and rebuild capacity should be budgeted as normal operating cost.
## The mechanical reason An Elasticsearch index is a set of Lucene segments, and **segments are write-once**. When a document is indexed, the field's mapping decides what actually gets written to disk: a `text` field is run through an analyzer and its tokens land in the inverted index; a `keyword` field is stored as a single undivided term plus a doc-values column; a `long` is encoded into a numeric point structure (a BKD tree) and a numeric doc-values column; a `date` becomes a number of milliseconds. Those encodings are not interchangeable and there is no faithful record of the original input to re-derive them from. `_source` holds the original JSON, but `_source` is a stored blob — it is not indexed, and Elasticsearch will not silently walk every segment re-deriving structures. So a mapping type change would mean: new documents encoded one way, all existing documents encoded another, and queries returning nonsense for half the index. Elasticsearch refuses the update instead, with an error saying the mapper cannot be changed from one type to the other. ## What you *can* change on a live index The update mapping API (`PUT /<index>/_mapping`) accepts genuinely additive changes: - **Adding a new field** to `properties`. Costs nothing for existing documents — they simply do not have it. - **Adding a sub-field** under an existing field's `fields` block, for example a `title.keyword` beside an analyzed `title`. - **`ignore_above`** on a keyword field. - **`dynamic`** (`true`/`false`/`strict`/`runtime`) at the object or index level. - **`meta`**, free-form metadata. - **`search_analyzer`** on a text field — legal precisely because it affects only query-time analysis, not what is already on disk. ## What you cannot change - The **type** of an existing field (`text` → `keyword`, `long` → `double`, `keyword` → `date`). - The **index-time `analyzer`** of an existing `text` field — the tokens are already written. - Turning `index`, `doc_values` or `norms` on or off after the fact. - Promoting an `object` field to `nested`, which is a type change and additionally changes how documents are physically stored (nested objects become hidden sub-documents). - The **number of primary shards**, which is not a mapping concern but shows up in the same conversation because it has the same remedy. ## The sub-field trap Adding a multi-field is legal, but it is easy to misread what that means. `PUT _mapping` updates the mapping only; it does not touch stored documents. Documents indexed **before** the change have no data under `title.keyword`, so a `terms` aggregation or a sort on it silently returns nothing for them — no error, just wrong results. New documents work, which makes the bug look intermittent. The fix is `_update_by_query` on the index: it rewrites each document in place through the current mapping, populating the new sub-field. It is a real reindex of the whole index in terms of I/O — do not think of it as a metadata touch-up. ## The remedy for a real type change The standard sequence: 1. Create a new index — `products-v2` — with the corrected mapping, ideally derived from a versioned index template so the mapping is code-reviewed. 2. `POST /_reindex` from `products-v1` into `products-v2`, using a `script` if values must be converted (parsing a numeric string, splitting a field). 3. Verify document counts and spot-check queries against the new index directly. 4. Move the alias from v1 to v2 in one atomic `_aliases` call. That is why aliases and mapping immutability are always taught together: immutability makes rebuilds routine, and the alias is what makes a rebuild invisible to clients. ## Escape hatches worth naming If you only need a differently-typed *view* of data you already have, a runtime field computed at query time can avoid the rebuild — at the cost of evaluating it per query per document, and with no help from the inverted index. It is a good stopgap while a reindex is scheduled, not a substitute for correct mapping. The deeper lesson interviewers are fishing for: because fixing a type is expensive, **mappings should be explicit from the start**. An index that let dynamic mapping guess `long` for what turns out to be a version string like `1.10` has bought itself a full rebuild.
- You add title.keyword to an index that already holds ten million documents. What must you do before sorting on it?Run `_update_by_query` on the index. The mapping update only affects documents indexed afterwards; existing documents have nothing stored under the new sub-field, so a sort or terms aggregation quietly omits them. `_update_by_query` rewrites every document in place through the current mapping. It is full-index I/O, so throttle it and expect segment churn until merges catch up.
- Why is search_analyzer updatable on an existing text field when analyzer is not?`search_analyzer` only affects how the query string is analyzed at request time; nothing already written to disk depends on it. The index-time `analyzer` determined which tokens were written into the inverted index for every existing document, and those tokens cannot be regenerated without rewriting the documents — so Elasticsearch rejects changing it in place.
- When is a runtime field a reasonable alternative to reindexing after a mapping mistake?When the correct value can be derived from `_source` or an existing field and query volume on it is low — for example exposing a numeric view of a field that was mapped as text. It costs per-document evaluation at query time and cannot use the inverted index, so it is a stopgap that buys scheduling room, not a permanent fix for a hot field.
Indexing a document is like printing it: the type decides the ink and layout. You can print an extra page, but you cannot re-typeset the pages already in the box — you reprint the book.
saying these in an interview costs you the question
- Believes PUT _mapping can convert a text field to keyword
- Thinks Elasticsearch re-encodes existing documents from _source automatically
- Adds a multi-field and assumes old documents gain values
- Confuses updating search_analyzer with changing index-time analysis
- Says the fix is deleting the field from the mapping and re-adding it