Why is a string field in Elasticsearch usually mapped as text with a keyword sub-field?
answer
- one source value, several index entries
- the fields mapping parameter
- why Kibana keeps saying .keyword
- two types cannot live in one mapping
- analyzed for match, verbatim for aggregations
basics
~20 sA multi-field indexes the same source value more than once under different mappings. The text half serves match queries, while title.keyword keeps the value verbatim for exact filters, sorting and aggregations — two capabilities one mapping cannot provide.
solid answer
~40 sThe `fields` mapping parameter lets one JSON value be indexed several times under different types or analyzers. The standard pairing is a `text` parent for full-text `match` queries plus a `keyword` sub-field, addressed as `title.keyword`, that holds the value verbatim with `doc_values` so you can filter exactly, sort, and run `terms` aggregations. Dynamic mapping creates exactly this shape for any new string, with `ignore_above: 256` on the `keyword` half. Multi-fields are not limited to that pair: a second `text` sub-field with the `english` analyzer gives you stemmed recall alongside the exact-token parent. The cost is real — each sub-field is a separate set of index structures on disk — so on fields you truly never filter or aggregate on, drop the sub-field explicitly rather than inheriting it.
code
json · 11 lines{
"properties": {
"title": {
"type": "text",
"fields": {
"keyword": { "type": "keyword", "ignore_above": 256 },
"english": { "type": "text", "analyzer": "english" }
}
}
}
}go deeper
Know that title.keyword is a sub-field of title holding the raw value, and that filters, sorting and aggregations must target it while match queries target the analyzed parent.
Explain the fields mapping parameter: one source value indexed under several mappings, the dynamic default of text plus keyword with ignore_above, and the extra disk and indexing cost each sub-field carries.
Demonstrate the operational edge: a mapping update adds the sub-field but backfills nothing, so existing documents need update_by_query or a reindex before facets are trustworthy. Justify dropping inherited sub-fields on large text fields.
Own the mapping policy across teams: explicit mappings instead of dynamic multi-fields on wide documents, a convention for analyzer sub-fields used in relevance work, and the storage bill that doubling every string field quietly creates.
## What a multi-field actually is A multi-field is a field mapping that declares extra sub-mappings under the `fields` parameter. When a document is indexed, Elasticsearch takes the single JSON value and feeds it to the parent mapping *and* to every sub-field mapping, writing several independent sets of index structures from one source value. The document itself is unchanged; `_source` still contains one `title`. Only the index gains extra entries. You address a sub-field by dotted path: if `title` is `text` with a sub-field named `keyword`, then `title` is the analyzed field and `title.keyword` is the verbatim one. The sub-field name is arbitrary — `keyword` is only a convention that dynamic mapping happens to use. ## Why the text-plus-keyword pair is the default The two string types answer different questions and neither can answer both. An analyzed `text` field supports `match`, phrase queries and relevance ranking, but stores no whole-value term and no `doc_values`, so it cannot be filtered exactly, sorted, or aggregated. A `keyword` field can do all of those and none of the full-text ones. When Elasticsearch meets a string with no explicit mapping, it refuses to guess and gives you both: ```json "title": { "type": "text", "fields": { "keyword": { "type": "keyword", "ignore_above": 256 } } } ``` That is why so many real queries and Kibana visualizations refer to `something.keyword`: the analyzed parent cannot serve a bucket aggregation, and the sub-field can. ## Beyond the default pair Multi-fields are a general mechanism for indexing the same text several ways: - **Different analyzers.** A parent `text` field with the `standard` analyzer plus a `title.english` sub-field using the `english` analyzer lets you query both and combine them: the stemmed sub-field boosts recall ("running" finds "run"), the exact parent boosts precision, and a `multi_match` across both usually ranks better than either alone. - **A different type.** A numeric-looking string can be indexed both as `keyword` and as a numeric sub-field if some queries need ranges and others need exact term lookups. - **Search-as-you-type or n-gram sub-fields.** Prefix matching gets its own analysis chain without polluting the main field. ## What it costs Each sub-field is a full set of index structures: its own postings, its own `doc_values` where applicable, its own norms. A `keyword` sub-field on a long free-text field is not free — you are storing the entire value again as a single term (up to `ignore_above`) plus doc values for it. On a large index of article bodies, an inherited `body.keyword` can be pure waste: nobody aggregates on article bodies, and any body over the character limit is silently absent from that sub-field anyway. Each sub-field also counts toward the index's field count, so dynamically mapping a wide, unpredictable document doubles your field growth. ## The operational gotcha Adding a new sub-field to an existing mapping is allowed — sub-fields are additive, and unlike a type change the request succeeds. But it only affects documents indexed *after* the change. Existing documents have no data in the new sub-field, so filters and aggregations on it silently under-report. To populate it you must rewrite the existing documents, typically with `update_by_query` on the same index or a reindex into a new one. Interviewers like this question because "the mapping update succeeded" and "the data is there" are two different things. The mirror-image gotcha is querying the wrong half. `{"term": {"title": "Wireless Mouse"}}` returns nothing because `title` is analyzed; `{"term": {"title.keyword": "Wireless Mouse"}}` works. A sort on `title` errors or is meaningless; a sort on `title.keyword` is correct. Half of "Elasticsearch returns no results" tickets are a missing `.keyword` suffix. ## How to decide, per field Start from usage, not from the default: - Searched full-text only (article body, description): plain `text`, drop the sub-field. - Filtered, sorted, faceted only (status, SKU, tenant ID, host name): plain `keyword`, no analyzed parent. - Genuinely both (product title, person name, tag that users also type into a search box): `text` with an explicit `keyword` sub-field, and set `ignore_above` deliberately. Writing the mapping explicitly costs a few minutes and removes a whole class of surprises: no silent 256-character truncation you did not choose, no doubled index for fields nobody facets on, and no `.keyword` archaeology when someone reads the queries a year later.
- You add a keyword sub-field to a mapping on a live index. Why do aggregations on it still come back nearly empty?The mapping update applies to future indexing only. Documents already in the index were written before the sub-field existed, so they contribute nothing to it. You have to rewrite them — update_by_query over the same index, or a reindex into a fresh one — before the sub-field reflects the full corpus.
- Does a multi-field increase the size of _source?No. _source stores the original JSON document exactly as it was sent, and a multi-field adds no keys to it. What grows is the index: each sub-field writes its own postings, doc values and norms, so disk usage and indexing time rise even though the stored document is unchanged.
- When would you add a second text sub-field with a different analyzer?When exact and stemmed matching pull in opposite directions. Keep the parent on a light analyzer for precision, add an english-analyzed sub-field for recall, then query both in a multi_match and boost the exact field higher. Documents matching the precise form rank above those matching only the stem.
saying these in an interview costs you the question
- Thinks a multi-field duplicates the value inside _source
- Assumes adding a sub-field backfills existing documents
- Says you can sort on the analyzed parent field
- Believes every string needs a keyword sub-field
- Cannot say what the .keyword suffix in a query refers to