skip to content

Why is a string field in Elasticsearch usually mapped as text with a keyword sub-field?

level: middleimportance: must knowfreq 70%

answer

  1. one source value, several index entries
  2. the fields mapping parameter
  3. why Kibana keeps saying .keyword
  4. two types cannot live in one mapping
  5. analyzed for match, verbatim for aggregations

basics

~20 s

A multi-field indexes the same source value more than once under different mappings. The text half serves match queries, while title.keyword keeps the value verbatim for exact filters, sorting and aggregations — two capabilities one mapping cannot provide.

solid answer

~40 s

The `fields` mapping parameter lets one JSON value be indexed several times under different types or analyzers. The standard pairing is a `text` parent for full-text `match` queries plus a `keyword` sub-field, addressed as `title.keyword`, that holds the value verbatim with `doc_values` so you can filter exactly, sort, and run `terms` aggregations. Dynamic mapping creates exactly this shape for any new string, with `ignore_above: 256` on the `keyword` half. Multi-fields are not limited to that pair: a second `text` sub-field with the `english` analyzer gives you stemmed recall alongside the exact-token parent. The cost is real — each sub-field is a separate set of index structures on disk — so on fields you truly never filter or aggregate on, drop the sub-field explicitly rather than inheriting it.

code

json · 11 lines
json
{
  "properties": {
    "title": {
      "type": "text",
      "fields": {
        "keyword": { "type": "keyword", "ignore_above": 256 },
        "english": { "type": "text", "analyzer": "english" }
      }
    }
  }
}

go deeper

for a junior

Know that title.keyword is a sub-field of title holding the raw value, and that filters, sorting and aggregations must target it while match queries target the analyzed parent.

for a middle

Explain the fields mapping parameter: one source value indexed under several mappings, the dynamic default of text plus keyword with ignore_above, and the extra disk and indexing cost each sub-field carries.

for a senior

Demonstrate the operational edge: a mapping update adds the sub-field but backfills nothing, so existing documents need update_by_query or a reindex before facets are trustworthy. Justify dropping inherited sub-fields on large text fields.

for a principal

Own the mapping policy across teams: explicit mappings instead of dynamic multi-fields on wide documents, a convention for analyzer sub-fields used in relevance work, and the storage bill that doubling every string field quietly creates.

## What a multi-field actually is A multi-field is a field mapping that declares extra sub-mappings under the `fields` parameter. When a document is indexed, Elasticsearch takes the single JSON value and feeds it to the parent mapping *and* to every sub-field mapping, writing several independent sets of index structures from one source value. The document itself is unchanged; `_source` still contains one `title`. Only the index gains extra entries. You address a sub-field by dotted path: if `title` is `text` with a sub-field named `keyword`, then `title` is the analyzed field and `title.keyword` is the verbatim one. The sub-field name is arbitrary — `keyword` is only a convention that dynamic mapping happens to use. ## Why the text-plus-keyword pair is the default The two string types answer different questions and neither can answer both. An analyzed `text` field supports `match`, phrase queries and relevance ranking, but stores no whole-value term and no `doc_values`, so it cannot be filtered exactly, sorted, or aggregated. A `keyword` field can do all of those and none of the full-text ones. When Elasticsearch meets a string with no explicit mapping, it refuses to guess and gives you both: ```json "title": { "type": "text", "fields": { "keyword": { "type": "keyword", "ignore_above": 256 } } } ``` That is why so many real queries and Kibana visualizations refer to `something.keyword`: the analyzed parent cannot serve a bucket aggregation, and the sub-field can. ## Beyond the default pair Multi-fields are a general mechanism for indexing the same text several ways: - **Different analyzers.** A parent `text` field with the `standard` analyzer plus a `title.english` sub-field using the `english` analyzer lets you query both and combine them: the stemmed sub-field boosts recall ("running" finds "run"), the exact parent boosts precision, and a `multi_match` across both usually ranks better than either alone. - **A different type.** A numeric-looking string can be indexed both as `keyword` and as a numeric sub-field if some queries need ranges and others need exact term lookups. - **Search-as-you-type or n-gram sub-fields.** Prefix matching gets its own analysis chain without polluting the main field. ## What it costs Each sub-field is a full set of index structures: its own postings, its own `doc_values` where applicable, its own norms. A `keyword` sub-field on a long free-text field is not free — you are storing the entire value again as a single term (up to `ignore_above`) plus doc values for it. On a large index of article bodies, an inherited `body.keyword` can be pure waste: nobody aggregates on article bodies, and any body over the character limit is silently absent from that sub-field anyway. Each sub-field also counts toward the index's field count, so dynamically mapping a wide, unpredictable document doubles your field growth. ## The operational gotcha Adding a new sub-field to an existing mapping is allowed — sub-fields are additive, and unlike a type change the request succeeds. But it only affects documents indexed *after* the change. Existing documents have no data in the new sub-field, so filters and aggregations on it silently under-report. To populate it you must rewrite the existing documents, typically with `update_by_query` on the same index or a reindex into a new one. Interviewers like this question because "the mapping update succeeded" and "the data is there" are two different things. The mirror-image gotcha is querying the wrong half. `{"term": {"title": "Wireless Mouse"}}` returns nothing because `title` is analyzed; `{"term": {"title.keyword": "Wireless Mouse"}}` works. A sort on `title` errors or is meaningless; a sort on `title.keyword` is correct. Half of "Elasticsearch returns no results" tickets are a missing `.keyword` suffix. ## How to decide, per field Start from usage, not from the default: - Searched full-text only (article body, description): plain `text`, drop the sub-field. - Filtered, sorted, faceted only (status, SKU, tenant ID, host name): plain `keyword`, no analyzed parent. - Genuinely both (product title, person name, tag that users also type into a search box): `text` with an explicit `keyword` sub-field, and set `ignore_above` deliberately. Writing the mapping explicitly costs a few minutes and removes a whole class of surprises: no silent 256-character truncation you did not choose, no doubled index for fields nobody facets on, and no `.keyword` archaeology when someone reads the queries a year later.

  • You add a keyword sub-field to a mapping on a live index. Why do aggregations on it still come back nearly empty?
    The mapping update applies to future indexing only. Documents already in the index were written before the sub-field existed, so they contribute nothing to it. You have to rewrite them — update_by_query over the same index, or a reindex into a fresh one — before the sub-field reflects the full corpus.
  • Does a multi-field increase the size of _source?
    No. _source stores the original JSON document exactly as it was sent, and a multi-field adds no keys to it. What grows is the index: each sub-field writes its own postings, doc values and norms, so disk usage and indexing time rise even though the stored document is unchanged.
  • When would you add a second text sub-field with a different analyzer?
    When exact and stemmed matching pull in opposite directions. Keep the parent on a light analyzer for precision, add an english-analyzed sub-field for recall, then query both in a multi_match and boost the exact field higher. Documents matching the precise form rank above those matching only the stem.

saying these in an interview costs you the question

  • Thinks a multi-field duplicates the value inside _source
  • Assumes adding a sub-field backfills existing documents
  • Says you can sort on the analyzed parent field
  • Believes every string needs a keyword sub-field
  • Cannot say what the .keyword suffix in a query refers to

context