skip to content

In an Elasticsearch mapping, how do the text and keyword field types differ?

level: juniorimportance: must knowfreq 88%

answer

  1. one is analyzed, one is not
  2. think match versus term query
  3. only one of them can be sorted
  4. which type gets doc_values by default?
  5. the whole string stays one term

basics

~10 s

text is analyzed into tokens for full-text search and cannot be sorted or aggregated by default. keyword stores the value verbatim as one term, which is what filters, sorts and aggregations need.

solid answer

~40 s

They are the two string types and they are indexed completely differently. A `text` field is run through an analyzer at index time, so "Wireless Mouse" becomes the terms `wireless` and `mouse`; that makes `match` queries work but means the original string no longer exists as a searchable unit, and `text` has no `doc_values`, so sorting and aggregating on it are disabled by default. A `keyword` field is not analyzed at all: the whole string, case and punctuation included, becomes a single term, and it gets `doc_values`, so term filters, sorts, and `terms` aggregations all work. Rule of thumb: human-readable prose that people search inside is `text`; identifiers, status codes, tags, e-mail addresses and anything you filter, sort or group by is `keyword`. Dynamic mapping hedges by giving you both.

code

json · 9 lines
json
{
  "mappings": {
    "properties": {
      "description": { "type": "text" },
      "status":      { "type": "keyword" },
      "sku":         { "type": "keyword" }
    }
  }
}

go deeper

for a junior

Be ready to state the split in one breath: text is analyzed for full-text search, keyword is stored verbatim for exact filters, sorting and aggregations. Know that term queries on analyzed fields usually return nothing.

for a middle

Explain the mechanics: the analyzer runs at index time and stores tokens, text has no doc_values so aggregations are disabled, keyword produces one term plus doc_values. Say why dynamic mapping creates both halves.

for a senior

Show you have debugged the empty-result-set version of this in production, and know the fix is a new index plus a reindex because field types are immutable. Discuss the index-size and heap cost of each choice.

for a principal

Own the policy: explicit mappings over dynamic guesses, a naming and typing convention across teams, and the cost of getting it wrong at scale — wasted analyzed copies of machine fields versus reindexing petabytes to correct a type.

## Two ways to store a string Elasticsearch offers two general-purpose string field types, and choosing between them is the single most consequential mapping decision most teams make. The difference is not cosmetic: it changes what is written into the index, and therefore what queries can ever work. ## What happens to a text field A field mapped as `text` is passed through an **analyzer** at index time. The analyzer splits the string into tokens, lowercases them, and may apply further filters. With the default `standard` analyzer, the value `"Wireless Mouse"` is stored in the inverted index as two terms: `wireless` and `mouse`. The original string is *not* stored as a term anywhere in the index — it survives only in `_source`, the JSON document Elasticsearch keeps for retrieval. That is exactly what you want for full-text search. A `match` query on that field runs the same analyzer over the query string, so searching `mouse`, `MOUSE` or `wireless mouse` all find the document. It is also why a `term` query — which does no analysis and looks the query string up verbatim — usually returns nothing on a `text` field: nobody ever indexed the term `"Wireless Mouse"`. Because `text` fields are tokenized, there is no single value per document to sort by or group on. Elasticsearch reflects this by giving `text` fields no `doc_values` (the columnar, per-document structure that powers sorting and aggregations). Aggregating directly on a `text` field fails unless you enable `fielddata`, an in-heap structure that is disabled by default precisely because it is expensive. ## What happens to a keyword field A field mapped as `keyword` is not analyzed. The exact string — including case, spaces and punctuation — becomes one term. `"Wireless Mouse"` is a single term `Wireless Mouse`, and only a `term` query for that exact string matches it. A `match` query still works mechanically, but since the field has no analyzer, the query text is compared as a whole, so partial words do not match. `keyword` fields have `doc_values` enabled by default, so they can be sorted on, aggregated on, and read in scripts cheaply from disk rather than heap. They also disable `norms` by default, because length normalization is meaningless for an exact-match field. ## Choosing between them - Prose a human searches inside — product descriptions, article bodies, titles, comments — is `text`. - Values a machine matches exactly — user IDs, SKUs, status codes, tags, host names, e-mail addresses, country codes, log levels — are `keyword`. - Values you sort by, group by, or build a facet from must be `keyword` (or a numeric/date type). The classic production bug is mapping a status field as `text`: `{"term": {"status": "IN_PROGRESS"}}` silently returns zero hits because the analyzer lowercased and possibly split the value, and the dashboard's `terms` aggregation either errors or buckets fragments instead of whole statuses. ## Why dynamic mapping gives you both When Elasticsearch maps a string for the first time without an explicit mapping, it does not choose — it creates a `text` field with a `keyword` **sub-field** (a multi-field) so that `title` works for `match` and `title.keyword` works for term filters, sorts and aggregations. That is a convenience, not an endorsement: on a field you never search full-text, the analyzed half is wasted index; on a field with long values the `keyword` half carries `ignore_above: 256`, so long values are quietly absent from it. ## Cost and immutability The two types cost different things. `text` costs analysis time, a larger postings list (many terms per document), and optional positions for phrase matching. `keyword` costs one term plus `doc_values` per document, and for very high-cardinality fields the `doc_values` and global-ordinal work behind aggregations dominates. Crucially, you cannot change the type of an existing field. Once `status` is `text` in a live index, turning it into `keyword` means creating a new index with the corrected mapping and reindexing into it, then swapping over. That is why the type decision is worth getting right at design time rather than discovering it from an empty result set in production. ## A quick mental check Ask of every string field: *will a human type a fragment of this and expect a hit?* If yes, `text`. *Will code compare it, sort by it, or count occurrences of it?* If yes, `keyword`. If genuinely both, map it as `text` with an explicit `keyword` sub-field and be deliberate about the sub-field's settings rather than inheriting the dynamic default.

  • You already have a live index where status was mapped as text. How do you fix it?
    You cannot change a field's type in place. Create a new index with the corrected mapping, reindex the data into it, and swap traffic over — ideally through an alias so readers never see the switch. If the field is only wrong for filtering, a cheaper stopgap is to add a keyword sub-field and repopulate it, but that still requires rewriting existing documents.
  • Can you aggregate on a text field at all?
    Only by setting fielddata: true, which builds an in-heap, per-segment structure of the field's terms at query time. It is disabled by default because it can consume enormous heap and trip circuit breakers on high-cardinality fields. The correct answer in production is almost always a keyword sub-field with doc_values instead.
  • Does a keyword field ignore case?
    No — the value is stored byte-for-byte, so "Active" and "active" are different terms and a term query for one will not find the other. If you need case-insensitive exact matching, either normalize the data before indexing, use a keyword normalizer such as lowercase, or use the case_insensitive option on term-level queries.

text is like the index at the back of a book — the prose has been chopped into lookup words. keyword is the label on a filing-cabinet drawer: one exact string you either match or you don't.

saying these in an interview costs you the question

  • Says a term query on a text field does exact matching
  • Claims keyword fields are analyzed by the standard analyzer
  • Thinks a text field can be aggregated with no extra setting
  • Believes you can change text to keyword in place
  • Maps IDs and status codes as text because it is the default

context