What does Elasticsearch do with a keyword value longer than the field's ignore_above setting?
answer
- it is about length, not content
- the document still indexes fine
- _source keeps the value, the index does not
- aggregation counts quietly under-report
- dynamic keyword sub-fields carry this limit
basics
~20 sThe value is skipped for that field: it is neither indexed nor written to doc_values, so term queries, sorts and aggregations never see it. It still appears in _source, which makes the loss easy to miss.
solid answer
~50 s`ignore_above` sets a maximum character length for a `keyword` value. Anything longer is silently dropped from that field's index structures — no term in the inverted index, no entry in `doc_values` — so it cannot be matched by a `term` query, cannot be sorted on, and never appears in a `terms` aggregation bucket. The document itself is still indexed and still returned by other queries, and the full value is still in `_source`, so the search hit *looks* complete while the field is invisible to structured queries. The setting exists partly to protect you: Lucene rejects any term above a hard byte limit, so an unbounded `keyword` on user-supplied text can fail indexing outright. Dynamic mapping applies `ignore_above: 256` to the `keyword` sub-field it generates, which is where most people meet this behaviour without having chosen it.
code
json · 10 linesPUT /docs
{ "mappings": { "properties": {
"tag": { "type": "keyword", "ignore_above": 10 } } } }
PUT /docs/_doc/1
{ "tag": "short" }
PUT /docs/_doc/2
{ "tag": "this-tag-is-far-too-long" }
// doc 2 indexes fine, but tag is absent from the index for itgo deeper
Remember that ignore_above is a length cap on keyword values and that anything longer is simply not indexed for that field, even though the document itself is stored and returned.
Explain the mechanics: no term, no doc values, so term queries, exists, sorting and terms aggregations all skip the value while _source still shows it. Know that dynamic mapping applies 256 to generated keyword sub-fields.
Recognize the field-visible-but-missing-from-facets symptom quickly, and know the fix requires rewriting documents because mapping parameters only bind at index time. Explain why the guard beats a hard Lucene term-length failure on ingest.
Frame it as a data-quality boundary: decide per field whether oversized input should be dropped, normalized upstream or rejected loudly, and make sure dashboards built on aggregations disclose that they can silently under-count.
## The setting `ignore_above` is a mapping parameter on `keyword` fields. It takes a character count. When a document arrives whose value for that field is longer than the limit, Elasticsearch does not reject the document and does not truncate the value — it simply **does not add that value to the field's index structures**. Nothing goes into the inverted index, nothing goes into `doc_values`. For an array of strings, the rule applies per element: short elements are indexed, long ones are skipped, and the rest of the array is unaffected. ## Why it is so easy to miss The document is still indexed. It is still findable through every other field. And `_source` — the original JSON Elasticsearch stores for retrieval — is untouched, so when you fetch the document the long value is right there in the response. Everything looks normal. What breaks is structured access to that one field: - `{"term": {"url": "<the long value>"}}` returns nothing. - `{"exists": {"field": "url"}}` does not consider the document to have the field, because `exists` is driven by index structures. - Sorting on the field puts the document in the missing-value position. - A `terms` aggregation never produces a bucket for it, so counts under-report — and the total of the buckets no longer matches the document count. That combination — data visible in the hit, absent from the aggregation — is the diagnostic signature of `ignore_above`, and it is why the parameter shows up in interviews as a debugging scenario rather than a definition. ## Where the default comes from Most people never type `ignore_above`. They meet it because dynamic mapping applies it. When Elasticsearch maps a previously unseen string field, it creates a `text` field with a `keyword` sub-field carrying `ignore_above: 256`: ```json "url": { "type": "text", "fields": { "keyword": { "type": "keyword", "ignore_above": 256 } } } ``` So a URL, a stack trace, a log message or a long product description is fully searchable through the analyzed parent and completely absent from `url.keyword` the moment it passes 256 characters. Aggregating on `url.keyword` then quietly reports on the short values only. ## Why the guard exists at all Lucene imposes a hard maximum on the byte length of a single indexed term. A `keyword` field with no limit that receives an arbitrarily long user-supplied string will hit it, and the indexing request fails with a document-level error. Setting `ignore_above` converts a hard failure into a silent skip, which is usually the better operational trade-off for machine-generated data of unpredictable length: one enormous log line should not fail an ingest batch. Note the units differ between the two: `ignore_above` counts characters, while Lucene's ceiling is in bytes after UTF-8 encoding. With multi-byte text, a value that passes a generous character limit can still exceed the byte ceiling, so leaving a comfortable margin matters for non-ASCII corpora. ## Getting it right Decide per field, from how the field is used: - **Identifiers, codes, host names, e-mails.** These have a known maximum. Set `ignore_above` a little above it so genuinely malformed input is dropped rather than blowing up the shard, and legitimate values are never lost. - **Long free text you also want to facet on exactly.** Reconsider the requirement. A faceted value that can be thousands of characters is rarely a useful bucket; either normalize it upstream (hash it, extract a category) or accept the analyzed field only. - **Fields you never filter, sort or aggregate on.** Drop the `keyword` sub-field entirely instead of tuning its limit. You save the storage and the confusion. ## Changing it later `ignore_above` can be updated on an existing mapping — it is one of the few mapping parameters that is not frozen. But like any mapping change it takes effect at index time, so documents already written with the old limit keep the old outcome: previously skipped values are not retroactively indexed. Raising the limit therefore requires rewriting the affected documents with `update_by_query` or a reindex before aggregations are trustworthy again. ## How to confirm it in a live system The cheap check is a mismatch test: run a query that returns the document by another field, confirm the value in `_source`, then run a `term` query on the exact value and an `exists` query on the field. If the document comes back from the first and not the others, and the mapping shows an `ignore_above` shorter than the value, you have your answer without touching the data.
- You raise ignore_above from 256 to 2048 on a live index. Do the previously skipped values become searchable?No. Mapping parameters take effect when a document is indexed, so documents already written kept the old outcome and their long values were never added to the index. You need to rewrite them — update_by_query on the same index or a reindex into a new one — before the field is complete.
- What happens on a keyword field with no ignore_above when a very long value arrives?Elasticsearch tries to index the whole string as one term and Lucene rejects it once it exceeds its hard per-term byte limit, so the indexing request fails with a document-level error. On unpredictable machine-generated input that turns one oversized value into a failed bulk item, which is precisely what setting the parameter avoids.
saying these in an interview costs you the question
- Says the long value is truncated to the limit
- Thinks the whole document is rejected
- Assumes _source missing the value is the symptom
- Believes raising the limit backfills old documents
- Confuses the character limit with a byte limit