skip to content

Why does aggregating on an Elasticsearch text field require fielddata, and why is it off by default?

level: middleimportance: must knowfreq 72%

answer

  1. Aggregations need doc-to-value, not term-to-doc
  2. Analyzed fields store tokens, not the original value
  3. One structure is on disk, the other on heap
  4. Uninverting the index at query time costs JVM memory
  5. The sub-field with doc_values is the fix

basics

~20 s

Analyzed text fields have no doc_values, so aggregating on them makes Elasticsearch uninvert the index into heap-resident fielddata, whose size grows with token count. That risks the heap, so it is disabled by default; aggregate on a keyword sub-field instead.

solid answer

~50 s

Aggregations need a document-to-value lookup, which Elasticsearch normally reads from **doc_values** — a columnar structure written at index time and read off-heap through the OS page cache. `text` fields are analyzed into tokens and have no doc_values at all. Setting `"fielddata": true` enables a fallback where, at query time, Elasticsearch uninverts each segment's inverted index into a doc-to-terms structure **on the JVM heap**, sized by the number of distinct tokens and documents. That is why it is off by default: it is the classic route to long GC pauses and a tripped fielddata circuit breaker. The normal fix is to aggregate on a `keyword` multi-field such as `title.keyword`, which has doc_values. Fielddata is only legitimate when you genuinely want to bucket on analyzed tokens — a word cloud or `significant_terms` over body text — and then you bound it with `fielddata_frequency_filter`.

code

json · 12 lines
json
{
  "mappings": {
    "properties": {
      "title": {
        "type": "text",
        "fields": {
          "keyword": { "type": "keyword", "ignore_above": 256 }
        }
      }
    }
  }
}

go deeper

for a junior

Recall that you aggregate and sort on keyword fields, not text fields, and that dynamic mapping gives you a .keyword sub-field for free. Know the error message when you see it.

for a middle

Explain the mechanics: doc_values are columnar and written at index time, text has none because it is analyzed, and fielddata uninverts the index into the heap on demand. Name the keyword multi-field as the fix.

for a senior

Show you have operated this. Tie fielddata to GC pauses and breaker trips, know how to inspect and clear the field data cache, and be able to say when fielddata is genuinely correct and how you would bound it.

for a principal

Own the policy: whether analyzed-token aggregation is allowed at all, how mappings and templates prevent an engineer from enabling fielddata on a hot index, and how the cost is moved to index time instead of query time.

## Two different data layouts An inverted index answers "which documents contain term T". An aggregation asks the opposite question — "for document D, what value does field F hold?" — and asks it for every matching document. Elasticsearch answers that from **doc_values**: a columnar, per-segment structure written at index time alongside the inverted index. Doc values live on disk and are read through the operating system's page cache, so they are effectively off-heap; a large terms aggregation streams them instead of filling the JVM. Doc values are on by default for the structured types — `keyword`, the numeric types, `date`, `boolean`, `ip`, `geo_point`. They are not available on `text`. ## Why text has no doc_values A `text` field is analyzed. `"The Quick Brown Fox"` is stored as the token stream the analyzer produced, and the original string is not retained in the index structure at all. There is no single value to place in a column, and bucketing on what is there would give you buckets named `the`, `quick`, `brown`, `fox` — almost never what a facet or a group-by wants. Asking for `"doc_values": true` on a `text` field is rejected by the mapping API. ## Fielddata is the in-heap fallback Setting `"fielddata": true` on a `text` field turns on the legacy path. At query time, per segment, Elasticsearch **uninverts** the inverted index: it walks every term's postings list and builds a document-to-terms structure **on the JVM heap**. It is built lazily on first use for each segment and cached in the field data cache, where it stays until the segment is merged away or the entry is evicted. The size of that structure scales with the number of distinct tokens and the number of documents, and it is heap the garbage collector must traverse on every major collection. On a large free-text field this is the textbook way to push a node into multi-second GC pauses and then out of the cluster. ## The guardrails around it - The **fielddata circuit breaker** (`indices.breaker.fielddata.limit`, 40% of the heap by default) estimates the cost before the load happens and throws `CircuitBreakingException` rather than letting the node OOM. - `indices.fielddata.cache.size` bounds the cache. Bounding it converts a memory problem into a latency problem: evicted entries mean the uninversion is paid again on the next query. - `GET _cat/fielddata?v` and `GET _nodes/stats/indices/fielddata?fields=*` show per-field heap consumption, and `POST /my-index/_cache/clear?fielddata=true` drops it. Note that the error message Elasticsearch returns is explicit — it tells you fielddata is disabled on the field and suggests either enabling it or using a keyword field. Candidates who reflexively follow the first suggestion are the ones who cause outages. ## The fix in practice: a keyword multi-field Index the field twice: `text` for matching, `keyword` for aggregating and sorting. Dynamic mapping already does this for strings, producing `title` plus `title.keyword` with `ignore_above: 256`. Aggregate on `title.keyword`. One nuance worth knowing: `ignore_above` means values longer than that many characters are not indexed into the keyword sub-field, so unusually long values silently vanish from the facet rather than erroring. If you are faceting over something that can be long, raise or remove `ignore_above` deliberately. ## When fielddata really is the right tool Sometimes tokens *are* the thing you want to count: a tag cloud of the words in a body field, or `significant_terms` over free text to surface distinguishing vocabulary. Then enable fielddata deliberately and bound it with `fielddata_frequency_filter`, which loads only terms whose document frequency falls between a minimum and maximum, and skips small segments below `min_segment_size`. That cuts the long tail of one-off tokens, which is where most of the memory goes. The alternative — often better — is to make the analysis explicit at index time: index a second `keyword` field populated by an ingest pipeline, or use a dedicated analyzed-but-aggregatable representation, so the cost is paid once during indexing instead of on every query. ## What an interviewer is listening for That you distinguish heap from page cache; that you know doc_values are the normal aggregation source and `text` simply lacks them; that your first fix is the keyword sub-field and not `fielddata: true`; and that raising the fielddata circuit breaker limit is not a fix, it just removes the seatbelt.

  • If fielddata is loaded per segment and cached, why does a heavily-indexed index make it worse rather than better?
    New segments arrive constantly and merges retire old ones, so cache entries are invalidated and rebuilt continuously. Instead of paying the uninversion once, you pay it repeatedly, and the cache holds structures for segments that are about to disappear. Fielddata suits static or read-mostly indices far better than a hot write path.
  • You added a keyword multi-field, but a few facet values are missing from the results. What is the likely cause?
    `ignore_above` on the keyword sub-field. Dynamic mapping sets it to 256 characters, and any value longer than that is not indexed into the keyword field at all — the document still exists and still matches text queries, but it contributes no bucket. Raise or remove `ignore_above` and reindex.
  • Where do global ordinals for a keyword field show up in node stats?
    Under field data. Even though keyword aggregations read doc_values from disk, the shard-level global ordinal map that a terms aggregation builds is held in memory and accounted for as field data, so `_cat/fielddata` on a keyword-heavy cluster is not necessarily evidence that someone enabled fielddata on text.

Doc values are a spreadsheet column you read straight down. Fielddata is reconstructing that column at query time by reading the book's back-of-book index and noting, for every word, which pages it appeared on.

saying these in an interview costs you the question

  • Enabling fielddata: true as the first response to the error
  • Claiming text fields have doc_values that are just disabled
  • Thinking doc_values live on the JVM heap
  • Raising the fielddata circuit breaker limit to make the error go away
  • Expecting a text aggregation to bucket whole field values

context