In Elasticsearch, what does a terms aggregation return and what does its size parameter control?
answer
- Think SQL GROUP BY
- One bucket per distinct field value
- A default caps how many buckets return
- Default is ten, ordered by count
- Leftovers reported as sum_other_doc_count
basics
~20 sA terms aggregation groups matching documents by the distinct values of a field and returns one bucket per value with a doc_count. size caps how many buckets come back, defaulting to 10, ordered by descending count.
solid answer
~40 sA `terms` aggregation is the group-by of Elasticsearch: it looks at the values of one aggregatable field across all documents the query matched and returns a bucket per distinct value, each with a `doc_count`. By default it returns the 10 buckets with the highest counts, ties broken by the term's ascending order. `size` changes how many buckets you get back — it is a limit on the *response*, not on how much data is scanned, so a large `size` still costs memory and CPU. The response also carries `sum_other_doc_count`, the number of matching documents that fell into terms you did not receive. Typical extras are `order`, `include`/`exclude`, `min_doc_count` and `missing`, and any metric or bucket aggregation can be nested inside each bucket via `aggs`.
code
json · 8 lines{
"size": 0,
"aggs": {
"top_tags": {
"terms": { "field": "tags", "size": 5 }
}
}
}go deeper
Be ready to write a terms aggregation from memory, name its default bucket limit, and read a bucket's key and doc_count out of a response.
Explain that size is a response limit, not a scan limit, and that the ordering default is descending doc_count with ties broken by term. Know why the keyword sub-field is the aggregatable one.
Show you can predict the memory cost of a large or nested terms aggregation and reach for include/exclude, min_doc_count and shard_size deliberately rather than just raising size until the numbers look right.
Own the guidance for your teams on when aggregation is the wrong tool at all — high-cardinality group-bys that belong in a rollup, a data stream with pre-aggregated indices, or a downstream analytics store.
## What a bucket aggregation is Elasticsearch splits aggregations into three families. *Metric* aggregations compute a number over a set of documents (`avg`, `sum`, `cardinality`). *Bucket* aggregations do not compute a number — they partition the documents the query matched into groups, and then let you nest more aggregations inside each group. *Pipeline* aggregations post-process the output of other aggregations. The `terms` aggregation is the most-used bucket aggregation and is the direct analogue of a SQL `GROUP BY column`. ## What terms actually returns Given `{"terms": {"field": "tags"}}`, Elasticsearch inspects the values of `tags` in every document that matched the query and produces one bucket per distinct value: ``` "buckets": [ { "key": "kotlin", "doc_count": 190 }, { "key": "java", "doc_count": 155 } ] ``` `key` is the term; `doc_count` is how many matching documents contained it. Because a field can be multi-valued (an array), a single document can contribute to several buckets — the sum of `doc_count` across buckets is therefore not necessarily the number of matching documents. Alongside `buckets` the response carries two bookkeeping fields: `sum_other_doc_count`, the number of matching documents whose term did not make it into the returned buckets, and `doc_count_error_upper_bound`, a bound on how far the returned counts might be off on a multi-shard index. ## size, and what it does not do `size` controls how many buckets are returned, and defaults to 10. Raising it does **not** make Elasticsearch scan less or more of the index — every matching document is visited either way. What `size` changes is how many terms are kept, sorted and shipped back, and how much per-bucket state the coordinating node must hold while merging shard results. That is why a very large `size` on a high-cardinality field is a classic way to trip a circuit breaker: you are asking one node to materialise millions of buckets, not just to read more data. If you want counts only and no documents, set the request-level `"size": 0` so the `hits` array is empty. Note the collision of names: the top-level `size` is the number of search hits, the `size` inside the `terms` object is the number of buckets. ## Ordering and ties The default `order` is `{"_count": "desc"}` — the most frequent terms first — with ties broken by the term ordered ascending. You can order by `_key` instead to get alphabetical or numeric ordering, which is stable and cheap, or by the value of a nested metric sub-aggregation (`{"order": {"avg_price": "desc"}}`). Ordering by a sub-aggregation is legal but produces the least reliable results on a sharded index, because each shard picks its local candidates using that metric before the coordinating node ever sees the whole picture. ## Which fields can you aggregate A `terms` aggregation reads columnar `doc_values`, which are on by default for `keyword`, numeric, `date`, `boolean` and `ip` fields. It does not work out of the box on an analyzed `text` field: `text` has no doc values, and aggregating one requires explicitly enabling in-memory fielddata, which is discouraged. The idiomatic mapping is a `text` field for searching with a `keyword` sub-field (commonly `tags.keyword`) for aggregating and sorting. ## The parameters you will actually use - `order` — as above. - `include` / `exclude` — keep or drop terms by regular expression or by an explicit array of exact values. Useful for stripping a dominant junk term without post-processing. - `min_doc_count` — the minimum count a term needs to be returned; the default is 1, so terms with no matching documents never appear. - `missing` — a value to bucket documents that have no value for the field at all; without it they are simply absent from every bucket. - `shard_size` — how many candidate terms each shard contributes before merging, the main accuracy dial. ## Nesting The real power is composition: any bucket aggregation accepts an `aggs` block that runs once per bucket. `terms` on `status` with a nested `avg` on `amount` gives average order value per status; a `terms` inside a `terms` gives a two-level breakdown, at the cost of multiplying bucket counts (100 outer × 100 inner = 10,000 buckets). Keep an eye on that product — nested `terms` is the usual cause of an aggregation that suddenly needs gigabytes.
- Why does a terms aggregation on a text field fail, and what is the standard mapping fix?Analyzed `text` fields have no doc values, so there is nothing columnar to group over; Elasticsearch refuses unless you explicitly enable fielddata, which is memory-hungry and discouraged. The standard fix is a multi-field: index the value as `text` for search and add a `keyword` sub-field (`tags.keyword`) for aggregating, sorting and exact matching.
- How does significant_terms differ from terms, and when would you use it?`terms` ranks by raw frequency, so it surfaces whatever is simply common. `significant_terms` ranks by how much more frequent a term is in the query's result set (the foreground) than in the wider index (the background), so it surfaces what is *distinctive* about those documents rather than what is merely popular. It is the right choice for suggesting related terms, anomaly hunting or automatic tagging, and it costs more because it needs background frequencies too.
- Why can the doc_count values in a terms aggregation add up to more than the number of matching documents?Because the field can be multi-valued. A document with `tags: ["java", "kotlin"]` is counted once in the `java` bucket and once in the `kotlin` bucket, so the buckets overlap. Summing `doc_count` is only equal to the hit count for single-valued fields; use `cardinality` or a separate `value_count` when you need document-level totals.
saying these in an interview costs you the question
- Thinking size limits how much data Elasticsearch scans
- Assuming terms returns every distinct value by default
- Aggregating directly on a text field instead of its keyword sub-field
- Confusing the request-level size with the terms size
- Assuming bucket doc_counts always sum to the total hits