skip to content

Bucket Aggregations

Bucket aggregations group documents — by term, by date interval, by range — and are the foundation everything else nests inside. Interviewers ask about terms accuracy: each shard returns only its own top N, so global counts can legitimately be wrong.

part ofElasticsearchoverview, primer and where to startread it →
on this pageshow

questions

6

In Elasticsearch, what does a terms aggregation return and what does its size parameter control?

level: juniorimportance: must knowfreq 68%

answer

  1. Think SQL GROUP BY
  2. One bucket per distinct field value
  3. A default caps how many buckets return
  4. Default is ten, ordered by count
  5. Leftovers reported as sum_other_doc_count

basics

~20 s

A terms aggregation groups matching documents by the distinct values of a field and returns one bucket per value with a doc_count. size caps how many buckets come back, defaulting to 10, ordered by descending count.

solid answer

~40 s

A `terms` aggregation is the group-by of Elasticsearch: it looks at the values of one aggregatable field across all documents the query matched and returns a bucket per distinct value, each with a `doc_count`. By default it returns the 10 buckets with the highest counts, ties broken by the term's ascending order. `size` changes how many buckets you get back — it is a limit on the *response*, not on how much data is scanned, so a large `size` still costs memory and CPU. The response also carries `sum_other_doc_count`, the number of matching documents that fell into terms you did not receive. Typical extras are `order`, `include`/`exclude`, `min_doc_count` and `missing`, and any metric or bucket aggregation can be nested inside each bucket via `aggs`.

code

json · 8 lines
json
{
  "size": 0,
  "aggs": {
    "top_tags": {
      "terms": { "field": "tags", "size": 5 }
    }
  }
}

go deeper

for a junior

Be ready to write a terms aggregation from memory, name its default bucket limit, and read a bucket's key and doc_count out of a response.

for a middle

Explain that size is a response limit, not a scan limit, and that the ordering default is descending doc_count with ties broken by term. Know why the keyword sub-field is the aggregatable one.

for a senior

Show you can predict the memory cost of a large or nested terms aggregation and reach for include/exclude, min_doc_count and shard_size deliberately rather than just raising size until the numbers look right.

for a principal

Own the guidance for your teams on when aggregation is the wrong tool at all — high-cardinality group-bys that belong in a rollup, a data stream with pre-aggregated indices, or a downstream analytics store.

## What a bucket aggregation is Elasticsearch splits aggregations into three families. *Metric* aggregations compute a number over a set of documents (`avg`, `sum`, `cardinality`). *Bucket* aggregations do not compute a number — they partition the documents the query matched into groups, and then let you nest more aggregations inside each group. *Pipeline* aggregations post-process the output of other aggregations. The `terms` aggregation is the most-used bucket aggregation and is the direct analogue of a SQL `GROUP BY column`. ## What terms actually returns Given `{"terms": {"field": "tags"}}`, Elasticsearch inspects the values of `tags` in every document that matched the query and produces one bucket per distinct value: ``` "buckets": [ { "key": "kotlin", "doc_count": 190 }, { "key": "java", "doc_count": 155 } ] ``` `key` is the term; `doc_count` is how many matching documents contained it. Because a field can be multi-valued (an array), a single document can contribute to several buckets — the sum of `doc_count` across buckets is therefore not necessarily the number of matching documents. Alongside `buckets` the response carries two bookkeeping fields: `sum_other_doc_count`, the number of matching documents whose term did not make it into the returned buckets, and `doc_count_error_upper_bound`, a bound on how far the returned counts might be off on a multi-shard index. ## size, and what it does not do `size` controls how many buckets are returned, and defaults to 10. Raising it does **not** make Elasticsearch scan less or more of the index — every matching document is visited either way. What `size` changes is how many terms are kept, sorted and shipped back, and how much per-bucket state the coordinating node must hold while merging shard results. That is why a very large `size` on a high-cardinality field is a classic way to trip a circuit breaker: you are asking one node to materialise millions of buckets, not just to read more data. If you want counts only and no documents, set the request-level `"size": 0` so the `hits` array is empty. Note the collision of names: the top-level `size` is the number of search hits, the `size` inside the `terms` object is the number of buckets. ## Ordering and ties The default `order` is `{"_count": "desc"}` — the most frequent terms first — with ties broken by the term ordered ascending. You can order by `_key` instead to get alphabetical or numeric ordering, which is stable and cheap, or by the value of a nested metric sub-aggregation (`{"order": {"avg_price": "desc"}}`). Ordering by a sub-aggregation is legal but produces the least reliable results on a sharded index, because each shard picks its local candidates using that metric before the coordinating node ever sees the whole picture. ## Which fields can you aggregate A `terms` aggregation reads columnar `doc_values`, which are on by default for `keyword`, numeric, `date`, `boolean` and `ip` fields. It does not work out of the box on an analyzed `text` field: `text` has no doc values, and aggregating one requires explicitly enabling in-memory fielddata, which is discouraged. The idiomatic mapping is a `text` field for searching with a `keyword` sub-field (commonly `tags.keyword`) for aggregating and sorting. ## The parameters you will actually use - `order` — as above. - `include` / `exclude` — keep or drop terms by regular expression or by an explicit array of exact values. Useful for stripping a dominant junk term without post-processing. - `min_doc_count` — the minimum count a term needs to be returned; the default is 1, so terms with no matching documents never appear. - `missing` — a value to bucket documents that have no value for the field at all; without it they are simply absent from every bucket. - `shard_size` — how many candidate terms each shard contributes before merging, the main accuracy dial. ## Nesting The real power is composition: any bucket aggregation accepts an `aggs` block that runs once per bucket. `terms` on `status` with a nested `avg` on `amount` gives average order value per status; a `terms` inside a `terms` gives a two-level breakdown, at the cost of multiplying bucket counts (100 outer × 100 inner = 10,000 buckets). Keep an eye on that product — nested `terms` is the usual cause of an aggregation that suddenly needs gigabytes.

  • Why does a terms aggregation on a text field fail, and what is the standard mapping fix?
    Analyzed `text` fields have no doc values, so there is nothing columnar to group over; Elasticsearch refuses unless you explicitly enable fielddata, which is memory-hungry and discouraged. The standard fix is a multi-field: index the value as `text` for search and add a `keyword` sub-field (`tags.keyword`) for aggregating, sorting and exact matching.
  • How does significant_terms differ from terms, and when would you use it?
    `terms` ranks by raw frequency, so it surfaces whatever is simply common. `significant_terms` ranks by how much more frequent a term is in the query's result set (the foreground) than in the wider index (the background), so it surfaces what is *distinctive* about those documents rather than what is merely popular. It is the right choice for suggesting related terms, anomaly hunting or automatic tagging, and it costs more because it needs background frequencies too.
  • Why can the doc_count values in a terms aggregation add up to more than the number of matching documents?
    Because the field can be multi-valued. A document with `tags: ["java", "kotlin"]` is counted once in the `java` bucket and once in the `kotlin` bucket, so the buckets overlap. Summing `doc_count` is only equal to the hit count for single-valued fields; use `cardinality` or a separate `value_count` when you need document-level totals.

saying these in an interview costs you the question

  • Thinking size limits how much data Elasticsearch scans
  • Assuming terms returns every distinct value by default
  • Aggregating directly on a text field instead of its keyword sub-field
  • Confusing the request-level size with the terms size
  • Assuming bucket doc_counts always sum to the total hits

context

open as a page

Why can an Elasticsearch terms aggregation report wrong doc_count values on a multi-shard index?

level: middleimportance: must knowfreq 72%

basics

~20 s

Each shard independently returns only its own top candidate terms, and the coordinating node sums those partial lists. A term ranked low on one shard contributes nothing from it, so counts can be undercounted or the term missed entirely.

open as a page

What is the difference between calendar_interval and fixed_interval in an Elasticsearch date_histogram?

level: middleimportance: should knowfreq 58%

basics

~20 s

calendar_interval buckets by calendar units whose real length varies — months differ in length and daylight-saving days are not 24 hours — and accepts only a single unit. fixed_interval buckets by an exact multiple of milliseconds, so every bucket is identical.

open as a page

Why does a terms aggregation over an Elasticsearch nested field return no buckets on its own?

level: middleimportance: should knowfreq 38%

basics

~20 s

Objects under a nested field are indexed as separate hidden Lucene documents. A root-level aggregation only sees root documents, which carry no values for those subfields, so it finds nothing. Wrapping it in a nested aggregation switches to that context.

open as a page

When do you need an Elasticsearch composite aggregation instead of a plain terms aggregation?

level: seniorimportance: should knowfreq 40%

basics

~20 s

When you must walk every bucket rather than the top few. A terms aggregation returns only its top size buckets and offers no next page; composite streams all buckets in composite-key order and pages through them with after_key.

open as a page

How does the time_zone parameter change bucket boundaries in an Elasticsearch date_histogram aggregation?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Elasticsearch stores dates as UTC milliseconds; time_zone makes the rounding to bucket boundaries happen in that zone instead. Day and month buckets then start at local midnight, and daylight-saving transitions make calendar day buckets 23 or 25 hours long.

open as a page