skip to content

Aggregations

Elasticsearch's analytics half: bucket documents by term, range, or date, compute metrics inside each bucket, and nest further aggregations underneath. Interviewers ask because dashboards are built out of these, and because a terms aggregation on a high-cardinality field is a classic way to take a cluster down.

part ofElasticsearchoverview, primer and where to startread it →
on this pageshow

explore

questions

24

In Elasticsearch, what does a terms aggregation return and what does its size parameter control?

level: juniorimportance: must knowfreq 68%

answer

  1. Think SQL GROUP BY
  2. One bucket per distinct field value
  3. A default caps how many buckets return
  4. Default is ten, ordered by count
  5. Leftovers reported as sum_other_doc_count

basics

~20 s

A terms aggregation groups matching documents by the distinct values of a field and returns one bucket per value with a doc_count. size caps how many buckets come back, defaulting to 10, ordered by descending count.

solid answer

~40 s

A `terms` aggregation is the group-by of Elasticsearch: it looks at the values of one aggregatable field across all documents the query matched and returns a bucket per distinct value, each with a `doc_count`. By default it returns the 10 buckets with the highest counts, ties broken by the term's ascending order. `size` changes how many buckets you get back — it is a limit on the *response*, not on how much data is scanned, so a large `size` still costs memory and CPU. The response also carries `sum_other_doc_count`, the number of matching documents that fell into terms you did not receive. Typical extras are `order`, `include`/`exclude`, `min_doc_count` and `missing`, and any metric or bucket aggregation can be nested inside each bucket via `aggs`.

code

json · 8 lines
json
{
  "size": 0,
  "aggs": {
    "top_tags": {
      "terms": { "field": "tags", "size": 5 }
    }
  }
}

go deeper

for a junior

Be ready to write a terms aggregation from memory, name its default bucket limit, and read a bucket's key and doc_count out of a response.

for a middle

Explain that size is a response limit, not a scan limit, and that the ordering default is descending doc_count with ties broken by term. Know why the keyword sub-field is the aggregatable one.

for a senior

Show you can predict the memory cost of a large or nested terms aggregation and reach for include/exclude, min_doc_count and shard_size deliberately rather than just raising size until the numbers look right.

for a principal

Own the guidance for your teams on when aggregation is the wrong tool at all — high-cardinality group-bys that belong in a rollup, a data stream with pre-aggregated indices, or a downstream analytics store.

## What a bucket aggregation is Elasticsearch splits aggregations into three families. *Metric* aggregations compute a number over a set of documents (`avg`, `sum`, `cardinality`). *Bucket* aggregations do not compute a number — they partition the documents the query matched into groups, and then let you nest more aggregations inside each group. *Pipeline* aggregations post-process the output of other aggregations. The `terms` aggregation is the most-used bucket aggregation and is the direct analogue of a SQL `GROUP BY column`. ## What terms actually returns Given `{"terms": {"field": "tags"}}`, Elasticsearch inspects the values of `tags` in every document that matched the query and produces one bucket per distinct value: ``` "buckets": [ { "key": "kotlin", "doc_count": 190 }, { "key": "java", "doc_count": 155 } ] ``` `key` is the term; `doc_count` is how many matching documents contained it. Because a field can be multi-valued (an array), a single document can contribute to several buckets — the sum of `doc_count` across buckets is therefore not necessarily the number of matching documents. Alongside `buckets` the response carries two bookkeeping fields: `sum_other_doc_count`, the number of matching documents whose term did not make it into the returned buckets, and `doc_count_error_upper_bound`, a bound on how far the returned counts might be off on a multi-shard index. ## size, and what it does not do `size` controls how many buckets are returned, and defaults to 10. Raising it does **not** make Elasticsearch scan less or more of the index — every matching document is visited either way. What `size` changes is how many terms are kept, sorted and shipped back, and how much per-bucket state the coordinating node must hold while merging shard results. That is why a very large `size` on a high-cardinality field is a classic way to trip a circuit breaker: you are asking one node to materialise millions of buckets, not just to read more data. If you want counts only and no documents, set the request-level `"size": 0` so the `hits` array is empty. Note the collision of names: the top-level `size` is the number of search hits, the `size` inside the `terms` object is the number of buckets. ## Ordering and ties The default `order` is `{"_count": "desc"}` — the most frequent terms first — with ties broken by the term ordered ascending. You can order by `_key` instead to get alphabetical or numeric ordering, which is stable and cheap, or by the value of a nested metric sub-aggregation (`{"order": {"avg_price": "desc"}}`). Ordering by a sub-aggregation is legal but produces the least reliable results on a sharded index, because each shard picks its local candidates using that metric before the coordinating node ever sees the whole picture. ## Which fields can you aggregate A `terms` aggregation reads columnar `doc_values`, which are on by default for `keyword`, numeric, `date`, `boolean` and `ip` fields. It does not work out of the box on an analyzed `text` field: `text` has no doc values, and aggregating one requires explicitly enabling in-memory fielddata, which is discouraged. The idiomatic mapping is a `text` field for searching with a `keyword` sub-field (commonly `tags.keyword`) for aggregating and sorting. ## The parameters you will actually use - `order` — as above. - `include` / `exclude` — keep or drop terms by regular expression or by an explicit array of exact values. Useful for stripping a dominant junk term without post-processing. - `min_doc_count` — the minimum count a term needs to be returned; the default is 1, so terms with no matching documents never appear. - `missing` — a value to bucket documents that have no value for the field at all; without it they are simply absent from every bucket. - `shard_size` — how many candidate terms each shard contributes before merging, the main accuracy dial. ## Nesting The real power is composition: any bucket aggregation accepts an `aggs` block that runs once per bucket. `terms` on `status` with a nested `avg` on `amount` gives average order value per status; a `terms` inside a `terms` gives a two-level breakdown, at the cost of multiplying bucket counts (100 outer × 100 inner = 10,000 buckets). Keep an eye on that product — nested `terms` is the usual cause of an aggregation that suddenly needs gigabytes.

  • Why does a terms aggregation on a text field fail, and what is the standard mapping fix?
    Analyzed `text` fields have no doc values, so there is nothing columnar to group over; Elasticsearch refuses unless you explicitly enable fielddata, which is memory-hungry and discouraged. The standard fix is a multi-field: index the value as `text` for search and add a `keyword` sub-field (`tags.keyword`) for aggregating, sorting and exact matching.
  • How does significant_terms differ from terms, and when would you use it?
    `terms` ranks by raw frequency, so it surfaces whatever is simply common. `significant_terms` ranks by how much more frequent a term is in the query's result set (the foreground) than in the wider index (the background), so it surfaces what is *distinctive* about those documents rather than what is merely popular. It is the right choice for suggesting related terms, anomaly hunting or automatic tagging, and it costs more because it needs background frequencies too.
  • Why can the doc_count values in a terms aggregation add up to more than the number of matching documents?
    Because the field can be multi-valued. A document with `tags: ["java", "kotlin"]` is counted once in the `java` bucket and once in the `kotlin` bucket, so the buckets overlap. Summing `doc_count` is only equal to the hit count for single-valued fields; use `cardinality` or a separate `value_count` when you need document-level totals.

saying these in an interview costs you the question

  • Thinking size limits how much data Elasticsearch scans
  • Assuming terms returns every distinct value by default
  • Aggregating directly on a text field instead of its keyword sub-field
  • Confusing the request-level size with the terms size
  • Assuming bucket doc_counts always sum to the total hits

context

open as a page

Why can an Elasticsearch terms aggregation report wrong doc_count values on a multi-shard index?

level: middleimportance: must knowfreq 72%

basics

~20 s

Each shard independently returns only its own top candidate terms, and the coordinating node sums those partial lists. A term ranked low on one shard contributes nothing from it, so counts can be undercounted or the term missed entirely.

open as a page

Why is Elasticsearch's cardinality aggregation approximate, and what does precision_threshold trade off?

level: middleimportance: must knowfreq 70%

basics

~20 s

The cardinality aggregation estimates distinct counts with a HyperLogLog++ sketch of fixed size instead of tracking every value, so memory stays bounded as cardinality grows. precision_threshold trades memory for accuracy: higher means near-exact counts up to a larger unique count.

open as a page

Why can Elasticsearch's percentiles aggregation report a p99 that differs from the exact value?

level: middleimportance: must knowfreq 55%

basics

~20 s

The percentiles aggregation summarizes the value distribution in a bounded TDigest sketch rather than sorting every value, so results are approximate. TDigest is most accurate at extreme percentiles and least accurate near the median, and per-shard sketches are merged.

open as a page

Why does aggregating on an Elasticsearch text field require fielddata, and why is it off by default?

level: middleimportance: must knowfreq 72%

basics

~20 s

Analyzed text fields have no doc_values, so aggregating on them makes Elasticsearch uninvert the index into heap-resident fielddata, whose size grows with token count. That risks the heap, so it is disabled by default; aggregate on a keyword sub-field instead.

open as a page

How does buckets_path locate the value a pipeline aggregation consumes?

level: middleimportance: must knowfreq 55%

basics

~20 s

buckets_path is a small path language over the aggregation tree: a name refers to an aggregation, > descends into a sub-aggregation, . picks one value from a multi-value metric, and keywords like _count and _key address the bucket itself.

open as a page

In Elasticsearch, how do parent and sibling pipeline aggregations differ?

level: middleimportance: must knowfreq 68%

basics

~20 s

A parent pipeline is declared inside the multi-bucket aggregation it processes and acts on each of its buckets, as derivative and bucket_script do. A sibling is declared beside that aggregation and returns one result over all its buckets, as avg_bucket and max_bucket do.

open as a page

A terms aggregation fails with CircuitBreakingException — how do you diagnose and fix it?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Read which breaker tripped from the exception: request means the aggregation's own structures, fielddata means uninverted text or ordinals, parent means cluster-wide heap pressure. Fix the query — smaller size, composite paging, narrower filter, pre-aggregation — rather than raising the limit.

open as a page

In Elasticsearch, what does the stats aggregation return, and how does extended_stats extend it?

level: juniorimportance: should knowfreq 55%

basics

~10 s

The stats aggregation returns count, min, max, avg and sum for a numeric field in a single pass. extended_stats adds sum_of_squares, variance, standard deviation and standard-deviation bounds around the mean.

open as a page

What input does an Elasticsearch pipeline aggregation consume, and when does it run?

level: juniorimportance: should knowfreq 42%

basics

~20 s

A pipeline aggregation consumes the output of another aggregation — its buckets or its metric values — instead of documents. It is pointed at that output with buckets_path and is computed during the reduce phase, after the source aggregation has produced results.

open as a page

What is the difference between calendar_interval and fixed_interval in an Elasticsearch date_histogram?

level: middleimportance: should knowfreq 58%

basics

~20 s

calendar_interval buckets by calendar units whose real length varies — months differ in length and daylight-saving days are not 24 hours — and accepts only a single unit. fixed_interval buckets by an exact multiple of milliseconds, so every bucket is identical.

open as a page

Why does a terms aggregation over an Elasticsearch nested field return no buckets on its own?

level: middleimportance: should knowfreq 38%

basics

~20 s

Objects under a nested field are indexed as separate hidden Lucene documents. A root-level aggregation only sees root documents, which carry no values for those subfields, so it finds nothing. Wrapping it in a nested aggregation switches to that context.

open as a page

Why does an Elasticsearch value_count aggregation return more than the bucket's doc_count?

level: middleimportance: should knowfreq 45%

basics

~10 s

value_count counts extracted field values, not documents. A multi-valued field contributes one count per array element, so a document with three tags adds three, while a document missing the field adds nothing.

open as a page

In an Elasticsearch search request, what does size: 0 change and why does the shard request cache need it?

level: middleimportance: should knowfreq 50%

basics

~20 s

size: 0 returns no documents, so Elasticsearch skips the fetch phase and only computes aggregations and the hit count. The shard request cache only stores shard-level results for such requests, making repeated identical dashboard queries nearly free until the next refresh.

open as a page

Why does a derivative over a date_histogram misreport change when intervals have no documents?

level: middleimportance: should knowfreq 46%

basics

~20 s

A derivative subtracts consecutive buckets, so what counts as consecutive depends on which buckets came back. Metrics like avg yield no value on an empty interval, and min_doc_count above zero drops those intervals entirely, so the derivative silently spans a gap.

open as a page

When do you need an Elasticsearch composite aggregation instead of a plain terms aggregation?

level: seniorimportance: should knowfreq 40%

basics

~20 s

When you must walk every bucket rather than the top few. A terms aggregation returns only its top size buckets and offers no next page; composite streams all buckets in composite-key order and pages through them with after_key.

open as a page

How does the time_zone parameter change bucket boundaries in an Elasticsearch date_histogram aggregation?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Elasticsearch stores dates as UTC milliseconds; time_zone makes the rounding to bucket boundaries happen in that zone instead. Day and month buckets then start at local midnight, and daylight-saving transitions make calendar day buckets 23 or 25 hours long.

open as a page

What are global ordinals in Elasticsearch, and when does eager_global_ordinals pay off?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Global ordinals map each segment's local keyword term numbering onto one shard-wide numbering so terms aggregations can count into integer slots instead of hashing strings. They are rebuilt lazily after each refresh; eager_global_ordinals moves that build into refresh, trading write cost for stable search latency.

open as a page

How do you sort and page Elasticsearch aggregation buckets by a computed value?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Compute the value per bucket with bucket_script, then order the buckets with bucket_sort, which also takes from and size. A terms aggregation's own order clause cannot reference a pipeline aggregation, so bucket_sort is the mechanism.

open as a page

When does bucket_selector run, and why doesn't it make a terms aggregation cheaper?

level: seniorimportance: should knowfreq 50%

basics

~20 s

bucket_selector runs during the reduce phase, after the parent aggregation has collected documents and already chosen its buckets. It only removes buckets from the response, so all the collection work and memory has already been spent, and it can never bring back a bucket the parent did not return.

open as a page

Your billing report counts unique customers with a cardinality aggregation — how do you decide whether that approximation is acceptable?

level: principalimportance: should knowfreq 32%

basics

~20 s

Classify the consumer first: dashboards and trends tolerate estimates, invoices and audits do not. The cardinality aggregation returns no error bound, so if the number must be defensible, restructure so an exact metric answers it rather than tuning precision.

open as a page

How would you stop dashboard aggregations over 90 days of logs from destabilizing a shared Elasticsearch cluster?

level: principalimportance: should knowfreq 32%

basics

~20 s

Measure which panels cost what, then shrink the input with tiering and retention, pre-aggregate the recurring analyses into summary indices, make the remaining queries cacheable, and bound the blast radius with bucket limits and workload isolation before buying heap.

open as a page

How do init_script, map_script, combine_script and reduce_script divide work in a scripted_metric aggregation?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

init_script runs once per shard to seed a state object, map_script runs once per matching document to accumulate into it, combine_script runs once per shard to produce the value sent over the wire, and reduce_script runs once on the coordinating node over the list of shard results.

open as a page

How do the sampler and diversified_sampler aggregations bound the cost of expensive sub-aggregations?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Both are single-bucket aggregations that restrict their sub-aggregations to a capped set of documents per shard — sampler keeps the top-scoring shard_size hits, diversified_sampler additionally limits how many hits share a given field value, so one prolific source cannot dominate the sample.

open as a page