skip to content

Pipeline Aggregations

Pipeline aggregations consume the output of other aggregations instead of documents, which is how you get derivatives, moving averages, and per-bucket filtering. Interviewers use them to check you understand the parent vs sibling distinction and its second pass.

part ofElasticsearchoverview, primer and where to startread it →
on this pageshow

questions

6

How does buckets_path locate the value a pipeline aggregation consumes?

level: middleimportance: must knowfreq 55%

answer

  1. It addresses aggregations by name, not fields
  2. One character descends the tree
  3. One character picks a value from stats
  4. Two underscore keywords address the bucket
  5. Scripted pipelines take a map, not a string

basics

~20 s

buckets_path is a small path language over the aggregation tree: a name refers to an aggregation, > descends into a sub-aggregation, . picks one value from a multi-value metric, and keywords like _count and _key address the bucket itself.

solid answer

~40 s

`buckets_path` is a miniature path language over the aggregation tree, resolved relative to where the pipeline aggregation is declared. A bare name refers to an aggregation at that level — `"sales"`. The `>` separator descends into a sub-aggregation, which is how a sibling pipeline reaches inside a bucket aggregation: `"sales_per_month>sales"`. A `.` selects one value out of a multi-value metric aggregation, as in `"sales_stats.sum"`. Special keywords address the bucket rather than an aggregation: `_count` is the bucket's `doc_count`, `_key` its key, and `_bucket_count` the number of buckets in a named multi-bucket sub-aggregation. `bucket_script` and `bucket_selector` take `buckets_path` as a **map** of variable names to paths, and the Painless script reads them as `params.<name>`. An unresolvable path is a request error, not a silent null.

code

json · 17 lines
json
{
  "size": 0,
  "aggs": {
    "by_category": {
      "terms": { "field": "category", "size": 20 },
      "aggs": {
        "revenue": { "sum": { "field": "price" } },
        "revenue_per_doc": {
          "bucket_script": {
            "buckets_path": { "total": "revenue", "docs": "_count" },
            "script": "params.docs == 0 ? 0 : params.total / params.docs"
          }
        }
      }
    }
  }
}

go deeper

for a junior

Recall that buckets_path names other aggregations, that > walks into a sub-aggregation, and that _count is the bucket's document count.

for a middle

Explain path relativity — inside the bucket aggregation a bare name suffices, outside it you must descend — and the map form that feeds params in bucket_script.

for a senior

Distinguish a path error from a gap: one fails the request, the other is governed by gap_policy, and insert_zeros quietly fabricates data points for averages and rates.

for a principal

Treat these paths as an API surface dashboards depend on; renaming an aggregation silently breaks every pipeline path referencing it, so naming conventions and query review matter at team scale.

## Why a path language exists A pipeline aggregation consumes a value that some other aggregation produced. Since aggregations form a tree of named nodes, referencing that value needs an addressing scheme, and `buckets_path` is it. Getting it right is most of the practical difficulty of pipeline aggregations, and path errors are the most common failure people hit. ## The grammar **A bare name** refers to an aggregation reachable from the pipeline's own position. Inside a `date_histogram`'s `aggs` block, `"buckets_path": "sales"` means "the `sales` aggregation in this same block, evaluated for this bucket". **`>` descends** one level into a sub-aggregation. A sibling pipeline declared outside a bucket aggregation must use it to get in: `"sales_per_month>sales"` means "walk into the `sales_per_month` aggregation, and for each of its buckets take the `sales` metric". Paths can descend several levels: `"by_region>by_month>sales"`. **`.` selects a value** from a multi-value metric aggregation. A `stats` aggregation produces `count`, `min`, `max`, `avg` and `sum` in one node, so a pipeline aggregation must say which one it wants: `"sales_stats.sum"`. Single-value metrics such as `sum` or `avg` need no suffix. **Special keywords** address the bucket itself instead of a named aggregation. `_count` is the bucket's `doc_count` — invaluable, because it lets you compute per-document averages or thresholds without adding a redundant `value_count` aggregation. `_key` is the bucket key. `_bucket_count` gives the number of buckets in a named multi-bucket sub-aggregation, which is how you filter parent buckets by how many children they have — for example, keeping only sessions that visited more than three distinct pages. ## The map form Most pipeline aggregations consume exactly one value and take `buckets_path` as a string. The script-driven ones, `bucket_script` and `bucket_selector`, consume several and take a **map** from variable name to path: ```json "revenue_per_doc": { "bucket_script": { "buckets_path": { "total": "revenue", "docs": "_count" }, "script": "params.docs == 0 ? 0 : params.total / params.docs" } } ``` The keys of the map become the names under `params` inside the Painless script. This is worth stressing in an interview: the script does not name aggregations, it names the map keys, and mixing the two up produces a null-pointer failure at execution rather than a helpful path error. ## Relativity: the mistake everyone makes A path is resolved **relative to where the pipeline aggregation is declared**. Inside the bucket aggregation, `"sales"` is correct and `"sales_per_month>sales"` is wrong; outside it, exactly the reverse. Half of all "my pipeline aggregation does not work" problems are a path written for the other placement. The error message names the path that could not be resolved, which points straight at the fix once you know the rule. ## Missing values versus broken paths These two are different and are often confused. A **broken path** — naming an aggregation that does not exist, or descending into something that is not a bucket aggregation — is a request-level error; the search fails. A **missing value** — the path resolves fine, but a particular bucket has no value for it, as an `avg` yields on an empty bucket — is handled by `gap_policy`. The default `skip` treats that bucket as though it did not exist and continues from the next available value; `insert_zeros` substitutes zero, which is right for counts and sums but silently invents data for averages and rates. ## Practical notes Paths reference aggregation **names**, not field names; if the metric is named `sales` over field `price`, the path is `sales`. Because paths are resolved on already-computed results, a pipeline aggregation may reference **another pipeline aggregation's** output — a `bucket_sort` sorting on a `bucket_script` value is the everyday case. And on a bucket aggregation that truncates, such as `terms` with a `size`, the path resolves only over the buckets that came back; nothing in the path language reaches data the source aggregation did not return.

  • How do you reference a bucket's document count from a pipeline aggregation?
    Use the special path `_count`, which resolves to the bucket's `doc_count`. It saves adding a redundant `value_count` aggregation and is the usual denominator in a `bucket_script` — revenue per order, errors per request. The related `_key` addresses the bucket key, and `_bucket_count` gives the number of buckets in a named multi-bucket sub-aggregation.
  • Inside a bucket_script, why does params.revenue fail when buckets_path maps 'total' to the revenue aggregation?
    Because `params` is keyed by the **map keys**, not by aggregation names. With `"buckets_path": { "total": "revenue" }` the script must read `params.total`; `params.revenue` is simply absent and the script fails at execution. Keeping the map key and the aggregation name identical avoids the confusion entirely.
  • What is the difference between an unresolvable buckets_path and a bucket with no value?
    An unresolvable path — a name that does not exist, or a descent into something that is not a bucket aggregation — is a request error and the search fails. A resolvable path whose value happens to be missing for a bucket, such as an `avg` over an empty bucket, is governed by `gap_policy`: `skip` by default, or `insert_zeros` to substitute zero.

saying these in an interview costs you the question

  • Writes a field name in buckets_path instead of the aggregation name
  • Uses a full > path from inside the bucket aggregation
  • Reads params by aggregation name rather than map key
  • Expects a bad path to yield null instead of failing the request
  • Adds a value_count aggregation instead of using _count

context

open as a page

In Elasticsearch, how do parent and sibling pipeline aggregations differ?

level: middleimportance: must knowfreq 68%

basics

~20 s

A parent pipeline is declared inside the multi-bucket aggregation it processes and acts on each of its buckets, as derivative and bucket_script do. A sibling is declared beside that aggregation and returns one result over all its buckets, as avg_bucket and max_bucket do.

open as a page

What input does an Elasticsearch pipeline aggregation consume, and when does it run?

level: juniorimportance: should knowfreq 42%

basics

~20 s

A pipeline aggregation consumes the output of another aggregation — its buckets or its metric values — instead of documents. It is pointed at that output with buckets_path and is computed during the reduce phase, after the source aggregation has produced results.

open as a page

Why does a derivative over a date_histogram misreport change when intervals have no documents?

level: middleimportance: should knowfreq 46%

basics

~20 s

A derivative subtracts consecutive buckets, so what counts as consecutive depends on which buckets came back. Metrics like avg yield no value on an empty interval, and min_doc_count above zero drops those intervals entirely, so the derivative silently spans a gap.

open as a page

How do you sort and page Elasticsearch aggregation buckets by a computed value?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Compute the value per bucket with bucket_script, then order the buckets with bucket_sort, which also takes from and size. A terms aggregation's own order clause cannot reference a pipeline aggregation, so bucket_sort is the mechanism.

open as a page

When does bucket_selector run, and why doesn't it make a terms aggregation cheaper?

level: seniorimportance: should knowfreq 50%

basics

~20 s

bucket_selector runs during the reduce phase, after the parent aggregation has collected documents and already chosen its buckets. It only removes buckets from the response, so all the collection work and memory has already been spent, and it can never bring back a bucket the parent did not return.

open as a page