How does buckets_path locate the value a pipeline aggregation consumes?
answer
- It addresses aggregations by name, not fields
- One character descends the tree
- One character picks a value from stats
- Two underscore keywords address the bucket
- Scripted pipelines take a map, not a string
basics
~20 sbuckets_path is a small path language over the aggregation tree: a name refers to an aggregation, > descends into a sub-aggregation, . picks one value from a multi-value metric, and keywords like _count and _key address the bucket itself.
solid answer
~40 s`buckets_path` is a miniature path language over the aggregation tree, resolved relative to where the pipeline aggregation is declared. A bare name refers to an aggregation at that level — `"sales"`. The `>` separator descends into a sub-aggregation, which is how a sibling pipeline reaches inside a bucket aggregation: `"sales_per_month>sales"`. A `.` selects one value out of a multi-value metric aggregation, as in `"sales_stats.sum"`. Special keywords address the bucket rather than an aggregation: `_count` is the bucket's `doc_count`, `_key` its key, and `_bucket_count` the number of buckets in a named multi-bucket sub-aggregation. `bucket_script` and `bucket_selector` take `buckets_path` as a **map** of variable names to paths, and the Painless script reads them as `params.<name>`. An unresolvable path is a request error, not a silent null.
code
json · 17 lines{
"size": 0,
"aggs": {
"by_category": {
"terms": { "field": "category", "size": 20 },
"aggs": {
"revenue": { "sum": { "field": "price" } },
"revenue_per_doc": {
"bucket_script": {
"buckets_path": { "total": "revenue", "docs": "_count" },
"script": "params.docs == 0 ? 0 : params.total / params.docs"
}
}
}
}
}
}go deeper
Recall that buckets_path names other aggregations, that > walks into a sub-aggregation, and that _count is the bucket's document count.
Explain path relativity — inside the bucket aggregation a bare name suffices, outside it you must descend — and the map form that feeds params in bucket_script.
Distinguish a path error from a gap: one fails the request, the other is governed by gap_policy, and insert_zeros quietly fabricates data points for averages and rates.
Treat these paths as an API surface dashboards depend on; renaming an aggregation silently breaks every pipeline path referencing it, so naming conventions and query review matter at team scale.
## Why a path language exists A pipeline aggregation consumes a value that some other aggregation produced. Since aggregations form a tree of named nodes, referencing that value needs an addressing scheme, and `buckets_path` is it. Getting it right is most of the practical difficulty of pipeline aggregations, and path errors are the most common failure people hit. ## The grammar **A bare name** refers to an aggregation reachable from the pipeline's own position. Inside a `date_histogram`'s `aggs` block, `"buckets_path": "sales"` means "the `sales` aggregation in this same block, evaluated for this bucket". **`>` descends** one level into a sub-aggregation. A sibling pipeline declared outside a bucket aggregation must use it to get in: `"sales_per_month>sales"` means "walk into the `sales_per_month` aggregation, and for each of its buckets take the `sales` metric". Paths can descend several levels: `"by_region>by_month>sales"`. **`.` selects a value** from a multi-value metric aggregation. A `stats` aggregation produces `count`, `min`, `max`, `avg` and `sum` in one node, so a pipeline aggregation must say which one it wants: `"sales_stats.sum"`. Single-value metrics such as `sum` or `avg` need no suffix. **Special keywords** address the bucket itself instead of a named aggregation. `_count` is the bucket's `doc_count` — invaluable, because it lets you compute per-document averages or thresholds without adding a redundant `value_count` aggregation. `_key` is the bucket key. `_bucket_count` gives the number of buckets in a named multi-bucket sub-aggregation, which is how you filter parent buckets by how many children they have — for example, keeping only sessions that visited more than three distinct pages. ## The map form Most pipeline aggregations consume exactly one value and take `buckets_path` as a string. The script-driven ones, `bucket_script` and `bucket_selector`, consume several and take a **map** from variable name to path: ```json "revenue_per_doc": { "bucket_script": { "buckets_path": { "total": "revenue", "docs": "_count" }, "script": "params.docs == 0 ? 0 : params.total / params.docs" } } ``` The keys of the map become the names under `params` inside the Painless script. This is worth stressing in an interview: the script does not name aggregations, it names the map keys, and mixing the two up produces a null-pointer failure at execution rather than a helpful path error. ## Relativity: the mistake everyone makes A path is resolved **relative to where the pipeline aggregation is declared**. Inside the bucket aggregation, `"sales"` is correct and `"sales_per_month>sales"` is wrong; outside it, exactly the reverse. Half of all "my pipeline aggregation does not work" problems are a path written for the other placement. The error message names the path that could not be resolved, which points straight at the fix once you know the rule. ## Missing values versus broken paths These two are different and are often confused. A **broken path** — naming an aggregation that does not exist, or descending into something that is not a bucket aggregation — is a request-level error; the search fails. A **missing value** — the path resolves fine, but a particular bucket has no value for it, as an `avg` yields on an empty bucket — is handled by `gap_policy`. The default `skip` treats that bucket as though it did not exist and continues from the next available value; `insert_zeros` substitutes zero, which is right for counts and sums but silently invents data for averages and rates. ## Practical notes Paths reference aggregation **names**, not field names; if the metric is named `sales` over field `price`, the path is `sales`. Because paths are resolved on already-computed results, a pipeline aggregation may reference **another pipeline aggregation's** output — a `bucket_sort` sorting on a `bucket_script` value is the everyday case. And on a bucket aggregation that truncates, such as `terms` with a `size`, the path resolves only over the buckets that came back; nothing in the path language reaches data the source aggregation did not return.
- How do you reference a bucket's document count from a pipeline aggregation?Use the special path `_count`, which resolves to the bucket's `doc_count`. It saves adding a redundant `value_count` aggregation and is the usual denominator in a `bucket_script` — revenue per order, errors per request. The related `_key` addresses the bucket key, and `_bucket_count` gives the number of buckets in a named multi-bucket sub-aggregation.
- Inside a bucket_script, why does params.revenue fail when buckets_path maps 'total' to the revenue aggregation?Because `params` is keyed by the **map keys**, not by aggregation names. With `"buckets_path": { "total": "revenue" }` the script must read `params.total`; `params.revenue` is simply absent and the script fails at execution. Keeping the map key and the aggregation name identical avoids the confusion entirely.
- What is the difference between an unresolvable buckets_path and a bucket with no value?An unresolvable path — a name that does not exist, or a descent into something that is not a bucket aggregation — is a request error and the search fails. A resolvable path whose value happens to be missing for a bucket, such as an `avg` over an empty bucket, is governed by `gap_policy`: `skip` by default, or `insert_zeros` to substitute zero.
saying these in an interview costs you the question
- Writes a field name in buckets_path instead of the aggregation name
- Uses a full > path from inside the bucket aggregation
- Reads params by aggregation name rather than map key
- Expects a bad path to yield null instead of failing the request
- Adds a value_count aggregation instead of using _count