skip to content

In Elasticsearch, how do parent and sibling pipeline aggregations differ?

level: middleimportance: must knowfreq 68%

answer

  1. Where you declare it decides everything
  2. Inside the bucket agg, or beside it
  3. One value per bucket vs one value total
  4. avg_bucket and max_bucket sit outside
  5. Inside, the path is relative to the bucket

basics

~20 s

A parent pipeline is declared inside the multi-bucket aggregation it processes and acts on each of its buckets, as derivative and bucket_script do. A sibling is declared beside that aggregation and returns one result over all its buckets, as avg_bucket and max_bucket do.

solid answer

~40 s

The difference is **where you declare it and what shape the output has**. A *parent* pipeline lives inside the `aggs` block of the multi-bucket aggregation it consumes, and its `buckets_path` is relative to each bucket — so `derivative`, `cumulative_sum`, `moving_fn` and `bucket_script` add a new value to every bucket, while `bucket_selector` and `bucket_sort` remove or reorder buckets instead. A *sibling* pipeline is declared next to that aggregation, at the same level, and its `buckets_path` must walk down into it with `>` — for example `sales_per_month>sales`. It produces one new aggregation result beside the original, not a per-bucket value: `avg_bucket`, `sum_bucket`, `min_bucket`, `max_bucket`, `stats_bucket`, `percentiles_bucket`. Some parent pipelines also constrain their parent: `derivative` and `cumulative_sum` require a `histogram` or `date_histogram`, because they assume ordered, evenly spaced buckets.

code

json · 13 lines
json
{
  "size": 0,
  "aggs": {
    "sales_per_month": {
      "date_histogram": { "field": "date", "calendar_interval": "month" },
      "aggs": {
        "sales": { "sum": { "field": "price" } },
        "mom_change": { "derivative": { "buckets_path": "sales" } }
      }
    },
    "best_month": { "max_bucket": { "buckets_path": "sales_per_month>sales" } }
  }
}

go deeper

for a junior

Recall that some pipeline aggregations go inside the bucket aggregation and some go next to it, and be able to name derivative and avg_bucket as examples of each.

for a middle

Explain the mechanics: placement decides whether buckets_path is relative or must descend with >, and whether the output is a value per bucket or one value overall.

for a senior

Show judgment about the failure modes — a sibling over a truncated terms aggregation summarises only the returned buckets, and histogram-only pipelines reject a terms parent.

for a principal

Own how these compose in dashboards at scale: every parent pipeline multiplies per-bucket work on the coordinating node, so bucket cardinality, not the pipeline itself, is the capacity question.

## Two placements, two output shapes Every pipeline aggregation is either a **parent** pipeline or a **sibling** pipeline, and the family decides three things at once: where the aggregation is written in the request body, how its `buckets_path` is interpreted, and what shape appears in the response. ### Parent pipelines A parent pipeline is written **inside** the `aggs` block of the multi-bucket aggregation whose buckets it consumes. It runs once per bucket, and its `buckets_path` is resolved *relative to that bucket* — so the path is usually just the name of a sibling metric aggregation in the same block, or a special keyword such as `_count`. ```json "sales_per_month": { "date_histogram": { "field": "date", "calendar_interval": "month" }, "aggs": { "sales": { "sum": { "field": "price" } }, "mom_change": { "derivative": { "buckets_path": "sales" } } } } ``` Each month bucket comes back with an extra `mom_change` value. The parent family splits further by what it does to the bucket: - **Adds a value:** `derivative`, `serial_diff`, `cumulative_sum`, `cumulative_cardinality`, `moving_fn`, `moving_percentiles`, `bucket_script`, `normalize`. - **Removes buckets:** `bucket_selector`, whose script returns a boolean and whose false buckets vanish from the response. - **Reorders and truncates buckets:** `bucket_sort`, which sorts the parent's bucket list and can apply `from`/`size`. Some parent pipelines are pickier about what they can be nested in. `derivative`, `cumulative_sum`, `moving_fn` and `serial_diff` assume an **ordered series of evenly spaced buckets**, so they must sit inside a `histogram` or `date_histogram`; putting a `derivative` inside a `terms` aggregation is rejected. `bucket_script`, `bucket_selector` and `bucket_sort` have no such restriction and are perfectly at home under `terms`. ### Sibling pipelines A sibling pipeline is written **next to** the aggregation it consumes, at the same level of the tree, and produces **one** new result rather than a per-bucket value. Because it is not inside the aggregation, its `buckets_path` has to walk down into it, using `>` to descend and the metric name at the end: ```json "best_month": { "max_bucket": { "buckets_path": "sales_per_month>sales" } } ``` The sibling family is the `_bucket` set: `avg_bucket`, `sum_bucket`, `min_bucket`, `max_bucket`, `stats_bucket`, `extended_stats_bucket`, `percentiles_bucket`. They answer questions of the form "across all the buckets of that aggregation, what is the average/max/spread of this metric?" — the average monthly revenue, the busiest hour, the distribution of per-customer spend. `min_bucket` and `max_bucket` have a response shape worth knowing: they return a `value` **and** a `keys` array, because more than one bucket can tie for the extreme. Candidates who expect a single scalar key are surprised by the array. ## Choosing between them The question to ask is: *do I want a number per bucket, or a number about the buckets?* "Show me the change from month to month" is per bucket, so it is a parent pipeline. "Which month was the best" or "what was the average month" is about the whole bucket set, so it is a sibling. The two compose freely in one request — a `derivative` inside the histogram and a `max_bucket` beside it, both reading the same `sales` metric. ## Common mistakes The classic error is placing a sibling-style aggregation inside the bucket aggregation, or vice versa. Declaring `avg_bucket` inside the `date_histogram` gives it a path that resolves to a single bucket's metric, which is not a bucket set, and the request fails. Declaring `derivative` at the top level with `buckets_path: "sales_per_month>sales"` fails too, because `derivative` validates that its parent is a histogram-family aggregation. The second mistake is forgetting that the two families differ in **path relativity**. Inside the histogram, `"buckets_path": "sales"` is correct and `"sales_per_month>sales"` is not; outside it, the opposite holds. Reading a path error message as "the aggregation is broken" rather than "the pipeline is on the wrong level" wastes a lot of debugging time. The third is assuming a sibling pipeline sees the full data. It sees the buckets the source aggregation returned — for a `terms` aggregation that means the top `size` terms, so `max_bucket` over a truncated `terms` is the maximum *of the returned buckets*, not of the index.

  • Why does max_bucket return a keys array instead of a single key?
    Because several buckets can tie for the maximum value. `max_bucket` returns `value` plus a `keys` array holding the key of every bucket that reached it, so a tie between two months yields two keys. Client code that reads `keys[0]` blindly silently hides ties; code that assumes a scalar breaks outright.
  • Can you put a derivative aggregation inside a terms aggregation?
    No. `derivative` assumes an ordered series of evenly spaced buckets, so it must be nested in a `histogram` or `date_histogram`; against a `terms` aggregation the request is rejected. The same restriction applies to `cumulative_sum`, `serial_diff` and `moving_fn`. `bucket_script`, `bucket_selector` and `bucket_sort` have no such requirement and work fine under `terms`.
  • How does buckets_path differ between the two families?
    A parent pipeline's path is resolved relative to each bucket, so it names a metric in the same `aggs` block — `"sales"`. A sibling pipeline sits outside the bucket aggregation, so its path must descend into it with `>` — `"sales_per_month>sales"`. Using the wrong form for the placement is the most common cause of an unresolvable-path error.

A parent pipeline is an extra column added to every row of a report; a sibling pipeline is the single summary line printed underneath the table.

saying these in an interview costs you the question

  • Declares avg_bucket inside the date_histogram it should summarise
  • Uses a full > path from inside the bucket aggregation
  • Expects max_bucket to return a single scalar key
  • Thinks derivative works under a terms aggregation
  • Calls bucket_selector a sibling because it filters the result

context