skip to content

Why does a terms aggregation over an Elasticsearch nested field return no buckets on its own?

level: middleimportance: should knowfreq 38%

answer

  1. Lucene documents are flat, not nested
  2. Nested objects become separate hidden documents
  3. Aggregations run over matching root documents
  4. A wrapper aggregation switches the document context
  5. Counting parents needs a second switch back

basics

~20 s

Objects under a nested field are indexed as separate hidden Lucene documents. A root-level aggregation only sees root documents, which carry no values for those subfields, so it finds nothing. Wrapping it in a nested aggregation switches to that context.

solid answer

~50 s

When a field is mapped as `nested`, each object in the array is indexed as its own hidden Lucene document, linked to its parent. Aggregations run over the root documents the query matched, and those roots hold no values for `variants.color`, so a plain `terms` on it produces an empty bucket list — no error, just nothing. You have to descend explicitly with `{"nested": {"path": "variants"}}` and put the `terms` inside it; then the aggregation runs over the nested documents. The consequence to state next is that inside that context `doc_count` counts **nested** documents, not parents — three matching variants on one product count as three. Add a `reverse_nested` sub-aggregation to climb back to the roots and count distinct products. And note that a nested aggregation sees *all* nested documents of the matching roots, not only the ones that satisfied a nested query, so filters usually have to be repeated inside it.

code

json · 13 lines
json
{
  "mappings": {
    "properties": {
      "variants": {
        "type": "nested",
        "properties": {
          "color": { "type": "keyword" },
          "price": { "type": "double" }
        }
      }
    }
  }
}

go deeper

for a junior

Recall that nested fields need a nested aggregation wrapper with a path, and that without it you get an empty result rather than an error.

for a middle

Explain that nested objects are separate hidden Lucene documents, that aggregations start from matching roots, and that reverse_nested is how you get back to parent counts.

for a senior

Catch the subtle version: a nested query filter does not constrain the nested aggregation, so the filter must be repeated inside it — and weigh nested indexing cost against denormalising the grouped attribute onto the root.

for a principal

Own the modelling decision. Nested arrays multiply document count and merge cost across the whole cluster; decide where per-object fidelity is worth that and where a denormalised root field or a separate index serves the query pattern better.

## What the nested field type does at index time Lucene has no concept of an object; a document is a flat set of fields. Elasticsearch flattens ordinary JSON objects, so `variants: [{color: red, size: S}, {color: blue, size: L}]` becomes `variants.color: [red, blue]` and `variants.size: [S, L]` — and the association between `red` and `S` is destroyed. That is the classic cross-object matching bug: a query for red *and* size L matches a product that has neither combination. Mapping `variants` as `"type": "nested"` fixes it by indexing each array element as a **separate hidden Lucene document**, stored in the same segment block as its parent root document and linked to it. Each nested document holds only that object's fields; the root document holds the rest. Correct per-object matching then falls out of a join between the two. ## Why the aggregation comes back empty Aggregations execute over the documents the query matched — and a search matches and returns *root* documents. The root document for a product carries `name`, `price`, `sku`, but not `variants.color`, because those values live in the child documents. So `{"terms": {"field": "variants.color"}}` at the top level of the `aggs` block is asking root documents for a field none of them has. The result is a valid response with an empty `buckets` array. There is no error and no warning, which is why this is a bug report rather than a stack trace: "the facet panel is blank in production." The `nested` aggregation is the explicit context switch. `{"nested": {"path": "variants"}}` takes the set of matching roots, follows the link down to all of their nested documents, and makes those the input to whatever you nest inside its `aggs`. A `terms` on `variants.color` there sees real values and produces buckets. ## doc_count changes meaning This is the second half of the answer and the part weak candidates miss. Inside a nested aggregation, `doc_count` counts nested documents. If one product has three red variants, the `red` bucket gets 3, not 1. For an e-commerce facet — "how many *products* have a red variant" — that number is wrong and inflates popular products. `reverse_nested` climbs back out. Placed as a sub-aggregation inside the nested context, `{"reverse_nested": {}}` returns to the root documents that own the nested documents in the current bucket, and its own `doc_count` is the number of distinct parents. So the correct facet is `nested → terms(variants.color) → reverse_nested`, reading the reverse-nested `doc_count` as the product count. `reverse_nested` also accepts a `path` to jump to an intermediate level in a multi-level nesting rather than all the way to the root. ## Query filters do not travel down automatically The third trap: a search whose query is `{"nested": {"path": "variants", "query": {"term": {"variants.color": "red"}}}}` matches products that have at least one red variant. The matching unit is still the root product. When a `nested` aggregation then descends, it descends to *all* of that product's variants — blue and green included. So a nested aggregation on price under that query reports statistics over every variant of every product that has a red one, which is almost never what was asked. The fix is to repeat the constraint inside the nested aggregation with a `filter` sub-aggregation: `nested(path) → filter(term variants.color=red) → avg(variants.price)`. Redundant-looking, entirely necessary. Interviewers use exactly this scenario because it separates people who have shipped a faceted search from people who have read the docs. ## Cost, and the alternatives Nested documents multiply the number of Lucene documents in the index — a product with 50 variants is 51 documents — which inflates index size, merge work and the per-shard document count that the index-level nested-object limits are there to protect. Aggregating over them is proportionally more expensive than over roots. So reach for `nested` only when per-object correctness genuinely matters. If you never need to correlate two subfields of the same object, plain `object` fields with their flattening are cheaper and aggregate directly with no wrapper. If you only need one grouped attribute, denormalising it onto the root (`colors: [red, blue]` as a top-level `keyword`) gives you a straightforward `terms` aggregation with root-level `doc_count` and none of this machinery — at the price of keeping the denormalised copy in sync on every update.

  • Inside a nested aggregation, why is doc_count often the wrong number for a facet count, and what fixes it?
    Inside the nested context `doc_count` counts nested documents, so a product with three red variants contributes 3 to the red bucket and inflates it. Add a `reverse_nested` sub-aggregation inside each term bucket; it climbs back to the owning root documents and its own `doc_count` is the number of distinct parents, which is what a facet should display.
  • A search filters for red variants and a nested aggregation averages variant price. Why is the average wrong?
    The query matches root products that have at least one red variant, but the nested aggregation then descends to *all* variants of those products — blue and green included. Query-level nested filters do not propagate into the aggregation context. Repeat the constraint as a `filter` sub-aggregation inside the nested aggregation so only red variants feed the average.
  • When is denormalising onto the root document better than using a nested field?
    When you never need to correlate two attributes of the same object. Copying just the values you group on to a root-level keyword array gives a plain terms aggregation whose doc_count already counts products, with no nested wrapper, no reverse_nested and none of the document multiplication that inflates index size and merge cost. The tradeoff is keeping the copy in sync on every update.

saying these in an interview costs you the question

  • Expecting a root-level terms aggregation to see nested fields
  • Reading nested doc_count as a count of parent documents
  • Assuming a nested query filter constrains the nested aggregation
  • Thinking the empty result must be a mapping error
  • Mapping everything nested without weighing the document multiplication

context