skip to content

How does Elasticsearch's nested field type change the way an array of objects is indexed and queried?

level: middleimportance: must knowfreq 72%

answer

  1. each object becomes its own hidden document
  2. stored in one contiguous block with the root
  3. ordinary queries can't see those inner documents
  4. one query names the path and correlates clauses
  5. negation belongs outside, at the parent level

basics

~20 s

Mapping a field as nested indexes each object in the array as its own hidden Lucene document, stored in one block with its parent. A nested query with a path then requires all clauses to match the same inner object.

solid answer

~50 s

`"type": "nested"` tells Elasticsearch to index every object in the array as a separate hidden Lucene document, written contiguously in a block with the root document last — Lucene's block join. Because those inner documents are hidden, an ordinary `term` or `match` on `comments.author` matches nothing; you must wrap it in a `nested` query that names the `path`. Inside that query all clauses are evaluated against one inner document, so `author: alice` AND `votes > 10` can only match if a single comment satisfies both. `score_mode` (default `avg`) decides how the matching children's scores roll up into the parent's `_score`. Add `inner_hits` to see which children matched. A common trap: to exclude documents that have a matching child, put the whole `nested` query inside `bool.must_not` — a `must_not` **inside** the nested query only asks for some child that differs.

code

json · 13 lines
json
{
  "mappings": {
    "properties": {
      "comments": {
        "type": "nested",
        "properties": {
          "author": { "type": "keyword" },
          "votes":  { "type": "integer" }
        }
      }
    }
  }
}

go deeper

for a junior

Know that nested is a field type you declare in the mapping, and that querying it needs a nested query naming the path so that all conditions apply to the same object.

for a middle

Explain the block-join storage — children as hidden documents written with the root — why ordinary queries no longer see them, and what score_mode does with multiple matching children.

for a senior

Demonstrate the operational consequences: the reindex needed to adopt nested, the silent zero-hit regression on existing queries, and the correct way to express negation over a one-to-many relation.

for a principal

Own when a relation deserves nested at all. Weigh child cardinality, update rate and query shape against denormalizing, and set a standard so teams don't reach for nested on unbounded arrays.

## Declaring it ```json { "mappings": { "properties": { "comments": { "type": "nested", "properties": { "author": { "type": "keyword" }, "votes": { "type": "integer" } }} }}} ``` The sub-fields are mapped exactly as they would be under an `object`; only the container's type changes. Nested must be declared in the mapping — dynamic mapping never chooses it for you, so an array of objects is flattened unless you say otherwise, and switching an existing field from `object` to `nested` requires a reindex. ## What the index looks like Each object in the array becomes its own Lucene document. Those inner documents and their root are written **contiguously as a block** in the same segment, children first and the root document last. That layout is what makes Lucene's block join cheap: for any matching child, the engine can find its parent by scanning forward to the next document marked as a parent, with no term lookup and no cross-shard traffic. It is also why the whole block must be rewritten together whenever anything in the document changes. Inner documents are excluded from ordinary search by construction — every normal query is filtered to root documents only. This is the single most confusing consequence for newcomers: after switching a field to `nested`, previously working `term` queries on its sub-fields silently return zero hits rather than raising an error. ## Querying it ```json { "query": { "nested": { "path": "comments", "query": { "bool": { "must": [ { "term": { "comments.author": "alice" } }, { "range": { "comments.votes": { "gt": 10 } } } ]}}, "score_mode": "avg", "inner_hits": {} }}} ``` Inside the `nested` query, clauses are predicates over one inner document, so both conditions must hold for the same comment. Field names inside must still use the **full path** (`comments.author`, not `author`). Useful parameters: `path` (required), `query`, `score_mode`, `inner_hits`, and `ignore_unmapped` to return no hits instead of an error when the path does not exist in one of several searched indices. ## Scoring A parent may have many matching children. `score_mode` decides how those child scores become the parent's `_score`: `avg` (the default), `max`, `min`, `sum`, or `none` — the last discards child scores entirely, which is what you want when the nested query is a pure filter. In a filter context scores are not computed at all, so `score_mode` is irrelevant there. ## The must_not trap These two are not the same: - `bool.must_not: { nested: { path: "comments", query: { term: { "comments.author": "spam" } } } }` — documents where **no** comment is by spam. Usually what you want. - `nested: { path: "comments", query: { bool: { must_not: { term: { "comments.author": "spam" } } } } }` — documents having **at least one** comment not by spam, which a document full of spam plus one clean comment satisfies. Negation over a one-to-many relation has to be phrased at the parent level, and getting this backwards is a classic interview probe. ## Sorting and aggregating Sorting on a nested sub-field requires telling Elasticsearch which children to consider, via the `nested` option on the sort with a `path` and an optional `filter`, plus a `mode` such as `max` or `min` to reduce many child values to one sort value. Aggregating requires a `nested` aggregation to descend into the inner documents, and a `reverse_nested` aggregation to climb back to the root — a plain `terms` aggregation on a nested sub-field sees nothing, for the same reason a plain query does. ## Multi-level nesting and retrieval Nested fields can contain nested fields; the inner `nested` query's `path` is the full dotted path (`orders.items`), and each level adds another set of hidden documents. Retrieval of the matched entries is not automatic: the hit's `_source` always contains the entire original document, every child included, so to know *which* comment matched you must request `inner_hits`. ## What nested still does not give you It is not a join. Children live inside the parent document, cannot be indexed, updated or deleted on their own, and cannot be searched as documents in their own right. If children change far more often than parents, or their number is unbounded, nested is the wrong shape and a join field or plain denormalization fits better.

  • What happens to an existing term query on comments.author after you change comments from object to nested?
    It stops matching. Nested inner documents are hidden from ordinary search, so the query returns zero hits rather than an error — a silent regression. Every query, sort and aggregation touching that path must be rewritten to use `nested` (or the `nested` sort option and `nested` aggregation), and the index must be reindexed because the mapping change is not applied in place.
  • What does score_mode default to in a nested query, and when would you change it?
    It defaults to `avg`, averaging the scores of all matching children into the parent's score. Use `max` when the single best-matching child should decide relevance, `sum` when more matching children should mean a better parent, and `none` when the nested query is a pure filter and child scores should be ignored entirely.
  • How do you sort parent documents by a value inside their nested children?
    Use the `nested` option on the sort clause with the child `path` and an optional `filter` to restrict which children count, plus a `mode` such as `max` or `min` to collapse many child values into one sort value. Without the `nested` option the sort has no valid values to read.

saying these in an interview costs you the question

  • Says nested is just a naming convention for dotted fields
  • Claims a plain term query on a nested sub-field still works
  • Puts must_not inside the nested query to exclude documents
  • Thinks inner documents are separately indexable or updatable
  • Believes changing object to nested applies without a reindex

context