What does inner_hits add to an Elasticsearch nested query's results?
answer
- the response otherwise shows every child
- you need to know which one matched
- a miniature search response per hit
- carries field and array offset metadata
- only the top few come back by default
basics
~20 sinner_hits attaches to each matching document the specific nested or child documents that caused the match, with their own source, scores, sorting and highlighting. Without it the hit's _source contains every child and you cannot tell which one matched.
solid answer
~50 sA `nested` query returns root documents, and the hit's `_source` is the whole original JSON — every comment, every variant, matched or not. `inner_hits` asks Elasticsearch to also return, per hit, the inner documents that actually matched, as a small nested search response with its own `hits`, `_score`, and optional `highlight`. Each inner hit carries `_nested` metadata giving the `field` and the array `offset`, so you can map it back to the right element of the original array. Useful options are `size` (only the top few come back by default — 3), `from`, `sort`, `_source` filtering, and `name` when a search contains more than one `inner_hits` block and the responses must be distinguishable. It also works on `has_child` and `has_parent` queries. It is not free: retrieving inner hits is extra per-hit fetch work on top of the main query.
code
json · 21 lines{
"query": {
"nested": {
"path": "variants",
"query": {
"bool": {
"must": [
{ "term": { "variants.color": "red" } },
{ "term": { "variants.in_stock": true } }
]
}
},
"inner_hits": {
"name": "matched_variants",
"size": 10,
"sort": [ { "variants.price": "asc" } ],
"_source": [ "variants.sku", "variants.price" ]
}
}
}
}go deeper
Recall that inner_hits is an option on a nested query that returns the child objects that matched, so the UI can display the right one instead of the whole document.
Explain the response shape — a miniature search response with _nested field and offset — and the size default that silently truncates when a document has many matching children.
Weigh its fetch-phase cost against page size and inner size, know it also serves has_child and has_parent, and reach for a nested aggregation when only counts are needed.
Decide where matched-child retrieval belongs at all: inline in the search response, in a second call, or designed away by denormalizing so the matched entity is the document being returned.
## The problem it solves A `nested` query answers a parent-level question — "which products have a variant in red, size M, in stock?" — and returns products. The hit's `_source` is the whole product document, all fifty variants included. The UI, however, usually needs to show the *matching* variant: its price, its SKU, its image. Nothing in an ordinary response tells you which of the fifty caused the match, and re-evaluating the predicate in application code duplicates the query logic and gets it wrong the moment analysis is involved. `inner_hits` closes that gap. Adding it to the `nested` query makes Elasticsearch run a small secondary retrieval, per returned hit, over that hit's inner documents, and return them inline. ## The response shape Each top-level hit gains an `inner_hits` object keyed by the nested path (or by the `name` you supply), containing a miniature search response: `total`, and a `hits` array whose entries have `_source` (the inner object as it appeared in the original JSON), `_score`, and a `_nested` block naming the `field` and the zero-based `offset` of that object within the array. The offset is what lets a client line the inner hit up with the element it came from, which matters when two variants share the same field values. ## Options worth knowing - `size` — how many inner documents to return per hit. Only the top few come back by default (3), which quietly truncates results when a document has many matching children; raise it deliberately, and remember every extra inner hit is more work and more payload. - `from` — offset within the inner hits, for paging inside a single parent. - `sort` — order the inner hits by a child field rather than by relevance; a variant list is usually more useful sorted by price than by score. - `_source` — include or exclude child fields, keeping payloads small when children are fat. - `highlight` — highlight on the child fields, which is the only way to get a snippet from the matched child rather than from the whole parent. - `name` — required in practice when a query has more than one `inner_hits` block (two nested queries, or a nested plus a child query); without distinct names the responses collide. ## Beyond nested `inner_hits` is not nested-specific. On a `has_child` query it returns the matching child documents alongside each parent hit; on a `has_parent` query it returns the matched parent alongside each child hit. That is often the only reason a parent-child model is bearable in a UI — otherwise a second round trip is needed just to fetch the related documents. ## Cost Inner hits are retrieved during the fetch phase, after the main query has selected the top documents. That bounds the work by page size rather than by the whole result set, but it is still real: for each returned hit, Elasticsearch runs the inner query against that document's children and materialises their source. A page of 50 hits with `size: 20` inner hits materialises up to a thousand child sources. Highlighting inner hits multiplies it again. If a client only needs a count of matching children rather than their content, it is cheaper to omit `inner_hits` and use a `nested` aggregation, or to accept the parent-level answer. ## Practical notes - The scores you see on inner hits are the child scores that `score_mode` rolled up into the parent's `_score`; they are useful for debugging why a parent scored the way it did. - Inner hits reflect the inner query, not the whole search. A filter applied outside the `nested` query does not restrict which children come back. - Because the inner `_source` is reconstructed from the parent's stored `_source`, disabling `_source` on the index removes inner hit content too.
- How do you tell which element of the original array an inner hit came from?Each inner hit carries a `_nested` block with the `field` (the nested path) and a zero-based `offset` into that array in the original `_source`. Use the offset rather than matching on field values, since two children can be identical in the fields you selected.
- Does inner_hits work with parent-child queries as well?Yes. On `has_child` it returns the matching child documents next to each parent hit; on `has_parent` it returns the matched parent next to each child hit. Without it, a parent-join model needs a second query just to fetch the related documents for display.
- When would you deliberately not ask for inner_hits?When the client only needs the parent, or only a count. Inner hits are fetched per returned hit and materialise child sources — a large page with a large inner size can dominate the response. A `nested` aggregation gives counts far more cheaply than retrieving every matching child.
saying these in an interview costs you the question
- Thinks the hit's _source already excludes non-matching children
- Assumes all matching children are returned by default
- Re-runs the predicate in application code to find the match
- Believes inner_hits changes which parent documents match
- Uses two inner_hits blocks without naming them