skip to content

Why does an Elasticsearch bool query on an array of objects match fields from different objects?

level: middleimportance: must knowfreq 78%

answer

  1. arrays of objects are not kept together
  2. Lucene documents have no inner structure
  3. each leaf path becomes one multi-valued field
  4. clauses are tested per document, not per entry
  5. the object type flattens; nested does not

basics

~20 s

By default Elasticsearch flattens an array of objects into independent multi-valued fields, losing which value belonged to which entry. A query for name alice and age 25 then matches even when those values came from different entries.

solid answer

~40 s

Lucene documents are flat: a document is a bag of field-to-values, with no notion of inner objects. When Elasticsearch indexes `users: [{name: alice, age: 30}, {name: john, age: 25}]` under a default `object` field, it stores `users.name: [alice, john]` and `users.age: [30, 25]` as two unrelated multi-valued fields. A `bool` with `must` clauses on both fields is evaluated against the **document**, not against an array entry, so `users.name: alice` AND `users.age: 25` both hold and the document matches. This is cross-object matching. `_source` still shows the original JSON, which is why the response looks structured and the bug looks impossible. The fixes are to map the field as `nested`, to denormalize into one document per entry, or to index a synthetic composite keyword such as `alice|30` when only one combination matters.

code

json · 7 lines
json
{
  "group": "fans",
  "users": [
    { "name": "alice", "age": 30 },
    { "name": "john",  "age": 25 }
  ]
}

go deeper

for a junior

Be able to state that Elasticsearch stores an array of objects as separate lists of values, so a query can match values that came from different entries in the array.

for a middle

Explain the internal representation — full dotted path becomes one multi-valued field per leaf — and why a bool must therefore evaluates against the whole document. Name nested and denormalization as the two real fixes.

for a senior

Show how you would diagnose this in production: reproduce with a two-entry document and an impossible combination, then choose between nested, one-document-per-entry, and a composite keyword based on array size and update rate.

for a principal

Own the modelling policy. Argue when the correlation is worth paying nested or denormalization costs across a whole index, and how you stop teams from shipping filters that silently over-match on flattened arrays.

## The setup Consider an index that leaves `users` as an ordinary object field — the default when Elasticsearch first sees an array of JSON objects — and a document: ```json { "group": "fans", "users": [ {"name": "alice", "age": 30}, {"name": "john", "age": 25} ] } ``` A `bool` query with two `must` clauses, `term users.name: alice` and `term users.age: 25`, matches this document even though no single user is both alice and 25. ## What Lucene actually stores Lucene has no concept of structure inside a document. A Lucene document is a flat set of field-to-values pairs. When Elasticsearch indexes an array of objects under an `object` field, it walks the JSON tree and appends every leaf value to the field named by its full dotted path. The document above becomes, internally, the equivalent of: ``` users.name: ["alice", "john"] users.age: [30, 25] ``` Two independent multi-valued fields. Each field remembers its own values — and, for analyzed text, the token positions inside that field — but nothing records that `alice` and `30` arrived in the same JSON object. Array boundaries are gone the moment the document is indexed. `_source` still holds the original JSON, which is why the API response looks perfectly structured, but `_source` is a stored blob that is returned after matching; it is not what queries run against. ## Why the query matches Query clauses are predicates over a document, not over an array element. `term users.name: alice` asks whether this document's `users.name` field contains the term alice — it does. `term users.age: 25` asks the same of `users.age` — it does too. `bool.must` requires both predicates to hold for the same document, and both hold. The more entries the array carries, the more spurious combinations become matchable, so a filter that reads as precise silently over-matches. The same flattening distorts analytics: a `terms` aggregation on `users.name` with an `avg` sub-aggregation on `users.age` averages *all* the document's ages inside *every* name bucket, because at document level every name is effectively paired with every age. ## When flattening is harmless Flattening only loses information when an array has more than one entry **and** a query correlates two or more fields from it. If the object always holds a single entry — an `address` with a `city` and a `zip` on a one-address document — or if queries only ever touch one field at a time, the flattened form answers correctly and costs far less than the alternatives. That is exactly why `object` is the default: most fields never need the extra machinery, and paying for it everywhere would be wasteful. ## The fixes, in the order to consider them **Denormalize.** Index one document per array entry — one document per user, carrying the group's fields alongside. Correlation becomes trivial because a single document now holds exactly one name and one age; queries are plain and fast and need no special syntax. The costs are duplicated parent fields, re-indexing every child when a parent field changes, and needing a `terms` aggregation on a group id (or a second query) to roll results back up to the group. **Map the field as `nested`.** This tells Elasticsearch to index each object in the array as its own hidden Lucene document, preserving the association, so a `nested` query can demand that all clauses match the *same* inner object. It is the direct fix, but it multiplies the physical document count and forces a full rewrite of the document and all of its children on any update. **Fake the correlation with a composite key.** When exactly one combination matters and it is a filter rather than open-ended search, index a synthetic keyword such as `users.name_age: ["alice|30", "john|25"]` and query it with one `term`. One field, one term, no ambiguity, none of the nested cost — but it only answers the combinations you anticipated at index time. ## What does not fix it The `flattened` field type does not preserve object identity either; it maps a whole subtree into a single field of key/value pairs so unpredictable keys cannot explode the mapping, and it is subject to the same cross-object matching. Adding more `must` clauses, tightening `minimum_should_match`, or switching from a `text` field to its `keyword` sub-field changes which terms match but never restores the lost association. ## Recognising it in the wild The signature is a filter returning documents that visibly fail it when you read `_source`, and only for documents whose arrays hold more than one entry. Reproduce it deliberately: index a two-entry document and query an impossible combination. If the document comes back, the field is flattened.

  • Does the flattening also affect aggregations, or only queries?
    It affects aggregations too. A `terms` aggregation on `users.name` with an `avg` sub-aggregation on `users.age` computes the average of every age in the document inside every name bucket, because at document level all names are paired with all ages. Correct per-entry analytics require a `nested` aggregation over a `nested` field, or one document per entry.
  • If _source shows the correct nested JSON, why can't the query use it?
    `_source` is a stored blob returned after matching; it is not indexed or searchable. Matching runs against the inverted index and doc values, which hold only the flattened field-to-terms mapping. Elasticsearch never re-parses `_source` to evaluate an ordinary query.
  • Would mapping the field as flattened solve the cross-object matching?
    No. The `flattened` type exists to stop unpredictable object keys from exploding the mapping — it indexes an entire subtree as one field of key/value pairs. Object identity is still lost, so the same false combinations match. Only `nested` (or denormalizing into separate documents) restores the association.

It is like filing two people's details as separate lists — first names in one column, ages in another — with no row numbers. You can prove the office contains an Alice and contains a 25-year-old, but not that they are the same person.

saying these in an interview costs you the question

  • Says bool must clauses are evaluated per array element
  • Claims _source structure is what queries actually match against
  • Thinks the analyzer or a keyword sub-field causes the mismatch
  • Believes adding more must clauses will narrow it correctly
  • Suggests flattened field type as the fix for cross-object matching

context