How does Elasticsearch's join field work with has_child and has_parent queries?
answer
- one index, two kinds of document
- the child must land on the parent's shard
- routing by parent id, always
- one query returns parents, the mirror returns children
- global ordinals make refresh more expensive
basics
~10 sA join field declares parent-child relations inside one index. Children are independent documents that must be routed to their parent's shard; has_child returns parents whose children match, and has_parent returns children whose parent matches.
solid answer
~50 sA `join` field maps relation names, for example `{"relations": {"question": "answer"}}`, and each document sets that field to its role — `"question"` for a parent, or `{"name": "answer", "parent": "<parent id>"}` for a child. Parents and children are separate documents in the **same index**, and every child index or get request must pass `routing` equal to the parent's id so the two land on the same shard; the join is resolved shard-locally and never crosses shards. `has_child` takes a child `type` and query and returns matching **parents** (with optional `score_mode`, `min_children`, `max_children`, `inner_hits`); `has_parent` takes `parent_type` and returns matching **children**; `parent_id` fetches children of a known parent. The payoff is updating a child without touching the parent. The price is real: an index may declare only one join field, a child has exactly one parent, the join field maintains eagerly-built global ordinals rebuilt after each refresh, and joining queries are markedly slower than the equivalent nested query.
code
json · 10 lines{
"mappings": {
"properties": {
"doc_relation": {
"type": "join",
"relations": { "question": "answer" }
}
}
}
}go deeper
Know that a join field puts parents and children in one index as separate documents, and that has_child finds parents by their children while has_parent finds children by their parent.
Explain the routing requirement and why the join must be shard-local, the direction each query returns, and the one-join-field, one-parent-per-child constraints.
Bring the operating cost: eagerly built global ordinals making refresh heavier, shard skew from uneven fan-out, expensive-query gating, and when the update pattern actually justifies the join over denormalizing.
Own the call between join and denormalization at scale — read volume versus child churn, shard-balance risk from hot parents, and the migration path if the join stops paying for itself.
## Declaring the relation ```json { "mappings": { "properties": { "doc_relation": { "type": "join", "relations": { "question": "answer" } } }}} ``` One field, a map of parent name to child name (or to an array of child names when a parent has several kinds of child). An index may declare **one** join field only, and a child may have exactly one parent. Multi-level hierarchies are possible — a grandparent relation by chaining names — but each level compounds the cost and the documentation discourages it. ## Indexing A parent sets the field to its relation name; a child sets a name plus its parent's id, and the request must carry a `routing` parameter equal to that parent id: ``` PUT qa/_doc/1 { "text": "...", "doc_relation": "question" } PUT qa/_doc/2?routing=1 { "text": "...", "doc_relation": { "name": "answer", "parent": "1" } } ``` Routing is not optional and not a performance hint: the whole design rests on parent and children living on the same shard, so the join can be resolved locally with no cross-shard traffic. Get, update and delete requests for a child also need the same `routing`, and forgetting it is the classic "document not found on a document I just indexed" bug. A consequence of routing by parent id is skew: a parent with a million children puts all of them on one shard. Parent-child models with wildly uneven fan-out produce hot, oversized shards no matter how many shards the index has. ## Querying **`has_child`** — give it the child `type` and a `query`; it returns the **parents** whose children match. Options: `score_mode` (`none` by default, plus `avg`, `max`, `min`, `sum`) to roll child scores into the parent score, `min_children` and `max_children` to require a count range, and `inner_hits` to return the matching children alongside each parent. **`has_parent`** — give it the `parent_type` and a `query`; it returns the **children** whose parent matches. Its `score` parameter (a boolean, false by default) decides whether the parent's score propagates to the child. **`parent_id`** — given a relation `type` and a parent `id`, returns that parent's children directly. It is the cheap option when the parent is already known and no predicate on the parent is needed. Both joining queries work in filter context, and both are classed among Elasticsearch's expensive queries, so a cluster with `search.allow_expensive_queries` disabled will reject them. ## What it costs The join field is backed by **global ordinals** — a per-shard mapping of the field's values that lets the join be resolved quickly. For join fields these are built eagerly, which means the work happens at refresh time rather than at query time: refreshes on an index with a large join field get measurably more expensive, and a heavy indexing rate with a short refresh interval compounds that. Query time is worse than nested too, because the parent and child sets have to be looked up and intersected rather than walked as one contiguous block. In exchange you get true document independence: a child can be indexed, updated and deleted on its own, with no rewrite of the parent or of the other children. That is the one thing nested cannot do, and it is the only reason to accept the rest. ## When it earns its keep Parent-join fits when children vastly outnumber parents, when children change far more often than parents, and when the read pattern tolerates slower joining queries — a question with thousands of answers that arrive continuously, a product with a stream of price updates. It does not fit when the relation is small and static (nested is cheaper and faster), when search latency is critical at high query rates (denormalize), or when the parent and child would naturally live in different indices, since a join field cannot span indices. ## The practical guidance Elasticsearch's own advice is to prefer denormalization: copy the parent's fields onto the child documents and search the children directly, resolving the parent-level view with an aggregation on a parent id or a second query. That removes the join entirely, keeps every document independent, and scales with shard count instead of fighting it. Reach for `join` only when the update pattern genuinely forbids rewriting parents and the query volume is low enough to absorb the cost.
- What breaks if you index a child without the routing parameter?The child is routed by its own id and can land on a different shard from its parent. Joining queries resolve shard-locally, so that child becomes invisible to `has_child` and `has_parent`, and later get, update or delete calls that do pass the correct routing will not find it. There is no error at index time — the data is simply wrong.
- Why does a join field make refreshes more expensive?The join is resolved through global ordinals on the join field, and for join fields those are built eagerly rather than lazily at first query. Each refresh creates new segments and triggers the rebuild, so a high indexing rate with a short refresh interval pays that cost repeatedly.
- When is a nested field a better fit than a join field for the same relation?When the many-side is small, bounded, and changes together with its parent. Nested keeps everything in one contiguous block, so queries are considerably faster and no routing discipline is needed. The join field's only real advantage — updating a child without rewriting the parent — is worthless if the children rarely change.
- Can a join field relate documents that live in two different indices?No. Parent and child must be documents in the same index and on the same shard; the relation is resolved shard-locally. Relating separate indices means denormalizing the needed fields onto one side, or performing the join in the application with a second query.
saying these in an interview costs you the question
- Thinks parents and children can live in different indices
- Omits routing when indexing or fetching a child document
- Says has_child returns the matching child documents
- Believes parent-child is faster than nested for the same data
- Assumes an index can declare several join fields