What does mapping a field as nested in Elasticsearch cost at index and update time?
answer
- structure is paid for in document count
- each entry is a real Lucene document
- documents are immutable, blocks are contiguous
- one field change rewrites the whole block
- two index settings cap fields and objects
basics
~20 sEvery nested object becomes its own Lucene document, so a document with 200 entries indexes as 201. Any update rewrites the whole block, deleting and re-adding every child, and dedicated limits cap nested fields per index and nested objects per document.
solid answer
~50 sNested is not free structure — it is document multiplication. A document with 200 nested entries produces 201 Lucene documents in one block, which inflates segment size, merge work, and the doc counts you see in `_cat/indices` (search totals still count only root documents, so the two disagree). Because Lucene documents are immutable and the block must stay contiguous, changing a single field means deleting and re-indexing the parent and every child — an update to one entry of a 500-entry array rewrites 501 documents. Two settings bound the damage: `index.mapping.nested_fields.limit` caps distinct nested mappings in an index (default 50) and `index.mapping.nested_objects.limit` caps nested objects in a single document (default 10000). Joining queries are also classed as expensive and can be disabled cluster-wide. The practical rule: nested suits small, bounded arrays that change with their parent.
code
bash · 5 lines# docs.count includes nested documents
GET _cat/indices/products?v&h=index,docs.count,store.size
# hits.total counts only root documents
GET products/_countgo deeper
Know that each nested object is stored as an extra hidden document, so nested arrays make an index physically much larger than the number of business records suggests.
Explain update amplification: documents are immutable and the block is contiguous, so touching one field rewrites the parent and every child. Name the two nested limit settings and what each bounds.
Bring the operational read: doc-count divergence in _cat/indices, merge and disk pressure, expensive-query gating, and the point at which a growing child array means the model must change rather than the limit.
Set the modelling rule for the organisation — bounded many-sides may nest, unbounded or high-churn ones may not — and tie it to shard-sizing assumptions so capacity planning is not silently invalidated.
## Document multiplication The cost model follows directly from the storage model. Each object in a nested array is indexed as its own Lucene document, written contiguously with the root in a block. A product with 200 variants is 201 Lucene documents. Ten million products with an average of 20 variants each is not 10 million documents but 210 million. The consequences are the ones you would expect from having ten or twenty times as many documents: larger segments, more postings, more merge work, more heap pressure during merges, and slower `forceMerge`. Shard sizing guidance that assumed one document per business entity becomes badly wrong. A reliable symptom: `_cat/indices` reports `docs.count` including nested documents, while a search's `hits.total` and the `_count` API report only root documents. An index that claims 210 million documents but returns 10 million on a `match_all` count is not broken — it has nested fields. ## Update amplification Lucene documents are immutable, and a nested block must remain contiguous and correctly ordered. So there is no such thing as updating one nested object. A partial update through `_update` reads the document's `_source`, applies the change, and re-indexes the entire document: the old block is marked deleted and a fresh block of root-plus-all-children is written. Appending one entry to a 500-entry array costs 501 deletions and 501 insertions. This is what makes nested wrong for high-churn children. A document whose array is appended to on every user action becomes a write amplifier: index throughput drops, deleted documents pile up until merges reclaim them, and disk usage inflates in the meantime. ## The limits Two index settings exist specifically to stop nested from destroying a cluster: - `index.mapping.nested_fields.limit` — the maximum number of distinct nested mappings an index may declare, default 50. It guards against a schema where dozens of fields are nested, each multiplying documents. - `index.mapping.nested_objects.limit` — the maximum number of nested objects a single document may contain, across all its nested fields, default 10000. Exceeding it rejects the indexing request rather than letting one pathological document explode a segment. Both are index settings, and both can be raised. Raising them is usually a signal to reconsider the model, not a fix: the limits exist because the underlying cost is real, and pushing them up moves the failure from a clean rejection to an out-of-memory event. ## Query and fetch cost Block join itself is efficient — finding the parent of a matching child is a forward scan within the block, not a term lookup. But the query still evaluates against many more documents than the root count suggests, and `inner_hits` adds per-hit fetch work that materialises child sources. Elasticsearch classes joining queries (nested and parent-child) among its *expensive queries*, which a cluster can disable wholesale with the `search.allow_expensive_queries` setting; on a cluster where that has been turned off, nested queries stop working entirely. Aggregating over nested fields also costs more than it looks: a `nested` aggregation descends into the child documents, so the bucket collection runs over the inflated document count, and getting back to parent-level metrics needs `reverse_nested`. ## Mitigations - **Bound the array.** Nested is a good fit when the many-side is small and has a natural ceiling — sizes of a garment, addresses of a customer. It is a bad fit when the many-side grows without limit. - **`include_in_parent` / `include_in_root`.** These mapping parameters additionally copy the child's values into the parent (or root) document so that ordinary, cheap queries can run on the flattened copy while correlated queries use the nested form. The cost is duplicated index data. - **Split the churn out.** If children change constantly and parents rarely, a parent-join field or separate child documents avoid rewriting the parent on every child change — at the price of slower queries and application-side stitching. - **Denormalize.** One document per child entry, with parent fields copied in, removes the multiplication problem entirely: the physical document count is the same as with nested, but there is no block rewrite and no special query syntax. ## How to check before shipping Measure with real data. Index a representative sample, compare `_cat/indices` `docs.count` against a `_count`, look at store size per business entity, and time a realistic update. The ratio you find — physical documents per entity, and documents rewritten per update — is the number that decides whether nested is affordable.
- Why does _cat/indices report far more documents than a match_all count on the same index?`_cat/indices` reports Lucene's physical document count, which includes hidden nested documents; search totals and `_count` report only root documents. A ratio of ten or twenty to one is normal for a nested-heavy mapping and is the quickest way to spot document multiplication.
- What does include_in_parent buy you, and what does it cost?It copies each nested child's field values into the parent document as well, so cheap ordinary queries and aggregations can run on the flattened copy while correlated queries use the nested form. The cost is a second copy of that data in the index — larger segments and more merge work — and the flattened copy still allows cross-object matching.
- A document's nested array keeps growing and indexing throughput has collapsed. What do you change?Stop rewriting the parent for every child. Either move the many-side out into its own documents (denormalized, or a parent-join child), or bound the array and archive older entries. Raising `index.mapping.nested_objects.limit` only postpones the failure, because the block rewrite cost grows with array size regardless of the cap.
saying these in an interview costs you the question
- Says nested only adds metadata, not extra documents
- Claims a single nested object can be updated in place
- Fixes a limit breach by raising it without remodelling
- Treats _cat/indices docs.count as the business entity count
- Uses nested for an unbounded, high-churn child array