skip to content

When is the Elasticsearch flattened field type the right choice for an object with unpredictable keys?

level: seniorimportance: nice to knowfreq 28%

answer

  1. one field, no matter how many keys
  2. aimed at mapping explosion
  3. everything inside becomes a keyword
  4. dotted paths still address sub-keys
  5. no analysis and no numeric ranges

basics

~20 s

Use flattened when an object's keys are open-ended and would otherwise create thousands of mappings. The whole object becomes one field whose leaf values are indexed as keywords — cheap and bounded, but with no per-key types and no full-text analysis.

solid answer

~50 s

`flattened` maps an entire JSON object as a **single** field. Every leaf value inside it is indexed as a keyword, and you query sub-keys by dotted path such as `attributes.color`. The win is mapping stability: user-supplied metadata, per-tenant custom attributes or arbitrary labels no longer add one mapping per key, so the field count stays constant no matter how varied the documents are. The price is that everything inside becomes a string. There is no numeric or date typing, so ranges are lexicographic rather than arithmetic; there is no analysis, so full-text matching inside the object does not work; only a limited set of query types is supported; and relevance scoring inside it is essentially flat. Use it for high-cardinality bookkeeping metadata you filter and facet on. Use explicit mappings for anything you actually rank, range over, or search as text.

code

json · 9 lines
json
PUT /events
{ "mappings": { "properties": {
    "message":    { "type": "text" },
    "attributes": { "type": "flattened" } } } }

PUT /events/_doc/1
{ "message": "ok",
  "attributes": { "color": "red", "weight": 42, "tenant": "acme" } }
// the mapping still contains exactly one attributes field

go deeper

for a junior

Recall that a flattened field maps a whole JSON object as one field whose values are all treated as keywords, and that you address sub-keys with a dotted path.

for a middle

Explain both halves of the trade: a constant mapping cost regardless of how many keys appear, paid for with string-only values, no analysis, lexicographic ranges and a reduced query surface.

for a senior

Diagnose when it is the right answer — open-ended metadata you only filter and facet on — and propose the hybrid design that promotes the few attributes needing real types while flattened absorbs the tail.

for a principal

Own the schema-governance angle: which teams may ship unbounded key sets, where the boundary between typed and flattened metadata sits, and how index templates carry that decision forward given field types cannot change in place.

## The problem it solves Some documents carry an object whose keys are not under your control: labels attached by users, per-tenant custom attributes, arbitrary tags from an ingestion pipeline. With ordinary `object` mapping, every new key that appears becomes a new field in the index mapping. Across a large corpus that grows without bound — every distinct key from every distinct tenant is a permanent entry in the mapping, the cluster state that carries it grows, and eventually indexing fails when the index's field limit is reached. This runaway growth is the classic mapping-explosion failure, and it is what `flattened` exists to prevent. ## What flattened does Declaring a field as `flattened` maps the *whole subtree* as one field: ```json "attributes": { "type": "flattened" } ``` At index time Elasticsearch walks the object and indexes each leaf value as a keyword, keeping track of which key it came from. The mapping never grows: no matter how many distinct keys the documents contain, the mapping still shows exactly one field. You then query in two ways. Against the field itself, a value matches if it appears **anywhere** in the object — useful for "does this document have this value at all?" Against a dotted path, the value must appear under that specific key: ```json { "query": { "term": { "attributes.color": "red" } } } { "query": { "term": { "attributes": "red" } } } ``` Doc values are maintained, so `terms` aggregations and sorting on a sub-key work, which is what makes the type genuinely useful for faceting over open-ended metadata. ## What you give up The simplification is not free, and articulating the cost is what separates a real answer from a memorized definition. - **Everything is a string.** A leaf value of `42` is indexed as the keyword `"42"`. A range query on it compares strings lexicographically, so `"9"` sorts after `"42"`. Dates are strings too, so no date math and no date histograms. - **No analysis.** Values are not tokenized, lowercased or stemmed, so full-text search inside a flattened object does not work the way a `text` field does. Matching is exact-term matching. - **Limited query surface.** Only a subset of query types is supported against flattened fields. Anything expecting a richly typed field — and any feature that needs per-field analysis — is unavailable. - **Flat relevance.** Because there are no per-field statistics of the usual kind, scoring inside a flattened field is not a useful ranking signal. Treat it as a filter surface, not a relevance surface. - **Object structure is not preserved for correlation.** Like ordinary object mapping, an array of sub-objects loses the association between keys within one element, so you cannot ask "the element whose colour is red *and* whose size is large". - **Depth is bounded.** There is a configurable limit on how deeply the object may nest, so pathologically deep structures still fail rather than being indexed silently. ## Choosing between the options Given an object with many keys, you have four realistic choices, and the interview answer is knowing when each applies: 1. **Explicit `object` mapping** with `dynamic: strict` — when the key set is known and stable. Best typing, best relevance, no surprises; requires you to own the schema. 2. **`flattened`** — when the keys are open-ended and you only filter and facet on the values. Constant mapping cost, string-only semantics. 3. **Key/value pair array** — model `[{"key": "color", "value": "red"}]` with typed `key` and `value` fields. Keeps proper types and, with the appropriate mapping, keeps key and value correlated, at the cost of clumsier queries. 4. **Store as an unindexed blob** — if you only display the metadata, keep it in `_source` and index nothing. ## A practical shape Most production designs are hybrids. Promote the handful of attributes that matter — the ones you range over, rank on, or search as text — into explicit, properly typed fields, and let `flattened` absorb the long tail nobody queries by name. That gives you correct semantics where they are needed and a bounded mapping everywhere else. One last operational note: because the field type cannot be changed in place, the choice between `object` and `flattened` is effectively permanent for a given index. On a data stream or a rolled-over time-series index that hurts less — new backing indices pick up the corrected template — but for a long-lived index it means a reindex. Decide before the first ingest, not after the field-limit error arrives during an incident.

  • Why can't you rely on relevance scoring inside a flattened field?
    All leaves live in one field with no per-key analysis or meaningful per-key statistics, so the usual term-weighting signals do not distinguish a match under one key from a match under another. Treat flattened as a filtering and faceting surface, and put anything you genuinely need to rank into a properly typed, analyzed field.
  • How would you handle an object where five keys matter and the rest are arbitrary?
    Map the five explicitly with correct types — numeric, date, keyword or text as appropriate — and let a separate flattened field absorb the remainder. You get real typing, ranges and relevance where they matter while the unpredictable tail costs exactly one mapping entry.
  • Does flattened solve the array-of-objects correlation problem?
    No. Like ordinary object mapping, it loses the association between keys inside a single array element, so a query cannot require that two values came from the same element. If cross-key correlation within an element is a requirement, flattened is the wrong tool for that part of the document.

An object mapping gives every attribute its own labelled drawer, which stops working when anyone can invent a new label. flattened throws all the notes into one drawer with the key written on each note: you can still find them, but they are all just paper now.

saying these in an interview costs you the question

  • Thinks flattened preserves numeric and date types
  • Believes it enables full-text search inside the object
  • Says it keeps array elements correlated
  • Uses it for fields that drive relevance ranking
  • Cannot name the mapping-growth problem it addresses

context