skip to content

Indices and Mappings

How an index declares the shape of its documents: which fields exist, what type each one has, and what Elasticsearch may do with it. Almost every hard Elasticsearch problem — wrong results, an exploding mapping, a reindex you cannot avoid — traces back to a decision made here, which is why interviewers start with it.

part ofElasticsearchoverview, primer and where to startread it →
on this pageshow

explore

questions

30

How does an Elasticsearch data stream differ from writing to a single regular index?

level: juniorimportance: must knowfreq 72%

answer

  1. One name, many indices underneath
  2. Only the newest one takes writes
  3. Append-only: create, not index
  4. Backing indices are .ds-…-000001, hidden
  5. Template with data_stream block creates it

basics

~20 s

A data stream is one name in front of an ordered set of hidden backing indices. Writes are append-only and always land in the newest backing index; searches and rollover target the stream name, not any single index.

solid answer

~50 s

A data stream is the append-only pattern for time-series data such as logs, metrics and traces. You index and search using the stream name, but behind it Elasticsearch keeps a sequence of hidden backing indices named `.ds-<stream>-<date>-<generation>`. Only the newest generation is the **write index**; older ones stay searchable but accept no new documents. A rollover creates the next generation and moves the write pointer to it. You cannot create a data stream ad hoc: a composable index template matching the name must contain a `data_stream` object, and every document needs an `@timestamp` field. Writes must use the `create` op type, so indexing or updating by ID through the stream name is rejected. In exchange you get bounded index sizes, mapping and shard-count changes that take effect at the next rollover, and retention by dropping whole backing indices instead of running an expensive delete-by-query.

code

json · 19 lines
json
// PUT _index_template/logs-app
{
  "index_patterns": ["logs-app-*"],
  "data_stream": {},
  "priority": 500,
  "template": {
    "settings": {
      "index.lifecycle.name": "logs-90d",
      "index.number_of_shards": 2
    },
    "mappings": {
      "properties": {
        "@timestamp": { "type": "date" },
        "message":    { "type": "text" },
        "service":    { "type": "keyword" }
      }
    }
  }
}

go deeper

for a junior

Recall the shape: one stream name, hidden .ds- backing indices, only the newest takes writes, and every document needs an @timestamp. Know that you append with create rather than PUT-by-ID.

for a middle

Explain the mechanics: how the index template with a data_stream block creates the stream, why mapping and shard-count changes only affect the next generation, and how updates are done via _update_by_query or the backing index.

for a senior

Be ready to justify the pattern operationally — dropping a backing index versus delete-by-query, evolving shard counts across generations, and the backfill trap where old timestamps still land in today's write index.

for a principal

Own the decision of which datasets belong in streams at all, and the standardisation that follows: naming conventions, component templates shared across teams, and who owns the template when a mapping change breaks a downstream dashboard.

## What a data stream actually is A data stream is a named entry point for **append-only, timestamped** data: application logs, metrics, traces, audit events. It is not itself an index. Behind the name Elasticsearch keeps an ordered list of **backing indices** named `.ds-<stream-name>-<yyyy.MM.dd>-<generation>` — for example `.ds-logs-app-default-2026.08.20-000004`. The generation is a zero-padded counter incremented on every rollover. Exactly one backing index is the **write index** — always the newest generation. Every other backing index is still fully searchable but no longer receives new documents. Client code normally never mentions a backing index: it indexes into `logs-app-default` and searches `logs-app-default`, and the coordinating node fans the search out across all backing indices. ## Creating one A data stream cannot be conjured from a bare mapping. You first create a composable index template whose `index_patterns` match the stream name and that contains a `data_stream` object (usually empty). Indexing the first document into a matching name then auto-creates the stream and generation `000001`, or you create it explicitly with `PUT _data_stream/<name>`. Every backing index inherits mappings and settings from that template, which is the key operational property: **you change the template, not the stream**. Field additions, analyzer changes, shard counts and the ILM policy attached via `index.lifecycle.name` all apply to indices created from the next rollover onwards, while existing backing indices keep the definition they were born with. Additive mapping updates can also be pushed to all current backing indices with `PUT /<stream>/_mapping`. Each document must carry an `@timestamp` field mapped as `date` or `date_nanos`; if the template does not define it, Elasticsearch adds the mapping itself. ## Append-only write semantics Only the `create` op type is accepted through the stream name. In practice: - `POST /<stream>/_doc` (auto-generated ID) works. - `PUT /<stream>/_create/<id>` works — create with an explicit ID. - `PUT /<stream>/_doc/<id>` is **rejected**, because that request format means index-or-overwrite. - `_bulk` must use `create` actions; `index`, `update` and `delete` actions against the stream name are rejected. To change or remove a document you either run `_update_by_query` / `_delete_by_query` against the stream, or address the specific backing index that holds the document directly (optionally guarded with `if_seq_no` and `if_primary_term`). That restriction is exactly what makes the pattern cheap: an append never has to work out which generation already holds a given `_id`. ## Why this beats one perpetually growing index 1. **Retention becomes a metadata operation.** Expiring a day of logs means deleting a whole backing index — near-instant, and the disk space comes back immediately. Deleting the same rows from one huge index means a delete-by-query that only marks documents deleted; the space returns later, when merges rewrite the affected segments. 2. **Index-level settings can evolve.** Shard count is fixed for the life of an index, so a single ever-growing index locks yesterday's sizing decision in forever. With a stream, tomorrow's generation can have a different shard count, codec or refresh interval. 3. **Old data becomes immutable.** Once a backing index stops receiving writes it can be force-merged, made read-only, moved to cheaper hardware or converted to a searchable snapshot — none of which is possible for an index that is still being written to. 4. **Lifecycle automation has something to act on.** ILM operates per index; only a rolling series of indices gives it distinct units to age and delete. ## Gotchas worth knowing - **Routing ignores the timestamp.** In a standard data stream a document is written to the current write index no matter what its `@timestamp` says, so a backfilled three-month-old event lands in today's generation. Only a time-series data stream (`index.mode: time_series`) constrains documents to an accepted time range and rejects those outside it. - **Backing indices are hidden.** `GET _cat/indices` will not list them unless you ask for hidden indices; `GET _data_stream/<name>` is the friendlier view and shows the generation list and current write index. - **Deleting the stream deletes its data.** `DELETE _data_stream/<name>` removes every backing index with it. - **Rollover targets the stream.** `POST /<stream>/_rollover` creates the next generation; nothing rolls over on its own unless an ILM policy (or another scheduler) triggers it. The short version for an interview: a data stream is an append-only façade over a rolling series of indices, which converts "delete old rows" into "drop an index" and "change the mapping" into "edit the template and wait for the next rollover".

  • How do you correct a single wrong document that is already in a data stream?
    Either run `_update_by_query` or `_delete_by_query` against the stream name with a query that isolates the document, or find which backing index holds it (the search hit's `_index`) and issue a normal update or delete against that index directly, optionally guarded with `if_seq_no` and `if_primary_term`. What you cannot do is update or delete by ID through the stream name itself.
  • You add a new field to the template of a live data stream. When does it start applying?
    New backing indices created by the next rollover pick it up immediately, because each one is built from the template as it stands at rollover time. Existing backing indices keep their current mapping. If the change is purely additive, you can also push it to all current backing indices with `PUT /<stream>/_mapping`; anything non-additive needs a reindex.
  • When would you still choose a plain index over a data stream?
    When the data is not time-series and not append-only: a product catalogue, a user directory, reference data that is updated and deleted by ID all day. Data streams forbid index-by-ID and update-by-ID, and their whole value — dropping whole generations for retention — is meaningless for a dataset with no time dimension.

It is like a bound volume of daily newspapers: you always write on today's edition, everything ever printed stays readable, and clearing archive space means throwing out whole back issues rather than erasing individual paragraphs.

saying these in an interview costs you the question

  • Says a data stream is just an alias you can write to normally
  • Thinks documents route to a backing index by their @timestamp
  • Claims you can update or delete a document by ID via the stream name
  • Believes editing the template retroactively changes existing backing indices
  • Expects a data stream to appear without an index template declaring it

context

open as a page

In an Elasticsearch mapping, what do the dynamic values true, runtime, false and strict each do?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Elasticsearch's dynamic setting decides what happens to an unmapped field: true indexes it and adds it to the mapping, runtime adds it as a query-time field only, false ignores it but keeps it in _source, strict rejects the document.

open as a page

In an Elasticsearch mapping, how do the text and keyword field types differ?

level: juniorimportance: must knowfreq 88%

basics

~10 s

text is analyzed into tokens for full-text search and cannot be sorted or aggregated by default. keyword stores the value verbatim as one term, which is what filters, sorts and aggregations need.

open as a page

In Elasticsearch, what does an index template configure and when is it applied?

level: juniorimportance: must knowfreq 68%

basics

~20 s

An index template attaches settings, mappings and aliases to any index whose name matches its index_patterns. It is applied once, at index creation time; indices that already exist keep whatever configuration they were created with.

open as a page

Why do Elasticsearch applications query an index alias instead of the concrete index name?

level: juniorimportance: must knowfreq 72%

basics

~20 s

An alias is a second name that points at one or more indices. Applications talk to the alias so the physical index behind it can be rebuilt, versioned and swapped atomically, with no client change and no downtime.

open as a page

How does Elasticsearch decide when to roll a data stream over to a new backing index?

level: middleimportance: must knowfreq 64%

basics

~10 s

A rollover action lists conditions such as max_age, max_docs, max_primary_shard_size and max_size. ILM checks them on a periodic poll and rolls when any max condition is met, provided every min condition is also satisfied.

open as a page

Why is a string field in Elasticsearch usually mapped as text with a keyword sub-field?

level: middleimportance: must knowfreq 70%

basics

~20 s

A multi-field indexes the same source value more than once under different mappings. The text half serves match queries, while title.keyword keeps the value verbatim for exact filters, sorting and aggregations — two capabilities one mapping cannot provide.

open as a page

When several Elasticsearch index templates match a new index name, which one wins?

level: middleimportance: must knowfreq 62%

basics

~20 s

Exactly one composable index template is applied: the matching one with the highest priority. Elasticsearch refuses to store two templates whose index_patterns overlap at the same priority, so ties are prevented when you save the template rather than resolved at index creation.

open as a page

Why can't you change an existing field's type in an Elasticsearch mapping?

level: middleimportance: must knowfreq 78%

basics

~20 s

A field's type and analyzer are baked into immutable Lucene segments when each document is indexed, so already-written data cannot be reinterpreted under a new type. Adding new fields is fine; changing an existing one requires a new index and a reindex.

open as a page

An Elasticsearch index rejects writes with "Limit of total fields [1000] has been exceeded" — what happened and how do you fix it?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Dynamic mapping turned unbounded data keys into permanent mapping entries and hit index.mapping.total_fields.limit, which defaults to 1000. Raising the limit is a stopgap; the fix is to stop the key space from becoming fields, using strict or scoped dynamic settings, templates, or a key/value structure.

open as a page

How do Elasticsearch's hot, warm, cold and frozen data tiers differ in what they store and cost?

level: middleimportance: should knowfreq 58%

basics

~20 s

Hot holds the actively written, most-queried indices on the fastest hardware. Warm holds recent read-only data on cheaper nodes, cold trades replicas for snapshot-backed storage, and frozen keeps only a local cache while the data itself lives in object storage.

open as a page

How do dynamic_templates control the types Elasticsearch assigns to newly seen fields?

level: middleimportance: should knowfreq 48%

basics

~20 s

dynamic_templates is an ordered list of rules in an Elasticsearch mapping. Each rule matches unmapped fields by detected JSON type, field name, or dotted path, and supplies the mapping to use. The first matching rule wins.

open as a page

What is an Elasticsearch runtime field, and when do you use one instead of reindexing?

level: middleimportance: should knowfreq 50%

basics

~20 s

A runtime field is defined in the mapping or in a search request and evaluated per document at query time, usually from _source, with nothing written to the index. It gives schema-on-read: add, change or drop it instantly, and pay the cost on every query instead.

open as a page

Under Elasticsearch dynamic mapping, what type is created for a JSON string like "19.99", and why?

level: middleimportance: should knowfreq 54%

basics

~20 s

It becomes a text field with a keyword sub-field, because Elasticsearch's numeric_detection defaults to off. Date detection is on by default, so a date-looking string would instead be mapped as date — often the more dangerous surprise.

open as a page

What does Elasticsearch do with a keyword value longer than the field's ignore_above setting?

level: middleimportance: should knowfreq 46%

basics

~20 s

The value is skipped for that field: it is neither indexed nor written to doc_values, so term queries, sorts and aggregations never see it. It still appears in _source, which makes the loss easy to miss.

open as a page

In what order does Elasticsearch merge composed_of component templates into an index's configuration?

level: middleimportance: should knowfreq 52%

basics

~20 s

Component templates named in composed_of are merged in array order, each overriding the previous one on conflicts. The index template's own template block is applied last and beats them all, and anything supplied in the create-index request overrides even that.

open as a page

How do you verify which Elasticsearch template will apply to an index before it is created?

level: middleimportance: should knowfreq 44%

basics

~20 s

Call POST _index_template/_simulate_index/<name>. Elasticsearch resolves that name against every stored template and returns the settings, mappings and aliases the index would actually receive, plus an overlapping list naming the templates that matched but lost on priority.

open as a page

When does _update_by_query suffice for a mapping change and when do you need _reindex?

level: middleimportance: should knowfreq 55%

basics

~20 s

Use _update_by_query when the mapping change was legal in place — a new field or sub-field — and existing documents just need rewriting to pick it up. Use _reindex into a new index whenever the mapping itself could not be updated, such as a field type change.

open as a page

Why does the ILM forcemerge action belong in the warm phase rather than on a live write index?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Force merge rewrites an index into very few segments, which is expensive and only pays off once the data stops changing. A single huge segment is never merged again, so deleted documents inside it are never reclaimed while writes continue.

open as a page

An index has stopped progressing through its ILM policy. How do you diagnose and unstick it?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Call the ILM explain API for the index to see its current phase, action, step and any failure details, fix the underlying cause, then retry the policy for that index. Also confirm ILM itself is running cluster-wide.

open as a page

Should an Elasticsearch field holding numeric order IDs be mapped as keyword or as a numeric type?

level: seniorimportance: should knowfreq 38%

basics

~10 s

Map it as keyword. Numeric types are optimized for range queries, keyword for exact-term lookups, and an identifier is only ever matched exactly. Numeric mapping also destroys leading zeros and invites meaningless arithmetic aggregations.

open as a page

When would you set index: false or doc_values: false on an Elasticsearch field?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Set index: false on fields you only display or aggregate on, never search by; set doc_values: false on fields you only search or retrieve, never sort, aggregate or script on. Each removes one index structure to save disk and indexing time.

open as a page

How do legacy _template definitions interact with composable _index_template definitions?

level: seniorimportance: should knowfreq 33%

basics

~20 s

Composable templates win outright: if any composable index template matches a new index name, Elasticsearch ignores every legacy _template for that index. Legacy templates are deprecated, merge all matches by order, and cannot back data streams.

open as a page

How do you run _reindex over a 500-million-document index without destabilising the cluster?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Pre-create the destination, then run the reindex asynchronously with wait_for_completion=false, parallelise it with slices, throttle it with requests_per_second, allow conflicts to proceed, and monitor the task id — rethrottling or cancelling it if live traffic suffers.

open as a page

How do you change an Elasticsearch index mapping with zero downtime while writes keep arriving?

level: seniorimportance: should knowfreq 65%

basics

~20 s

Build a new index with the corrected mapping beside the old one, reindex into it, keep it current with dual writes or repeated timestamp catch-up passes, verify counts and queries, then move the alias in one atomic call and keep the old index for rollback.

open as a page

How would you design a data stream and ILM policy to keep 400 GB of logs a day for 13 months on a fixed budget?

level: principalimportance: should knowfreq 38%

basics

~20 s

Work backwards from query patterns: keep a few days hot on fast disk, weeks warm and force-merged, months as snapshot-backed cold or frozen indices where object storage carries the bulk, and delete on schedule with a snapshot guard.

open as a page

How would you design mappings for a multi-tenant Elasticsearch log platform where tenants send arbitrary JSON keys?

level: principalimportance: should knowfreq 30%

basics

~20 s

Split the document into a declared core that is strictly mapped and a free-form subtree that is not allowed to mint unbounded fields. Isolate tenants across indices so one tenant's key space cannot poison a shared mapping, and enforce field-count budgets with alerting.

open as a page

How would you structure Elasticsearch component and index templates for many teams and index families?

level: principalimportance: should knowfreq 26%

basics

~20 s

Layer a small number of semantic component templates - platform defaults, shared schema, per-team fields, an override slot - and give each family one narrow index template with a priority from a documented band. Keep all of it in version control and gate changes on simulated output.

open as a page

When is the Elasticsearch flattened field type the right choice for an object with unpredictable keys?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Use flattened when an object's keys are open-ended and would otherwise create thousands of mappings. The whole object becomes one field whose leaf values are indexed as keywords — cheap and bounded, but with no per-key types and no full-text analysis.

open as a page

How would you make Elasticsearch mapping changes a routine operation rather than a multi-hour incident?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Treat every index as a disposable, versioned artifact behind an alias, keep the authoritative data outside Elasticsearch so any index can be rebuilt from source, automate create-reindex-verify-swap as one pipeline, and budget standing capacity for rebuilds.

open as a page