How does an Elasticsearch data stream differ from writing to a single regular index?
answer
- One name, many indices underneath
- Only the newest one takes writes
- Append-only: create, not index
- Backing indices are .ds-…-000001, hidden
- Template with data_stream block creates it
basics
~20 sA data stream is one name in front of an ordered set of hidden backing indices. Writes are append-only and always land in the newest backing index; searches and rollover target the stream name, not any single index.
solid answer
~50 sA data stream is the append-only pattern for time-series data such as logs, metrics and traces. You index and search using the stream name, but behind it Elasticsearch keeps a sequence of hidden backing indices named `.ds-<stream>-<date>-<generation>`. Only the newest generation is the **write index**; older ones stay searchable but accept no new documents. A rollover creates the next generation and moves the write pointer to it. You cannot create a data stream ad hoc: a composable index template matching the name must contain a `data_stream` object, and every document needs an `@timestamp` field. Writes must use the `create` op type, so indexing or updating by ID through the stream name is rejected. In exchange you get bounded index sizes, mapping and shard-count changes that take effect at the next rollover, and retention by dropping whole backing indices instead of running an expensive delete-by-query.
code
json · 19 lines// PUT _index_template/logs-app
{
"index_patterns": ["logs-app-*"],
"data_stream": {},
"priority": 500,
"template": {
"settings": {
"index.lifecycle.name": "logs-90d",
"index.number_of_shards": 2
},
"mappings": {
"properties": {
"@timestamp": { "type": "date" },
"message": { "type": "text" },
"service": { "type": "keyword" }
}
}
}
}go deeper
Recall the shape: one stream name, hidden .ds- backing indices, only the newest takes writes, and every document needs an @timestamp. Know that you append with create rather than PUT-by-ID.
Explain the mechanics: how the index template with a data_stream block creates the stream, why mapping and shard-count changes only affect the next generation, and how updates are done via _update_by_query or the backing index.
Be ready to justify the pattern operationally — dropping a backing index versus delete-by-query, evolving shard counts across generations, and the backfill trap where old timestamps still land in today's write index.
Own the decision of which datasets belong in streams at all, and the standardisation that follows: naming conventions, component templates shared across teams, and who owns the template when a mapping change breaks a downstream dashboard.
## What a data stream actually is A data stream is a named entry point for **append-only, timestamped** data: application logs, metrics, traces, audit events. It is not itself an index. Behind the name Elasticsearch keeps an ordered list of **backing indices** named `.ds-<stream-name>-<yyyy.MM.dd>-<generation>` — for example `.ds-logs-app-default-2026.08.20-000004`. The generation is a zero-padded counter incremented on every rollover. Exactly one backing index is the **write index** — always the newest generation. Every other backing index is still fully searchable but no longer receives new documents. Client code normally never mentions a backing index: it indexes into `logs-app-default` and searches `logs-app-default`, and the coordinating node fans the search out across all backing indices. ## Creating one A data stream cannot be conjured from a bare mapping. You first create a composable index template whose `index_patterns` match the stream name and that contains a `data_stream` object (usually empty). Indexing the first document into a matching name then auto-creates the stream and generation `000001`, or you create it explicitly with `PUT _data_stream/<name>`. Every backing index inherits mappings and settings from that template, which is the key operational property: **you change the template, not the stream**. Field additions, analyzer changes, shard counts and the ILM policy attached via `index.lifecycle.name` all apply to indices created from the next rollover onwards, while existing backing indices keep the definition they were born with. Additive mapping updates can also be pushed to all current backing indices with `PUT /<stream>/_mapping`. Each document must carry an `@timestamp` field mapped as `date` or `date_nanos`; if the template does not define it, Elasticsearch adds the mapping itself. ## Append-only write semantics Only the `create` op type is accepted through the stream name. In practice: - `POST /<stream>/_doc` (auto-generated ID) works. - `PUT /<stream>/_create/<id>` works — create with an explicit ID. - `PUT /<stream>/_doc/<id>` is **rejected**, because that request format means index-or-overwrite. - `_bulk` must use `create` actions; `index`, `update` and `delete` actions against the stream name are rejected. To change or remove a document you either run `_update_by_query` / `_delete_by_query` against the stream, or address the specific backing index that holds the document directly (optionally guarded with `if_seq_no` and `if_primary_term`). That restriction is exactly what makes the pattern cheap: an append never has to work out which generation already holds a given `_id`. ## Why this beats one perpetually growing index 1. **Retention becomes a metadata operation.** Expiring a day of logs means deleting a whole backing index — near-instant, and the disk space comes back immediately. Deleting the same rows from one huge index means a delete-by-query that only marks documents deleted; the space returns later, when merges rewrite the affected segments. 2. **Index-level settings can evolve.** Shard count is fixed for the life of an index, so a single ever-growing index locks yesterday's sizing decision in forever. With a stream, tomorrow's generation can have a different shard count, codec or refresh interval. 3. **Old data becomes immutable.** Once a backing index stops receiving writes it can be force-merged, made read-only, moved to cheaper hardware or converted to a searchable snapshot — none of which is possible for an index that is still being written to. 4. **Lifecycle automation has something to act on.** ILM operates per index; only a rolling series of indices gives it distinct units to age and delete. ## Gotchas worth knowing - **Routing ignores the timestamp.** In a standard data stream a document is written to the current write index no matter what its `@timestamp` says, so a backfilled three-month-old event lands in today's generation. Only a time-series data stream (`index.mode: time_series`) constrains documents to an accepted time range and rejects those outside it. - **Backing indices are hidden.** `GET _cat/indices` will not list them unless you ask for hidden indices; `GET _data_stream/<name>` is the friendlier view and shows the generation list and current write index. - **Deleting the stream deletes its data.** `DELETE _data_stream/<name>` removes every backing index with it. - **Rollover targets the stream.** `POST /<stream>/_rollover` creates the next generation; nothing rolls over on its own unless an ILM policy (or another scheduler) triggers it. The short version for an interview: a data stream is an append-only façade over a rolling series of indices, which converts "delete old rows" into "drop an index" and "change the mapping" into "edit the template and wait for the next rollover".
- How do you correct a single wrong document that is already in a data stream?Either run `_update_by_query` or `_delete_by_query` against the stream name with a query that isolates the document, or find which backing index holds it (the search hit's `_index`) and issue a normal update or delete against that index directly, optionally guarded with `if_seq_no` and `if_primary_term`. What you cannot do is update or delete by ID through the stream name itself.
- You add a new field to the template of a live data stream. When does it start applying?New backing indices created by the next rollover pick it up immediately, because each one is built from the template as it stands at rollover time. Existing backing indices keep their current mapping. If the change is purely additive, you can also push it to all current backing indices with `PUT /<stream>/_mapping`; anything non-additive needs a reindex.
- When would you still choose a plain index over a data stream?When the data is not time-series and not append-only: a product catalogue, a user directory, reference data that is updated and deleted by ID all day. Data streams forbid index-by-ID and update-by-ID, and their whole value — dropping whole generations for retention — is meaningless for a dataset with no time dimension.
It is like a bound volume of daily newspapers: you always write on today's edition, everything ever printed stays readable, and clearing archive space means throwing out whole back issues rather than erasing individual paragraphs.
saying these in an interview costs you the question
- Says a data stream is just an alias you can write to normally
- Thinks documents route to a backing index by their @timestamp
- Claims you can update or delete a document by ID via the stream name
- Believes editing the template retroactively changes existing backing indices
- Expects a data stream to appear without an index template declaring it