How does Elasticsearch decide when to roll a data stream over to a new backing index?
answer
- Something must ask; nothing rolls itself
- Two families of condition, combined differently
- Any max fires, every min must permit
- One size condition counts all primaries
- Checked on a poll, so indices overshoot
basics
~10 sA rollover action lists conditions such as max_age, max_docs, max_primary_shard_size and max_size. ILM checks them on a periodic poll and rolls when any max condition is met, provided every min condition is also satisfied.
solid answer
~50 sRollover is not automatic — something has to ask for it. In practice the hot phase of an ILM policy carries a `rollover` action whose conditions ILM evaluates on its periodic poll of managed indices; you can also fire `POST /<stream>/_rollover` yourself. The conditions come in two families. The `max_*` conditions — `max_age`, `max_docs`, `max_size`, `max_primary_shard_size`, `max_primary_shard_docs` — are the triggers, and **any one of them** being satisfied is enough. The `min_*` conditions are guards: **all** of them must hold before a rollover happens at all, which is how you stop a quiet stream from rolling into a parade of near-empty indices. Two details bite people. `max_size` is the combined size of all primaries, whereas `max_primary_shard_size` is the largest single primary — the latter is what you want if the goal is a predictable shard size. And because evaluation happens on a poll interval, an index routinely overshoots its threshold before it actually rolls.
code
json · 17 lines// PUT _ilm/policy/logs-90d (hot phase only)
{
"policy": {
"phases": {
"hot": {
"actions": {
"rollover": {
"max_primary_shard_size": "50gb",
"max_age": "1d",
"min_docs": 1
},
"set_priority": { "priority": 100 }
}
}
}
}
}go deeper
Recall that rollover starts a new backing index and that conditions such as max_age, max_docs and a size limit control when it happens. Know that documents already written stay where they are.
Explain how the max and min families combine, the difference between total-primaries size and largest-primary size, and why a periodic evaluation means indices overshoot their thresholds.
Be ready to pick thresholds for a real ingest rate, explain the empty-index problem on quiet streams, and handle the backfill case where index creation date and data age disagree.
Own the standard shape across many streams: one policy family rather than dozens of bespoke ones, the cluster-wide cost of evaluating thousands of managed indices, and how rollover cadence sets the granularity of every downstream retention promise.
## Rollover is an action, not a background rule An index does not roll over because it got big. Something must ask. There are two askers in practice: the `rollover` action inside the hot phase of an ILM policy, which Elasticsearch evaluates when it polls managed indices, and an operator or script calling `POST /<stream>/_rollover` directly. Everything below describes the conditions both use. When a rollover fires on a data stream, Elasticsearch creates the next backing index generation from the current index template and makes it the write index. Nothing is copied and nothing moves: documents already written stay where they are, and the previous generation remains searchable, now permanently read-only for appends. ## The condition families **Trigger conditions** (`max_*`), any one of which is sufficient: - `max_age` — time since the index was created. `30d`, `1d`, `1h`. - `max_docs` — document count in the index's primaries. - `max_size` — total store size of **all** primary shards combined. - `max_primary_shard_size` — store size of the **largest single** primary shard. - `max_primary_shard_docs` — document count in the largest primary shard. **Guard conditions** (`min_age`, `min_docs`, `min_size`, `min_primary_shard_size`, `min_primary_shard_docs`), **all** of which must hold before any trigger counts. Their job is to suppress rollovers on low-traffic streams: with `max_age: 1d` alone, a stream that receives nothing overnight still produces a fresh empty index every day, and each empty index costs cluster state, shard allocations and file handles for nothing. `min_docs: 1` removes that whole class of waste. The combined rule is easy to state and easy to get backwards in an interview: **roll when (any max is met) AND (every min is met)**. ## max_size versus max_primary_shard_size This is the single most common confusion. Suppose an index has 5 primary shards and `max_size: 50gb`. That fires when the five primaries together reach 50 GB — roughly 10 GB per shard. Change the shard count to 1 and the same condition now produces a 50 GB shard. The condition's meaning therefore depends on a setting somewhere else, which is exactly what you do not want in a lifecycle policy. `max_primary_shard_size: 50gb` measures the largest primary directly, so it produces the same per-shard outcome regardless of how many shards the index has. If your goal is predictable shard size — and it usually is, since shard size drives search latency, recovery time and merge cost — this is the condition to reach for. `max_age` is normally paired with it as a ceiling, so that a stream which never reaches the size limit still rolls on a sensible cadence and retention stays granular. ## Why indices overshoot ILM does not watch indices continuously. It walks managed indices on a periodic poll (governed by the cluster setting `indices.lifecycle.poll_interval`, ten minutes by default in recent 8.x releases) and evaluates whatever step each index is waiting on. An index ingesting 5 GB per minute will therefore sail well past a 50 GB threshold before the next check happens. This is normal and is why thresholds are targets, not hard limits: leave headroom, and do not size disks on the assumption that no shard ever exceeds the configured maximum. The same poll explains the other frequent surprise: after you attach a policy or change one, nothing appears to happen for several minutes. Shortening the poll interval cluster-wide to chase responsiveness is rarely the right answer, because every managed index is re-evaluated on each pass and the cost scales with index count. ## Age is measured from index creation `max_age` counts from the creation of the backing index being evaluated, not from the age of the data in it. That matters when you bootstrap a stream by reindexing historical data: the new backing index is minutes old, however old its documents are, so an age-based rollover will not fire for a long time. The index setting `index.lifecycle.origination_date` (or `parse_origination_date`, which reads the date out of the index name) exists to override the reference point for indices whose data is older than their creation. ## Manual rollover `POST /<stream>/_rollover` with no body rolls unconditionally — useful in a runbook when a mapping fix must take effect immediately, or before a planned bulk load with different characteristics. Supplying a `conditions` block instead makes it a no-op unless the conditions are met, which is how external schedulers implement rollover without ILM. Note that a manual roll interacts with an ILM policy perfectly well; ILM simply sees a new write index next time it polls. ## A practical starting shape For most log streams a hot phase with `max_primary_shard_size` in the tens of gigabytes, `max_age` of a day or a few days, and `min_docs` of at least 1 covers it. Then let the warm, cold and delete phases handle everything downstream of the roll.
- Why prefer max_primary_shard_size over max_size in a rollover action?`max_size` measures all primaries added together, so its per-shard effect changes whenever the index's shard count changes — 50 GB across five shards is a 10 GB shard, across one shard it is a 50 GB shard. `max_primary_shard_size` measures the largest primary directly and therefore produces the same shard size regardless of shard count, which is what actually drives search latency, recovery time and merge cost.
- A stream is configured with max_age 1d but receives almost no traffic. What goes wrong, and how do you fix it?It produces a brand-new, nearly empty backing index every single day. Each one costs cluster state, shard allocations and file handles while holding almost nothing, and retention becomes a long tail of tiny indices. Adding a `min_docs` guard (or `min_primary_shard_size`) suppresses the roll until there is actually something in the index, so quiet periods do not manufacture indices.
- You reindex a year of history into a fresh data stream. Why does an age-based rollover not fire?`max_age` is measured from the backing index's creation time, not from the timestamps of the documents inside it, and the bootstrap index was created minutes ago. Either roll over manually between chunks of the backfill, or set `index.lifecycle.origination_date` (or enable `parse_origination_date`) so ILM ages the index from the date its data belongs to.
saying these in an interview costs you the question
- Thinks all rollover conditions must be met simultaneously
- Treats max_size and max_primary_shard_size as the same measurement
- Expects the index to never exceed its configured threshold
- Believes rollover happens automatically without a policy or request
- Says max_age is measured from the age of the documents