An index has stopped progressing through its ILM policy. How do you diagnose and unstick it?
answer
- One API tells you the exact step
- Check the service is even running
- A failed step carries the error text
- Fix cause first, then resume
- Retry only helps an errored index
basics
~20 sCall the ILM explain API for the index to see its current phase, action, step and any failure details, fix the underlying cause, then retry the policy for that index. Also confirm ILM itself is running cluster-wide.
solid answer
~50 sStart with `GET /<index>/_ilm/explain?human`. It reports whether the index is managed, which policy it is on, and the exact phase, action and step it sits in — plus `failed_step` and a `step_info` object carrying the error when something broke. That single response usually names the problem. Then check `GET _ilm/status`: if someone ran `POST _ilm/stop` for an upgrade and never restarted it, every index is frozen mid-policy and no explain output looks broken. Common causes are a shrink step that cannot get all shards onto one node, a force merge that ran out of disk, a searchable-snapshot step pointing at a missing repository, a delete phase waiting on an SLM policy that never ran, and — for legacy alias-managed indices — a missing or misdirected `index.lifecycle.rollover_alias`. Fix the cause, then `POST /<index>/_ilm/retry`, which resumes from the failed step rather than restarting the phase.
code
bash · 8 lines# Where is every backing index of this stream, and why?
curl "$ES/logs-app-default/_ilm/explain?human&pretty"
# Is the lifecycle service itself running?
curl "$ES/_ilm/status?pretty"
# After fixing the cause, resume the failed step
curl -XPOST "$ES/.ds-logs-app-default-2026.07.02-000019/_ilm/retry"go deeper
Recall that ILM exposes a per-index explain endpoint showing the current phase, action and step, and that there is a retry call for indices that errored.
Explain how to read the explain output — managed flag, step and step_time, failed_step and step_info — and why a policy edit does not affect an index already inside a phase.
Be ready to walk a real stall end to end: shrink allocation, disk-starved merges, a missing snapshot repository, a wait-for-snapshot guard, a legacy rollover alias, then fix and retry.
Own the detection rather than the firefight: alerting on failed steps and on step times exceeding phase expectations, and a maintenance protocol that guarantees a stopped lifecycle service is always restarted.
## The one API that matters `GET /<index>/_ilm/explain?human` is the whole diagnosis in one call. For each index it returns: - `managed` — whether ILM controls it at all. If this is `false`, no policy is attached and everything else is moot: the index was created before the template carried `index.lifecycle.name`, or from a template that does not set it. - `policy` — which policy applies. - `phase`, `action`, `step` — exactly where the index is. A healthy index waiting for its next transition sits in a normal waiting step such as `check-rollover-ready` or a phase's completion step. - `age`, `phase_time`, `action_time`, `step_time` — how long it has been there. A step time from three weeks ago is the tell. - `failed_step` and `step_info` — present when a step errored, with the underlying exception message. This is usually the answer. - `phase_execution` — the cached copy of the phase definition the index is executing, including a version. This explains the frequent complaint that editing a policy "did nothing": an index runs the phase definition it entered with, and picks up the new one at a phase boundary. Run it across a whole stream (`GET /<stream>/_ilm/explain`) to see whether one generation is stuck or all of them. ## Then check that ILM is running at all `GET _ilm/status` returns `RUNNING`, `STOPPING` or `STOPPED`. ILM is deliberately stoppable — it is standard practice to `POST _ilm/stop` before a cluster upgrade or a large maintenance window so lifecycle actions do not fight the operation. The failure mode is forgetting `POST _ilm/start` afterwards. Nothing errors, no index reports a failed step, and everything simply ages in place. Check this early; it costs one request and explains an entire cluster's worth of stalled indices at once. ## The recurring causes **Shrink cannot allocate.** The shrink step must first gather every primary shard of the source index onto a single node. If no node has enough free disk, or allocation filtering and awareness rules prevent the co-location, the index waits indefinitely in the allocation check. `GET _cluster/allocation/explain` tells you which constraint is refusing. **Force merge ran out of disk.** A merge needs headroom for the new segment beside the old ones. A node near a watermark fails the merge and the index lands in an error step with a disk-related message. **Searchable snapshot cannot run.** The `snapshot_repository` named in the action does not exist, the repository is unreachable, credentials expired, or the cluster's licence does not cover searchable snapshots. `step_info` names it. **The delete phase is waiting for a snapshot.** A `wait_for_snapshot` action holds the index until the named SLM policy has completed a snapshot after the index entered the delete phase. If that SLM policy is disabled, failing, or was renamed, the index waits forever — which is the action working exactly as designed, refusing to delete unbacked-up data. **Legacy alias-based ILM is misconfigured.** For indices managed through a write alias rather than a data stream, the rollover step needs `index.lifecycle.rollover_alias` set on the index, and that alias must actually have this index as its write index. A missing setting or an alias whose write index is another index produces an explicit error in `step_info` — and it is the classic symptom of someone creating the bootstrap index by hand and forgetting `"is_write_index": true`. **A write block from disk watermarks.** When a node crosses the flood-stage watermark, indices on it get a read-only-allow-delete block, which then fails steps that need to write. Freeing disk is only half the fix; the block must be cleared too. ## Recovering Once the cause is fixed, `POST /<index>/_ilm/retry` moves the index back to the failed step and re-executes it. It applies only to indices actually in an error state; it is not a general "push the policy along" button, and it does not restart the phase from the beginning. For a data stream, retry the specific backing index, not the stream name. If a policy edit is what you need to take effect, remember that the index runs its cached phase definition. Removing and reattaching `index.lifecycle.name` is a blunt way to force re-evaluation, and it is worth being deliberate about, because an index that re-enters the policy is aged from its original origin and can jump several phases at once. ## Preventing the next one The practical monitoring signal is any index whose `step_time` is older than the phase's expected duration, plus any index with a non-null `failed_step`. Both come straight out of the explain API, which makes a cheap scheduled check. The second signal is `_ilm/status` not being `RUNNING`. Together they catch the overwhelming majority of stalls before someone notices the disk filling up.
- You edited an ILM policy but a running index still behaves the old way. Why?Each managed index executes a cached copy of the phase definition it entered, visible as `phase_execution` in the explain output. Edits are picked up when the index moves into a new phase, not immediately, which keeps a half-completed phase self-consistent. If the change must apply now, you can detach and reattach the policy — but be aware the index is aged from its original origin and may skip forward several phases at once.
- Every index in the cluster stopped progressing on the same day and none shows a failed step. Where do you look?`GET _ilm/status`. The likely answer is that ILM was stopped with `POST _ilm/stop` before an upgrade or maintenance window and never restarted. Nothing errors in that state — indices simply sit in their current step and age in place — so per-index diagnosis reveals nothing and the cluster-level status is the only signal. `POST _ilm/start` resumes everything.
- Does the ILM retry API restart the phase from the beginning?No. It re-executes the step that failed, keeping the work already completed in that phase. It also only applies to an index actually in an error step; calling it on a healthy index that is merely waiting for its `min_age` does nothing, because there is nothing to retry. For a data stream you retry the specific backing index, not the stream name.
saying these in an interview costs you the question
- Deletes and recreates the index instead of reading the error
- Calls the retry API before fixing the underlying cause
- Never checks whether the lifecycle service is running
- Expects a policy edit to affect a mid-phase index immediately
- Blames ILM when a wait-for-snapshot guard is doing its job