skip to content

How do you use GET _cluster/allocation/explain to diagnose an unassigned Elasticsearch shard?

level: seniorimportance: must knowfreq 62%

answer

  1. One API turns yellow into a named cause
  2. Two blocks: how it happened, and why not here
  3. Each rule reports YES, NO, or THROTTLE
  4. One verdict means later, not never
  5. Repeated failures eventually stop being retried

basics

~20 s

Call it with the index, shard number and whether it is the primary; the response gives unassigned_info explaining why the shard became unassigned, and per-node decider decisions where each NO or THROTTLE names the exact rule blocking allocation on that node.

solid answer

~40 s

`GET _cluster/allocation/explain` is the API that turns "my cluster is yellow" into a named cause. Pass a body naming `index`, `shard` and `primary`, or call it with no body to have Elasticsearch pick an arbitrary unassigned shard. Read the response in two parts. `unassigned_info` says **how** the shard became unassigned — `reason` values such as `NODE_LEFT`, `INDEX_CREATED`, `ALLOCATION_FAILED`, `CLUSTER_RECOVERED` — plus `failed_allocation_attempts` and the underlying exception. Then `node_allocation_decisions` lists every node with the **deciders** that ran there; each entry has a `decider` name, a `decision` of `YES`, `NO` or `THROTTLE`, and a human-readable explanation. Common blockers are `same_shard` (that node already holds a copy), `disk_threshold` (over the high watermark), `filter` or `awareness` (allocation rules no node satisfies), `data_tier`, and `max_retry` after repeated failures. THROTTLE means "later", not "never".

code

json · 5 lines
json
{
  "index": "orders-2026.08",
  "shard": 0,
  "primary": false
}

go deeper

for a junior

Know that Elasticsearch has an API that explains why a shard is unassigned, and that it is the first thing to run when the cluster turns yellow or red.

for a middle

Be able to read the output: the unassigned reason, and the per-node decider that returned NO. Name a few common deciders such as same_shard, disk_threshold and filter.

for a senior

Walk the full diagnosis loop from health to a specific fix, distinguish NO from THROTTLE, and explain when a missing primary means genuine data loss and a restore rather than a reroute.

for a principal

Turn the individual diagnosis into prevention: watermark headroom, decommission procedures that do not leave stale filters, tier and attribute hygiene, and runbooks that stop on-call engineers from reaching for accept_data_loss.

## When to reach for it Any time the cluster is yellow or red and you cannot immediately see why. Health tells you *that* a shard is unassigned; `_cat/shards` tells you *which*; the allocation explain API tells you *why*, in the allocator's own words. It is the single most useful answer to "how do you debug a red cluster", and interviewers ask it precisely because candidates who have never operated Elasticsearch reach for restarts instead. ## The request With no body, Elasticsearch explains an arbitrary unassigned shard — fine when there is only one problem. To target a specific shard: ```json { "index": "orders-2026.08", "shard": 0, "primary": false } ``` You can also explain a shard that *is* assigned, by adding `current_node`, to understand why it is not being moved or rebalanced somewhere else. ## Reading unassigned_info The `unassigned_info` block describes the event that left the shard without a home: - `INDEX_CREATED` / `CLUSTER_RECOVERED` / `NEW_INDEX_RESTORED` — normal startup states, usually transient. - `REPLICA_ADDED` — you increased `number_of_replicas` and there is nowhere for the new copy. - `NODE_LEFT` — the node holding the copy left; combined with delayed allocation this is often expected for a minute or so. - `ALLOCATION_FAILED` — an attempt was made and threw; `details` carries the exception, and `failed_allocation_attempts` counts how many times. - `PRIMARY_FAILED`, `FORCED_EMPTY_PRIMARY`, `MANUAL_ALLOCATION` — the less common cases. `last_allocation_status` and `at` (the timestamp) round out the picture. An `ALLOCATION_FAILED` with a corrupt-index exception in `details` is a very different problem from a `NODE_LEFT` on a cluster that is simply out of disk. ## Reading the deciders The allocator asks a chain of *deciders* whether a shard may go on each node, and the shard is allocated only if every decider says `YES`. The explain output shows each node with its verdict, so the diagnosis is usually a single line of text. The ones that come up in practice: - **same_shard** — the node already holds a copy of this shard. Expected on the primary's own node; if it is the reason on *every* node, you have more replicas configured than you have nodes. - **disk_threshold** — the node is above the high watermark (defaults: low 85%, high 90%, flood stage 95%), so no new shards may land there. The fix is disk or fewer shards, not more retries. At flood stage, Elasticsearch also applies a read-only-allow-delete block to indices on that node. - **filter** — index- or cluster-level allocation filtering (`index.routing.allocation.require.*`, `cluster.routing.allocation.exclude.*`) rules the node out. Very common after a decommission exclusion is left in place, or when an attribute like `box_type: warm` exists on no node. - **awareness** — allocation awareness would put too many copies in one zone. - **data_tier** — the index asks for a tier (`data_warm`, say) that no node in the cluster carries. - **max_retry** — allocation failed `index.allocation.max_retries` times (default 5) and the allocator has stopped trying. Fix the underlying cause, then `POST _cluster/reroute?retry_failed=true`. - **throttling** — a `THROTTLE` decision, meaning concurrent-recovery limits are already saturated. This resolves itself; it is not a fault. ## The no-valid-copy case For an unassigned **primary**, the output may instead report that it cannot allocate because no valid shard copy was found, listing each node's copy as stale or corrupt. That means every in-sync copy is gone — a genuine data-loss situation. The escape hatches are `_cluster/reroute` with `allocate_stale_primary` (promote an out-of-date copy, losing recent writes) or `allocate_empty_primary` (create an empty shard, losing everything in it); both require `accept_data_loss: true`. Restoring from a snapshot is almost always the better answer, and an interviewer will want to hear you say so before reaching for either command. ## A working diagnosis loop 1. `GET _cluster/health?level=indices` — which index. 2. `GET _cat/shards?v&h=index,shard,prirep,state,unassigned.reason` — which shard, and the coarse reason. 3. `GET _cluster/allocation/explain` on that shard — the decider that says NO. 4. Act on that decider specifically: free disk, remove the stale filter, add a node with the missing attribute or tier, lower `number_of_replicas`, or retry failed allocations. 5. Re-check health, and confirm the shard reaches `STARTED` rather than flapping between `INITIALIZING` and unassigned. The habit worth demonstrating is that every step narrows the problem and no step is "restart a node". Restarts frequently make things worse: they trigger further recoveries on a cluster that is already unable to place shards.

  • Allocation explain shows a NO from the max_retry decider. What does that mean and how do you clear it?
    Allocation was attempted and failed repeatedly — up to index.allocation.max_retries, which defaults to 5 — so the allocator stopped trying to avoid an endless loop. The retry counter is not the problem; the exception in unassigned_info.details is. Fix that cause, then run POST _cluster/reroute?retry_failed=true to reset the counter and trigger a fresh attempt.
  • What is the difference between a NO and a THROTTLE decision in Elasticsearch allocation explain output?
    NO means a rule forbids the shard on that node and it will never be placed there until something changes — disk freed, a filter removed, an attribute added. THROTTLE means the node is temporarily at its concurrent-recovery limit and the shard will be placed once in-flight recoveries finish. THROTTLE needs patience, not intervention.
  • Allocation explain says no valid shard copy can be found for an unassigned primary. What are your options?
    Every in-sync copy is gone, so this is real data loss for that shard. Prefer restoring the index from a snapshot. Failing that, _cluster/reroute offers allocate_stale_primary to promote an out-of-date copy — losing writes since it fell out of sync — or allocate_empty_primary to create an empty shard. Both demand accept_data_loss: true.

saying these in an interview costs you the question

  • Restarts nodes instead of asking the allocator why
  • Thinks a THROTTLE decision requires operator intervention
  • Reaches for allocate_empty_primary before considering a snapshot
  • Blames the retry limit rather than the underlying failure
  • Assumes unassigned always means the cluster is out of disk

context