skip to content

How do you restart an Elasticsearch data node for maintenance without triggering a full shard reallocation?

level: seniorimportance: should knowfreq 55%

answer

  1. A returning node brings its own shard copies
  2. One setting pauses replica allocation entirely
  3. There is a built-in grace period, about a minute
  4. Leaving versus returning want opposite treatment
  5. Drain first when the node is never coming back

basics

~20 s

Set cluster.routing.allocation.enable to primaries so replicas of the departing node are not rebuilt elsewhere, stop non-essential indexing and flush, restart the node, then set the setting back to all and wait for green. Delayed allocation covers short absences automatically.

solid answer

~50 s

When a data node leaves, its replicas go unassigned and Elasticsearch would normally rebuild them elsewhere — copying gigabytes for a node that is coming back in two minutes. Two mechanisms prevent that. **Delayed allocation** (`index.unassigned.node_left.delayed_timeout`, default `1m`) makes the allocator wait before reassigning replicas of a departed node; raising it to a few minutes covers a planned restart. For an explicit rolling restart, set `cluster.routing.allocation.enable: primaries` first, so no replica allocation happens at all while the node is out, stop non-essential indexing and call `POST /_flush` to make recovery cheaper, restart the node, then set the setting back to `all` and wait for green before moving to the next node. That is the opposite of a **permanent decommission**, where you want the data moved: set `cluster.routing.allocation.exclude._name` to the node and wait for it to drain to zero shards before shutting it down.

code

json · 5 lines
json
{
  "persistent": {
    "cluster.routing.allocation.enable": "primaries"
  }
}

go deeper

for a junior

Know that stopping a data node makes its shards unassigned and that Elasticsearch has settings to stop it rebuilding copies for a node that is coming right back.

for a middle

Describe the rolling-restart sequence in order and explain what allocation enable set to primaries actually prevents, plus the default grace period for a departed node's replicas.

for a senior

Distinguish a restart from a decommission and pick the right mechanism for each, reason about recovery throttles during a maintenance window, and name the failure modes of doing it wrong.

for a principal

Own the maintenance runbook: how many nodes may be out at once given the replica count and zone layout, what the throughput budget for recovery is, and how the procedure is automated so a forgotten setting cannot leave the fleet degraded.

## What happens when a node leaves The master notices the node is gone, marks its shards unassigned, and — for any primary it held — promotes an in-sync replica elsewhere so the index stays writable. Then the allocator wants to restore the configured replica count, which means building fresh copies on the remaining nodes. For a node that is genuinely dead, that is exactly right. For a node you are rebooting for a kernel patch, it is a large, pointless data movement that will be undone minutes later when the node returns and its own copies are found to be redundant. ## Delayed allocation `index.unassigned.node_left.delayed_timeout` is a per-index dynamic setting, default `1m`, that applies specifically to shards left unassigned because their node left. During the delay, the shards show as `unassigned` with `unassigned.reason: NODE_LEFT` and the cluster is yellow, but nothing is copied. If the node rejoins inside the window, it brings its shard copies with it and recovery is cheap: because replicas track the primary's operation history through retention leases, a returning replica usually replays only the operations it missed rather than copying whole segment files. Raising this to, say, `5m` across your indices before a maintenance window is the least invasive way to make short restarts free: ``` PUT _all/_settings { "settings": { "index.unassigned.node_left.delayed_timeout": "5m" } } ``` Note what it does **not** do: it never delays the promotion of a replica to primary. Data stays available immediately; only the rebuilding of redundancy is deferred. ## The rolling restart recipe The conventional sequence, one node at a time: 1. `PUT _cluster/settings` with `"persistent": { "cluster.routing.allocation.enable": "primaries" }`. While this is set, the allocator will allocate primary shards only — so replicas of the node you are about to stop are not rebuilt anywhere. 2. Stop non-essential indexing if you can, and run `POST /_flush` so that recent operations are committed to disk. That shortens recovery when the node comes back. 3. Stop the node, do the maintenance, start it again, and wait for it to rejoin (`GET _cat/nodes`). 4. Set `cluster.routing.allocation.enable` back to `all`. 5. Wait for the cluster to return to green before touching the next node. Moving on while still yellow stacks unassigned shards from two nodes and can turn the cluster red. The values for that setting are `all`, `primaries`, `new_primaries` and `none`; `all` is the default and the state you must return to. Leaving it at `primaries` after the window is a classic incident: the cluster looks fine but never restores replicas, and stays quietly yellow until someone notices. Managed Elasticsearch platforms drive this through the node shutdown API, which signals whether a node is going away temporarily or permanently. Elastic documents that API as intended for use by their orchestration products rather than by hand, so the settings-based recipe above remains the one to describe in an interview. ## Restart versus decommission The two operations want opposite behaviour, and conflating them is the mistake to avoid: - **Restart** (node returns): suppress reallocation. Delayed allocation plus `enable: primaries`. - **Decommission** (node never returns): *cause* reallocation, in a controlled way, before you power it off. Set `cluster.routing.allocation.exclude._name: "es-data-7"` at the cluster level, watch `GET _cat/allocation?v` until that node holds zero shards, then shut it down. The cluster never goes yellow, because every copy was moved while the source was still alive. Afterwards, clear the exclusion — a forgotten exclusion is one of the most common causes of a mysterious `filter` NO in allocation explain later. ## Recovery cost and throttles Recovery is deliberately rate-limited so that rebuilding shards cannot starve live traffic. `indices.recovery.max_bytes_per_sec` caps the byte rate per node (40mb by default on ordinary nodes in current versions, with higher defaults on dedicated cold and frozen tiers). `cluster.routing.allocation.node_concurrent_recoveries` limits how many recoveries a node participates in at once (default 2), with separate incoming and outgoing variants, and `cluster.routing.allocation.node_initial_primaries_recoveries` (default 4) governs the burst when a whole cluster restarts. On fast local disks and a fat network these defaults are conservative and are often raised temporarily to shorten a maintenance window — then put back, because leaving them high means the next unplanned node loss saturates the cluster. ## Pitfalls worth naming Restarting several nodes at once without waiting for green. Forgetting to re-enable allocation. Leaving a decommission exclusion in place. Setting `delayed_timeout` so high that a genuinely dead node's redundancy is never restored. And using a restart as a fix for unassigned shards — if the allocator is refusing to place a shard, restarting nodes adds recoveries to a cluster that already cannot place copies.

  • What does index.unassigned.node_left.delayed_timeout not protect you from?
    It never delays promoting a replica to primary, so it does not extend any window of data unavailability — that is deliberate. It also only applies to shards unassigned because a node left; shards unassigned for disk, filtering or allocation failures are unaffected. And if the node does not return within the window, the full rebuild proceeds as normal.
  • Why drain a node with allocation filtering before decommissioning it, instead of just shutting it down?
    Shutting it down makes its shards unassigned and forces the cluster through a yellow period while copies are rebuilt from the surviving replicas, under recovery throttles. Excluding the node by name relocates its shards while it is still serving, so redundancy is never reduced. Remember to clear the exclusion afterwards, or later allocations will keep avoiding that node.
  • An operator finishes a rolling restart but the cluster stays yellow for hours. What would you check first?
    Whether cluster.routing.allocation.enable is still set to primaries. It is a persistent cluster setting and does not reset itself, so replicas are never allocated and the cluster sits yellow indefinitely. Check GET _cluster/settings, set it back to all, and confirm with allocation explain that the enable decider is no longer returning NO.

saying these in an interview costs you the question

  • Restarts several data nodes without waiting for green
  • Forgets to reset allocation enable back to all
  • Powers off a decommissioned node without draining it first
  • Thinks delayed allocation postpones primary promotion too
  • Leaves recovery throttles raised permanently after maintenance

context