skip to content

How do you change the number of replicas on a live Elasticsearch index?

level: juniorimportance: should knowfreq 60%

answer

  1. It is a dynamic setting, not a final one
  2. One endpoint, one JSON body
  3. Expect a colour change while copies build
  4. A copy never shares a node with its primary

basics

~20 s

Send a PUT to the index's _settings endpoint with index.number_of_replicas. It is a dynamic setting, so it applies immediately without closing the index; new copies are allocated and recovered while health shows yellow until they are in sync.

solid answer

~40 s

`index.number_of_replicas` is a dynamic setting, so you update it with `PUT /my-index/_settings` and a body of `{"index": {"number_of_replicas": 2}}` — no downtime, no close, and it can be applied to several indices at once with a wildcard. Raising it makes the master add unassigned replica shards to the routing table and allocate them; each new copy runs peer recovery from its primary, copying segments and replaying recent translog operations. Cluster health is yellow while those copies initialize and returns to green when all are in sync. Lowering it deletes copies immediately. Remember that a replica costs disk, indexing work and heap like any other shard, and that a replica can never be allocated to the same node as its primary — so `number_of_replicas: 1` on a single-node cluster leaves the cluster permanently yellow.

code

bash · 8 lines
bash
# load with no redundancy, then restore copies
PUT /orders-2026/_settings
{ "index": { "number_of_replicas": 0 } }

# ... bulk load ...

PUT /orders-2026/_settings
{ "index": { "number_of_replicas": 1 } }

go deeper

for a junior

Know the call by heart: PUT to the index's _settings with index.number_of_replicas, applied live. Also know that yellow health simply means replicas are still being built.

for a middle

Explain the mechanics: new copies are allocated, peer recovery copies segments and replays translog, and each replica repeats the primary's indexing work so writes cost more per copy.

for a senior

Show judgment about when adding a replica actually helps — search throughput and node-failure tolerance — versus when the real bottleneck is a hot shard or a saturated cluster, and mention loading with zero replicas.

for a principal

Frame replica count as a cost/availability policy across the fleet: how many copies each data tier gets, what that multiplies in storage spend, and where auto_expand_replicas is acceptable.

## The request Replica count is one of the few genuinely live knobs on an Elasticsearch index: ``` PUT /orders/_settings { "index": { "number_of_replicas": 2 } } ``` The index keeps serving reads and writes throughout. You can target many indices at once (`PUT /logs-*/_settings`), and you can set the value in an index template so every new index is created with it. The default for a newly created index is one replica. ## What the cluster actually does Raising the count is a cluster-state change: the master adds N new replica shards per primary to the index's routing table, initially unassigned. The allocator then places them on eligible nodes, subject to the rule that a copy of a shard never shares a node with another copy of the same shard, plus any allocation awareness, filtering, or disk-watermark constraints. Each newly allocated replica performs **peer recovery** from its primary: it copies the primary's Lucene segment files, then replays operations from the primary's translog (or uses the retention leases / soft deletes history to catch up) until it is in sync. Only then does it join the in-sync set and start serving searches. While that is happening, `GET _cluster/health` reports **yellow**: all primaries are assigned, but not all replicas. That is expected and harmless. It goes green when the last copy finishes. Lowering the count is the reverse and is effectively instantaneous — the master removes copies from the routing table and the data directories are deleted. ## What replicas buy and what they cost Replicas buy two things: - **Availability.** If a node holding a primary dies, an in-sync replica is promoted, so no data is lost and writes resume. With zero replicas, losing a node means losing shards — the cluster goes red. - **Search throughput.** A search request can be served by the primary or any in-sync replica, so more copies mean more nodes able to answer the same query concurrently. Adaptive replica selection picks the copy expected to answer fastest. They cost: - **Disk**, linearly. Two replicas means three copies of the data. - **Indexing work.** Every write is applied on the primary and then forwarded to each replica, which indexes it too. Replicas do not receive finished segments on the write path; they do the same analysis and indexing work. So write cost scales with copies. - **Shard overhead.** Every replica is a shard on some node, counting against per-node shard limits and consuming heap, file handles and merge work. ## Common operational patterns - **Bulk-load with zero replicas.** For a big initial load or a reindex into a fresh index, set `number_of_replicas: 0` for the load and raise it afterwards. The load only pays for one copy of the indexing work, and the replicas are then built by segment copy during recovery, which is usually cheaper than re-indexing everything N times. Only do this when losing the in-flight index would be acceptable — during the load there is no redundancy. - **Single-node development clusters.** `number_of_replicas: 1` on one node yields a permanently yellow cluster, because the replica cannot be allocated to the node already holding the primary. Set it to 0 for local development. - **`index.auto_expand_replicas`.** Small lookup indices can be set to something like `0-1` or `0-all`, letting Elasticsearch expand the replica count as nodes join and shrink it as they leave. It is handy for tiny indices you want on every node, and a trap for large ones. - **Search-heavy fleets.** If searches are queueing but CPU on some nodes is idle, adding a replica adds a servable copy. It does not help if the bottleneck is a single hot shard receiving all the routing keys, or if the whole cluster is CPU-saturated. ## Sizing implication When you plan capacity, count *shard copies*, not primaries. An index with 10 primaries and 1 replica occupies 20 shards' worth of disk, heap and per-node shard budget. This matters when checking guidance such as keeping shard size in the tens of gigabytes and keeping the number of shards per node well below the heap-derived limit — those budgets are consumed by every copy, not just the primaries.

  • Why does a single-node cluster with the default settings report yellow?
    The default is one replica per primary, and Elasticsearch never allocates a shard copy to the node that already holds another copy of the same shard. With one node the replicas stay unassigned forever, so primaries are all assigned (not red) but replicas are not (yellow). Setting `number_of_replicas: 0` makes such a cluster green.
  • Why do teams set replicas to 0 during a large bulk load?
    Every replica indexes each document itself, so N replicas multiply the indexing CPU and network cost of the load. Loading with zero replicas and raising the count afterwards means the copies are built by peer recovery — largely a segment file copy — instead of by re-running analysis. The tradeoff is no redundancy while the load runs.

saying these in an interview costs you the question

  • Thinks changing replicas requires closing or recreating the index
  • Says yellow health during replica recovery means data loss
  • Assumes replicas receive finished segments instead of indexing writes
  • Believes more replicas always increase indexing throughput
  • Runs number_of_replicas 1 on one node and calls the cluster broken

context