skip to content

What does Elasticsearch's _split API require of the source index, and how does _shrink differ?

level: middleimportance: should knowfreq 55%

answer

  1. Neither one resizes the index in place
  2. Writes must stop before either can run
  3. One direction needs a multiple, the other a factor
  4. Only one of them gathers shards onto a single node
  5. Segments are hard-linked, documents are not re-analysed

basics

~20 s

Both create a new index rather than resizing in place, and both need the source made read-only first. _split needs a target primary count that is a multiple of the source count and compatible with number_of_routing_shards; _shrink needs a target count that divides the source count, green health, and a copy of every shard gathered on one node.

solid answer

~50 s

Neither API changes an existing index — each builds a new one from the source's segment files, which is why the source must first be blocked for writes with `index.blocks.write: true`. **`_split`** increases the primary count: the target must be a multiple of the source count and must be compatible with the source's `index.number_of_routing_shards`, which was fixed at creation and caps how far you can ever split. It hard-links the source segments into the new shards where the filesystem allows it, then re-hashes and deletes the documents that do not belong in each new shard. **`_shrink`** decreases the primary count: the target must be a factor of the source count, cluster health must be green, and a copy of every shard must first be relocated onto one node using allocation filtering, since the new shard is assembled from those local segment files. Afterwards you re-point aliases at the new index, remove the write block, and delete the source.

code

bash · 15 lines
bash
# 1. gather a copy of every shard on one node and block writes
PUT /logs-000042/_settings
{ "settings": {
    "index.number_of_replicas": 0,
    "index.routing.allocation.require._name": "data-node-7",
    "index.blocks.write": true
} }

# 2. wait for green, then shrink to a single primary
POST /logs-000042/_shrink/logs-000042-shrunk
{ "settings": {
    "index.number_of_shards": 1,
    "index.number_of_replicas": 1,
    "index.routing.allocation.require._name": null
} }

go deeper

for a junior

Know that these two APIs exist and that they produce a new index with a different primary shard count rather than resizing the current one.

for a middle

State the constraints precisely: read-only source for both, target a multiple of the source for split, a factor for shrink, plus green health and all shards on one node for shrink.

for a senior

Describe the whole runbook — allocation filtering, write block, waiting for green, the atomic alias swap, count verification, deleting the source — and say when reindex is the safer choice instead.

for a principal

Decide when resizing is even the right lever: for append-only data, standardising template shard counts and rollover usually beats a fleet-wide resize campaign, and shrink belongs in the cold-data path.

## Both APIs make a new index `POST /source/_split/target` and `POST /source/_shrink/target` never modify the source in place. They create a new index with a different primary shard count and populate it from the source's existing Lucene segment files — using hard links when the source and target data paths are on the same filesystem, and copying otherwise. Because the source's data must not change while that happens, the first step in both flows is the same: make the source read-only. ``` PUT /logs-000042/_settings { "settings": { "index.blocks.write": true } } ``` That block allows reads and metadata changes but rejects indexing. Note the practical consequence: for an index still receiving writes, neither API is a live operation. That is one reason the usual answer for append-only data is to roll over to a correctly sized new index instead. ## _split — more primaries Constraints: - The target primary count must be a **multiple** of the source count (3 → 6, 3 → 9, 1 → anything). - It must be compatible with the source's `index.number_of_routing_shards`, a final setting chosen at index creation. This is the real ceiling: routing shards define the granularity that Elasticsearch's routing formula can be re-divided at, so an index can only be split into counts that the routing-shard configuration supports. If you know an index may need to grow, set `number_of_routing_shards` explicitly at creation. - The source must be read-only and healthy. Mechanically, the target's shards hard-link (or copy) the source segments, then Elasticsearch re-hashes every document and deletes the ones that belong to a different target shard, then recovers the target index. Nothing is re-analysed, so it is much cheaper than a reindex — but it does temporarily use extra disk, and the deleted documents remain in the segments until merges reclaim them, so a force merge afterwards is common. ## _shrink — fewer primaries Constraints: - The target primary count must be a **factor** of the source count (8 → 4, 8 → 2, 8 → 1). Shrinking to 1 is always legal. - Cluster health must be **green**. - A copy of **every** shard of the source must be on a single node, because the target shard is built from those local segment files. You arrange that with allocation filtering: ``` PUT /logs-000042/_settings { "settings": { "index.number_of_replicas": 0, "index.routing.allocation.require._name": "data-node-7", "index.blocks.write": true } } ``` Wait for relocation to finish (health green again), then call `POST /logs-000042/_shrink/logs-000042-shrunk` with the target settings, typically `index.number_of_shards: 1`, `index.number_of_replicas: 1` and a cleared allocation requirement. Because the operation is essentially a local hard-link of segments plus a recovery, it is fast relative to reindexing the same data. ## Afterwards Neither API touches aliases for you. The full flow is: create the target, wait for it to go green, re-point any read aliases (and the write alias, if applicable) from source to target, remove the write block on the target if you copied it, verify document counts, then delete the source. Doing the alias swap atomically with a single `_aliases` request avoids a window where searches see neither index or both. ## When to reach for which - **`_shrink`** is the workhorse for time-series data: yesterday's index is no longer written to, so shrinking it to one primary and force-merging reduces shard count and improves compression on cold data. This is precisely the shape that index lifecycle tooling automates. - **`_split`** is for the case where an index was created far too small and now holds shards well beyond the tens-of-gigabytes range while still needing to grow. It is less common, because it requires a write block and the routing-shard ceiling was fixed at creation. - **Reindex** remains the general fallback: it has no shard-count arithmetic constraints, can change mappings and analysis at the same time, and can run while the source keeps serving reads — but it re-analyses every document and is the most expensive option. ## Interview framing A strong answer names the arithmetic constraint for each (multiple versus factor), the shared read-only requirement, and the shrink-specific "all shards on one node, cluster green" prerequisite — and then notes that both produce a new index, so the alias swap is part of the plan, not an afterthought.

  • Why can an index only be split into certain shard counts?
    Because `index.number_of_routing_shards` is a final setting chosen at creation, and it defines the granularity the routing formula can be re-divided at. The target count must be a multiple of the source count and compatible with that routing-shard configuration. If you expect an index to grow, set `number_of_routing_shards` deliberately at creation so the split factors you may need are available.
  • Why is _shrink so much cheaper than reindexing the same data into a one-shard index?
    Shrink assembles the target shard from the source's existing Lucene segment files on the same node, hard-linking them where the filesystem permits, then runs a recovery. Nothing is re-analysed, re-tokenised or re-scored. A reindex reads every document and indexes it again through the full analysis chain, which is typically orders of magnitude more CPU.
  • What has to happen after either operation completes?
    Point the aliases at the new index — ideally in a single atomic `_aliases` request that removes the old target and adds the new one — verify document counts match, clear any write block or allocation filter you set on the target, and only then delete the source index. Force-merging a shrunken read-only index is a common final step.

saying these in an interview costs you the question

  • Thinks _split or _shrink changes the shard count of the existing index
  • Forgets the source must be made read-only first
  • Says _shrink works with shards scattered across nodes
  • Claims split can target any shard count you like
  • Assumes aliases follow the new index automatically

context