How do you run _reindex over a 500-million-document index without destabilising the cluster?
answer
- Never hold it open on one HTTP connection
- Ask for a task id instead
- One knob parallelises, one knob slows it down
- You can change the rate mid-flight
- Mappings and settings do not come along
basics
~20 sPre-create the destination, then run the reindex asynchronously with wait_for_completion=false, parallelise it with slices, throttle it with requests_per_second, allow conflicts to proceed, and monitor the task id — rethrottling or cancelling it if live traffic suffers.
solid answer
~50 sNever run it synchronously. `POST /_reindex?wait_for_completion=false` returns a task id immediately; you poll `GET _tasks/<id>` for progress and the final result is recorded in the `.tasks` index. Set `"slices": "auto"` so the scroll is split into parallel sub-tasks, roughly one per source shard, instead of one serial stream. Set `requests_per_second` to cap the write rate, and adjust it live with `POST _reindex/<task_id>/_rethrottle?requests_per_second=…` — `-1` removes the limit. Add `conflicts: "proceed"` if the destination is also taking writes. Tune the destination for bulk load first: create it explicitly with the right mapping (reindex copies neither mappings nor settings), consider zero replicas and a relaxed refresh interval while it is not being queried, then restore both before the swap. Watch write-queue rejections, merge and disk pressure, and cancel via the tasks API if latency for real users moves.
code
json · 6 linesPOST /_reindex?wait_for_completion=false&slices=auto&requests_per_second=1000
{
"conflicts": "proceed",
"source": { "index": "events-v1", "size": 2000 },
"dest": { "index": "events-v2" }
}go deeper
Know that _reindex copies documents from one index to another and that the destination should be created first with the mapping you want. Operating a huge run is not expected at this level.
Explain the asynchronous task flow — wait_for_completion=false, a task id, GET _tasks — plus what slices and requests_per_second control and why the destination needs an explicit mapping.
Show you have run one: rethrottle instead of restarting, watch write-queue rejections and disk headroom, know cancellation leaves partial data, and bound each run so a restart is cheap.
Treat large rebuilds as planned capacity events — headroom, scheduling windows, an idempotent restartable job, and a standing decision about how much user-facing latency a background rebuild may consume.
## Run it as a task, not a request A half-billion-document reindex runs for hours. Issuing it synchronously ties the outcome to one HTTP connection, and a proxy timeout leaves you unsure whether the job died or is still writing. `POST /_reindex?wait_for_completion=false` returns `{"task": "<nodeId>:<taskId>"}` at once. `GET _tasks/<task_id>` reports the running status — documents created, updated, batches, version conflicts, throttling. When the task ends, its final result is stored as a document in the `.tasks` system index and remains readable there, which is what you want for an unattended overnight run. Cancellation is `POST _tasks/<task_id>/_cancel`. Cancelling does **not** roll back what was already written — a partially-populated destination is why you validate before swapping an alias, and why an idempotent restart matters. ## Parallelise with slices By default a reindex is a single scroll: one stream, one bulk pipeline, and on a large index it will take far longer than the cluster's capacity implies. `"slices": "auto"` splits the source scroll into independent sub-tasks — typically one per source shard — that run in parallel and are reported as child tasks under the parent. You can also pass an explicit number. Slicing multiplies throughput and multiplies load in the same breath. It is the main lever for finishing an overnight job by morning and the main way to saturate a cluster that is also serving users. ## Throttle deliberately `requests_per_second` limits the rate of batches, with the engine sleeping between them to hold the target. The value you can afford is discovered, not calculated: start conservative and raise it while watching search latency. The important operational detail is that this is adjustable **on a running task**: `POST _reindex/<task_id>/_rethrottle?requests_per_second=500`, or `-1` to remove the limit. You do not cancel and restart a six-hour job because you guessed the rate wrong. Speeding up applies promptly; a task already sleeping on a very low rate finishes its current wait before the new rate takes effect. ## Prepare the destination `_reindex` copies documents. It does **not** copy mappings, index settings, or aliases. Create the destination first, with the mapping you actually want — otherwise dynamic mapping guesses types from the first documents and you have rebuilt the same mistake. For the load itself, a destination nobody is querying does not need replicas or frequent refreshes: creating it with `number_of_replicas: 0` and a relaxed `refresh_interval` removes per-write replication and constant segment creation, then both are restored before the index takes traffic. Restoring replicas triggers a copy across nodes, so allow time for the index to go green before the swap. ## Tune the batch `source.size` sets how many documents each scroll batch pulls (the bulk write size follows it). Large documents with a big batch produce huge bulk requests and heap pressure; small batches waste round trips. Adjust it when documents are unusually large, and remember that a single oversized batch can trip the circuit breaker. ## Reindexing from another cluster `source.remote` pulls documents from a different cluster over HTTP: ``` "source": { "remote": { "host": "https://old-cluster:9200", "username": "…", "password": "…" }, "index": "products" } ``` The **destination** cluster's nodes must whitelist the source host via `reindex.remote.whitelist` in `elasticsearch.yml`, which is a node setting and therefore a restart. Throughput is bounded by the HTTP hop and the remote's scroll, and documents are buffered on heap during transfer — for large documents, reduce the batch size rather than discovering the limit under load. ## Handle failures and conflicts A bulk failure — mapping rejection, malformed value — aborts the reindex by default and is reported in the response `failures` array. Version conflicts on the destination are separate: `conflicts: "proceed"` counts them and carries on, which is what you want when the destination is also receiving live dual writes. Because a failed reindex leaves partial data, make restarts cheap: reindex by a bounded query (a date range or id range per pass) so a failed segment of work can be redone without re-copying everything. ## What to watch while it runs Search latency and the write thread pool queue and rejections; merge activity and disk free space (the destination grows, and merges need working room); heap and circuit-breaker trips; and the task's own progress counters. If user-facing latency moves, rethrottle down first — cancel only if that is not enough. ## The summary an interviewer wants "Pre-create the destination, run it as a throttled sliced task, monitor it, rethrottle rather than restart, verify before swapping, and know that cancelling leaves partial data behind."
- You realise mid-run that the reindex is hurting search latency. What do you do?Rethrottle it down rather than killing it: `POST _reindex/<task_id>/_rethrottle?requests_per_second=…` takes effect on the running task, so hours of completed work are preserved. Cancel only if throttling is insufficient — and remember cancellation leaves whatever was already written in the destination, so the restart plan has to tolerate partial data.
- What does cancelling a reindex task leave behind?Everything it had already written. There is no transaction and no rollback, so the destination holds a partial copy. That is why you verify before swapping an alias, and why bounding each run by a query — a date or id range — makes a restart cheap: you redo one bounded slice of work instead of half a billion documents.
- What must be configured before reindexing from a remote cluster?The destination cluster's nodes need the source host listed in `reindex.remote.whitelist` in `elasticsearch.yml`, which requires a restart since it is a node setting. Credentials go in the `remote` block, connectivity must exist from data nodes to the remote HTTP port, and batch size usually needs lowering for large documents because transferred batches are buffered in heap.
saying these in an interview costs you the question
- Runs a multi-hour reindex synchronously and treats the timeout as failure
- Thinks cancelling a reindex rolls back what it wrote
- Assumes the destination inherits the source mapping and settings
- Cancels and restarts instead of rethrottling the running task
- Leaves slices at the default and concludes reindex is inherently slow