skip to content

When does _update_by_query suffice for a mapping change and when do you need _reindex?

level: middleimportance: should knowfreq 55%

answer

  1. One writes back, one writes elsewhere
  2. Ask whether the mapping update was accepted
  3. Backfilling a new sub-field needs no new index
  4. Type changes cannot target the same index
  5. Both share conflicts, slices and throttling

basics

~20 s

Use _update_by_query when the mapping change was legal in place — a new field or sub-field — and existing documents just need rewriting to pick it up. Use _reindex into a new index whenever the mapping itself could not be updated, such as a field type change.

solid answer

~40 s

The dividing line is whether the **mapping update was accepted**. `_update_by_query` reindexes every document into the *same* index, re-running it through the current mapping and any ingest pipeline or script you supply. That is exactly what you need after adding a multi-field or a new dynamically-detected field, or to apply a script fix across documents. It cannot help when the change was rejected — a type change, an index-time analyzer change, `object` to `nested`, or a different shard count — because the index it writes back into still has the old mapping. Those need `_reindex` into a freshly created index with the corrected mapping, followed by an alias swap. Both APIs share the same machinery: `conflicts=proceed`, `slices`, `requests_per_second` throttling, and asynchronous execution via `wait_for_completion=false` plus the tasks API.

code

json · 4 lines
json
POST /products/_update_by_query?conflicts=proceed&wait_for_completion=false
{
  "query": { "exists": { "field": "title" } }
}

go deeper

for a junior

Recall the basic distinction: _update_by_query rewrites documents inside the same index, _reindex copies them into a different one. Knowing which API targets which destination is enough here.

for a middle

Explain the decision rule — was the mapping update accepted? — and name concrete cases on each side, plus the shared controls: conflicts, slices, throttling and asynchronous tasks.

for a senior

Demonstrate operational judgment: run these asynchronously and throttled, expect segment churn and temporary disk growth, and read version conflicts as evidence of live writes rather than failure.

for a principal

Frame it as capacity policy — full-index rewrites are a recurring cost, so cluster sizing, disk headroom and off-peak scheduling windows must assume them rather than treat each one as an incident.

## Two operations that look alike `POST /<index>/_update_by_query` and `POST /_reindex` are built on the same engine: scroll or search the source, run each document through an optional script and ingest pipeline, and bulk-write the result. The difference is the destination. `_update_by_query` writes back into the **same index**. `_reindex` writes into a **different index**, which may live in the same cluster or on a remote one. That single difference decides which one solves your problem. ## When `_update_by_query` is the right tool Use it when the mapping change was *legal* and only the stored documents lag behind: - **A multi-field was added.** `title.keyword` exists in the mapping but holds nothing for documents indexed before the change. Rewriting each document populates it. - **A new field arrived via a template or dynamic mapping** and older documents need it derived from data they already carry, via `script`. - **A data fix across the index** — normalising a status string, stripping a prefix, recomputing a denormalised value — expressed as a Painless script. - **An ingest pipeline should be applied to existing documents**; `_update_by_query` accepts a `pipeline` parameter. Mechanically it reads each document's `_source`, re-indexes it, and bumps `_version`. Because Lucene has no in-place update, every rewritten document becomes a new document plus a tombstone on the old one. An index can therefore roughly double on disk until merges reclaim the deleted docs. It is full-index I/O, not a metadata touch. ## When only `_reindex` will do Use it when the mapping could not be updated at all, so the target must be a new index: - **A field's type must change** — `text` to `keyword`, `keyword` to `date`, `long` to `double`. - **An index-time analyzer must change**, including adding a custom analyzer to a field that is already populated. - **`object` must become `nested`**, which changes both the type and the physical document layout. - **The primary shard count must change** (barring `_split`/`_shrink`, which have their own preconditions). - **Documents must move between clusters**, using `source.remote`. - **Only a subset should survive** — `_reindex` takes a `source.query`, so a rebuild can also drop stale documents by simply not copying them. ## Shared controls worth naming Both APIs accept the same operational knobs, and interviewers like to hear them: - `conflicts: proceed` — continue past version conflicts instead of aborting, counting them in the response. Essential on an index still taking live writes. - `slices: auto` or a number — run the scroll in parallel sub-tasks, typically one per source shard. - `requests_per_second` — throttle the write rate; adjustable on a running task with the `_rethrottle` endpoint. - `wait_for_completion=false` — return a task id immediately; poll `GET _tasks/<id>` and find the final result recorded in the `.tasks` index. Never run a multi-hour job synchronously over an HTTP connection. - `max_docs` — cap the work, which makes a dry run on a slice of the data easy. ## The version-conflict subtlety `_update_by_query` takes a snapshot of the index at the start and uses internal versioning on write. If the application updates a document between the snapshot and the rewrite, the rewrite sees a version conflict — and that is the correct outcome, because the live update is newer. With `conflicts=proceed` that document is skipped and counted; you then decide whether a second pass is needed. A candidate who says "conflicts mean it failed, rerun it from scratch" has missed that the conflict is protecting fresher data. ## Cost and scheduling Neither is free. Both rewrite every matched document, generate segment churn, and compete with live indexing for the write thread pool. On a large index, run them asynchronously, throttled, and preferably off-peak; watch for write-queue rejections and disk headroom. For `_update_by_query` remember the temporary size inflation from deleted docs awaiting merge. ## The decision in one line If `PUT _mapping` accepted your change, `_update_by_query` backfills it. If `PUT _mapping` rejected it, you are building a new index with `_reindex` and swapping an alias.

  • What does conflicts=proceed actually do on an _update_by_query request?
    It tells the task to keep going when a document's version changed between the initial snapshot and the rewrite, instead of aborting. Skipped documents are counted in the `version_conflicts` field of the response. The conflict means live traffic wrote a newer version, so skipping is usually correct; you inspect the counter afterwards and decide whether another pass is warranted.
  • Why can an index temporarily grow after _update_by_query even though no documents were added?
    Lucene segments are immutable, so rewriting a document writes a new copy and marks the old one deleted rather than editing it. Until merges run, both copies occupy disk. On a full-index rewrite that can approach double the original size, so check free space before starting and expect merge I/O afterwards.

saying these in an interview costs you the question

  • Thinks _update_by_query can apply a field type change
  • Believes it edits documents in place without rewriting them
  • Runs a multi-hour job synchronously and calls the timeout a failure
  • Treats version conflicts as corruption rather than newer live writes
  • Assumes _reindex copies the source index settings and mappings

context