skip to content

Mapping Changes, Reindex & Aliases

You cannot change the type of an existing field — the index has to be rebuilt. The production answer is an alias in front of versioned indices plus _reindex, and interviewers ask for exactly that zero-downtime cutover story.

part ofElasticsearchoverview, primer and where to startread it →
on this pageshow

questions

6

Why do Elasticsearch applications query an index alias instead of the concrete index name?

level: juniorimportance: must knowfreq 72%

answer

  1. A second name for an index
  2. Nothing is copied; it is a pointer
  3. Lets you rebuild and swap underneath clients
  4. One _aliases call, remove plus add, atomic
  5. Writes need exactly one is_write_index target

basics

~20 s

An alias is a second name that points at one or more indices. Applications talk to the alias so the physical index behind it can be rebuilt, versioned and swapped atomically, with no client change and no downtime.

solid answer

~50 s

An index alias is an indirection layer: `products` is a name that resolves to a concrete index such as `products-v3`. Because Elasticsearch indices cannot be renamed and most mapping changes force a rebuild, clients that hard-code the index name have to be redeployed every time the index is rebuilt. With an alias, you build `products-v4` beside the old one, reindex into it, and then move the alias in a **single** `POST /_aliases` call whose `actions` array removes the alias from v3 and adds it to v4. That call is atomic — there is no instant where `products` resolves to nothing or to both. Aliases can also carry a `filter` (a view over a subset of documents) and routing values, can span several indices for reads, and can designate exactly one backing index as the write target via `is_write_index`.

code

json · 7 lines
json
POST /_aliases
{
  "actions": [
    { "remove": { "index": "products-v1", "alias": "products" } },
    { "add":    { "index": "products-v2", "alias": "products" } }
  ]
}

go deeper

for a junior

Be ready to say what an alias is in one sentence and create one, and to explain that clients query the alias so the index behind it can be replaced without touching client code.

for a middle

Explain the atomic _aliases actions array versus a delete-then-add sequence, and why a multi-index alias needs is_write_index before it will accept an indexing request.

for a senior

Show that you would establish the alias convention before the first document is indexed, keep the previous index for rollback, and remember that reindex does not carry aliases across.

for a principal

Own the naming and indirection standard across teams — versioned index names, read and write aliases, dashboards and ingest pipelines all pointed at aliases — so no rebuild ever requires a coordinated client redeploy.

## What an alias actually is An Elasticsearch index alias is a **pointer stored in cluster state**, not a copy of data. Creating an alias moves no bytes and costs no disk. Any API that accepts an index name — search, index, delete, bulk, aggregations — accepts an alias name and resolves it to the concrete index or indices behind it at request time. The motivating fact is that **an Elasticsearch index cannot be renamed, and its mapping is largely immutable**. Changing a field's type, changing an index-time analyzer, or changing the shard count all require creating a *different* index and copying documents into it. If every application, dashboard and script names `products` directly, every one of them must be edited and redeployed at the moment the new index is ready. That turns a routine mapping change into a cross-team release. ## The indirection convention The standard production convention is: **no client ever names a concrete index.** You create `products-v1` (or `products-000001`) and give it the alias `products`. Clients only know `products`. When the mapping must change you create `products-v2`, populate it, verify it, and repoint the alias. It costs nothing to add the alias on day one and it is nearly impossible to add later without a coordinated redeploy, which is why interviewers treat "I would have put an alias in front of it from the start" as the expected answer. ## The atomic swap Repointing is done through the `_aliases` API, which takes an `actions` array applied as one cluster-state update: ``` POST /_aliases { "actions": [ { "remove": { "index": "products-v1", "alias": "products" } }, { "add": { "index": "products-v2", "alias": "products" } } ] } ``` All actions in one call succeed or fail together, so searches never see a window where `products` is missing or ambiguous. Doing the same thing as two separate calls — delete the alias, then create it — leaves a gap in which every query fails with `index_not_found_exception`. That distinction is the whole point of the API and is a common interview probe. The same call can carry a `remove_index` action, which deletes the old index in the same atomic update; in practice you normally keep the old index for a rollback window instead. ## Read aliases versus write aliases An alias may point at many indices. For **searches** that is a feature: one alias fans a query across `logs-2026-06`, `logs-2026-07` and so on. For **writes** it is ambiguous — Elasticsearch cannot guess which index a new document belongs in, so an index request through a multi-index alias is rejected unless exactly one of the backing indices is marked `"is_write_index": true`. That flag is what lets a rollover-style alias accept writes into the current index while still serving reads from all of them. A common pattern separates the two concerns explicitly: a read alias `products` over one or more indices, and a write alias `products-write` pointing at exactly one. During a cutover you can then flip reads and writes at different moments. ## Filtered and routed aliases An alias can attach a `filter`, so every search through it is implicitly ANDed with that query — a cheap tenant or category view over a shared index. It can also pin `routing` (or `index_routing`/`search_routing`) so requests through it hit a single shard. A filtered alias is a **convenience, not a security boundary**: a caller who is allowed to query the underlying index directly simply bypasses the filter. Real isolation needs the security features' document-level security, or separate indices. ## Practical gotchas - Aliases are cluster state; thousands of them are not free, though normal usage is nowhere near a problem. - `_reindex` does **not** copy aliases to the destination index — you add them yourself as part of the swap. - Deleting an index silently deletes its aliases with it. - An alias name may not collide with an existing index name. - Kibana index patterns, ingest destinations and Logstash outputs should all be pointed at the alias too, or the indirection leaks. ## Why interviewers ask it The alias question is really a proxy for "have you ever changed a mapping in production?" The candidate who has says *alias in front of versioned indices, atomic swap, keep the old one for rollback* without being prompted.

  • What happens if you index a document into an alias that points at three indices?
    The request is rejected as ambiguous unless exactly one of those indices is marked `"is_write_index": true` in the alias definition; then the document goes there. Reads are unaffected — a multi-index alias searches all of them. This is how rollover-style aliases serve reads across a whole set while writes land only in the current index.
  • Can you delete the old index in the same call that repoints the alias?
    Yes — the `_aliases` actions array supports a `remove_index` action alongside `add`/`remove`, so the swap and the deletion apply as one cluster-state update. In practice most teams do not: keeping the previous index around for a rollback window is cheap, and deleting it later is a one-line follow-up once the new index has proven itself.
  • Is a filtered alias a safe way to isolate one tenant's data?
    No. The filter applies only to requests that go through the alias; anyone permitted to query the underlying index directly sees everything. Filtered aliases are a convenience view — for real isolation use document-level security from the security features, or give each tenant its own index.

An alias is a DNS CNAME for your index: clients keep calling the same friendly name while you rebuild the machine behind it and cut over in one atomic change.

saying these in an interview costs you the question

  • Thinks an alias duplicates or copies the index data
  • Deletes the alias then recreates it, leaving a failing window
  • Assumes any alias accepts writes regardless of how many indices it spans
  • Treats a filtered alias as a security boundary between tenants
  • Plans to rename an index instead of swapping an alias

context

open as a page

Why can't you change an existing field's type in an Elasticsearch mapping?

level: middleimportance: must knowfreq 78%

basics

~20 s

A field's type and analyzer are baked into immutable Lucene segments when each document is indexed, so already-written data cannot be reinterpreted under a new type. Adding new fields is fine; changing an existing one requires a new index and a reindex.

open as a page

When does _update_by_query suffice for a mapping change and when do you need _reindex?

level: middleimportance: should knowfreq 55%

basics

~20 s

Use _update_by_query when the mapping change was legal in place — a new field or sub-field — and existing documents just need rewriting to pick it up. Use _reindex into a new index whenever the mapping itself could not be updated, such as a field type change.

open as a page

How do you run _reindex over a 500-million-document index without destabilising the cluster?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Pre-create the destination, then run the reindex asynchronously with wait_for_completion=false, parallelise it with slices, throttle it with requests_per_second, allow conflicts to proceed, and monitor the task id — rethrottling or cancelling it if live traffic suffers.

open as a page

How do you change an Elasticsearch index mapping with zero downtime while writes keep arriving?

level: seniorimportance: should knowfreq 65%

basics

~20 s

Build a new index with the corrected mapping beside the old one, reindex into it, keep it current with dual writes or repeated timestamp catch-up passes, verify counts and queries, then move the alias in one atomic call and keep the old index for rollback.

open as a page

How would you make Elasticsearch mapping changes a routine operation rather than a multi-hour incident?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Treat every index as a disposable, versioned artifact behind an alias, keep the authoritative data outside Elasticsearch so any index can be rebuilt from source, automate create-reindex-verify-swap as one pipeline, and budget standing capacity for rebuilds.

open as a page