skip to content

What does setting index.translog.durability to async cost you when an Elasticsearch node crashes?

level: seniorimportance: should knowfreq 44%

answer

  1. the acknowledgement stops meaning what it meant
  2. fsync moves from per request to a timer
  3. there is a sync_interval sized window
  4. process crash and power loss differ here
  5. replicas share the same exposure

basics

~20 s

With async durability the translog is fsynced on a timer rather than per request, so a node crash or power loss can lose every acknowledged write since the last sync interval. The gain is far fewer fsyncs and higher indexing throughput.

solid answer

~50 s

The default, `index.translog.durability: request`, fsyncs the translog on the primary and every in-sync replica before the write is acknowledged, so an acknowledged write survives a crash. Setting it to `async` decouples the two: operations are appended to the translog but fsynced only every `index.translog.sync_interval` (default `5s`). Throughput improves because you stop paying an fsync per bulk request, which on modest disks is often the bottleneck. The cost is a real data-loss window — if the node loses power, everything written since the last sync can be gone, even though clients received success responses. Note the failure mode is machine-level: an ordinary process crash still leaves the data in page cache for the OS to write, but a kernel panic or power loss does not. It is a reasonable trade for reindexable data such as logs or a search index rebuilt from a system of record; it is not acceptable when Elasticsearch is the only copy of the data.

go deeper

for a junior

Know that the translog is what makes an acknowledged Elasticsearch write survive a crash, and that its durability mode is a per-index setting.

for a middle

Explain the two modes and the sync interval, and be able to say exactly what an acknowledgement means under each.

for a senior

Argue the trade for a concrete workload: name the loss window, the failure class it exposes, why replicas do not neutralise it, and the safer throughput levers you would try first.

for a principal

Own it as a data-classification policy — which index classes may run async, whether a system of record exists upstream to replay from, and how the acknowledgement semantics are documented to the teams writing to the cluster.

## What the translog is for Each Elasticsearch shard keeps a **translog** — an append-only log of every indexing and delete operation. Its purpose is to close the gap between an acknowledged write and a Lucene commit. Segments are only fsynced at flush time, which happens relatively rarely; without a log, everything written since the last commit would vanish on a crash. On restart a shard opens its last Lucene commit and replays the translog forward to reach the acknowledged state. ## The durability setting `index.translog.durability` takes two values: - **`request`** (default): before responding to a write, the shard appends the operation to the translog and **fsyncs** it — and it does this on the primary *and* on every in-sync replica. An acknowledged write is therefore on stable storage in multiple places. - **`async`**: operations are appended to the translog and the response goes out immediately. The translog is fsynced on a timer, every `index.translog.sync_interval` (default `5s`). ## What you gain An fsync is a synchronous durability barrier on the storage device. Under a heavy bulk load, per-request fsyncs on every shard copy can dominate write latency, particularly on network-attached or spinning storage. Switching to `async` removes that barrier from the request path, batching many operations into one periodic sync. On fsync-bound hardware the throughput difference is large; on fast local NVMe it is often much smaller, which is worth measuring before accepting the risk. ## What you lose You lose the meaning of the acknowledgement. With `async`, a `200` response means "we have this in memory and in the page cache", not "this is on disk". If the machine loses power or the kernel panics, up to `sync_interval` worth of acknowledged writes is gone from that copy. Two refinements matter in an interview: 1. **Process crash versus machine crash.** If only the Elasticsearch process dies, unsynced translog data is still in the operating system's page cache and the OS will write it out; the shard recovers fine. The exposure is specifically to power loss, kernel panic, or a violent host termination. 2. **Replicas do not save you automatically.** With `async`, the replicas are also running async. A correlated failure — a rack losing power, a hypervisor host dying with several copies on it, an availability-zone outage — can take out every copy's unsynced window at once. Replication reduces the probability but does not turn `async` back into `request`. There is also a subtle behavioural point: if a shard is *removed and recovered*, and the operation was never fsynced anywhere, it simply never existed as far as the cluster is concerned. There is no reconciliation pass that notices a client was told "success". ## When it is a reasonable trade - **Reindexable data.** If Elasticsearch is a derived index over a relational database, an event log, or object storage, losing five seconds of writes costs a replay, not data. This is the canonical case. - **Observability and metrics pipelines.** Losing a few seconds of logs during a host power loss is usually tolerable, and these workloads are exactly the fsync-bound ones. - **Bulk backfills**, where you set it temporarily for the duration of the load and restore `request` afterwards. This is the safest use: the window of exposure is bounded and you can always re-run the load. ## When it is not - Elasticsearch is the system of record for the data. - The write is the acknowledgement a user or another system acts on — an order, a payment, an audit event. - Compliance requires that an acknowledged write is durable. In those cases keep `request` and buy throughput elsewhere: bigger bulk requests, more shards or nodes, faster storage, higher `refresh_interval`, fewer replicas during a load. ## Related knobs `index.translog.sync_interval` sets how often the async sync runs and therefore the size of the loss window. `index.translog.flush_threshold_size` governs when a flush commits segments and lets the translog be trimmed — that is about recovery time and disk usage, not about the acknowledgement guarantee. Both are dynamic index settings, so `async` can be turned on for a load and off again without recreating the index. ## How to present it A strong answer names the guarantee being sold — "acknowledged means fsynced on primary and in-sync replicas" — then says precisely what `async` replaces it with, the size of the window, the class of failure it exposes you to, and the workloads where that is an honest trade. Weak answers describe it as a general "performance setting" without naming the data loss.

  • If the Elasticsearch process is killed with SIGKILL while running async translog durability, is unsynced data lost?
    Usually not. The operations were already written to the translog file, so they sit in the OS page cache and the kernel flushes them even though Elasticsearch is gone. The exposure specific to `async` is host-level: power loss, kernel panic, or a hard hypervisor kill. This distinction is what separates a precise answer from a vague one.
  • Does having two replicas make async durability safe?
    No. Every copy is running async, so all of them have the same unsynced window. Independent host failures are unlikely to coincide, but correlated ones — a rack power event, a hypervisor failure, an availability-zone outage — can take out all copies' windows together. Replication lowers probability; it does not restore the guarantee.
  • What would you change instead of async if indexing throughput is too low but data loss is unacceptable?
    Raise `index.refresh_interval` so fewer segments are cut, use larger bulk requests with several concurrent bulk threads, drop replicas for the duration of an initial load and add them back, use auto-generated document ids to skip the existence lookup, and check for write thread-pool rejections. All of those buy throughput without weakening the acknowledgement.

saying these in an interview costs you the question

  • Describes async as purely a performance setting with no downside
  • Says replicas make async durability safe
  • Believes an ordinary process crash loses the unsynced translog
  • Confuses the sync interval with the refresh interval
  • Thinks a flush per request is what makes writes durable

context