skip to content

Write Model, Refresh & Durability

A write goes to the primary, then to the in-sync replicas, and becomes durable via the translog — but it is not searchable until a refresh. Interviewers use that near-real-time gap and refresh_interval tuning to check you know why a document you just indexed is missing.

part ofElasticsearchoverview, primer and where to startread it →
on this pageshow

questions

6

Why is a document Elasticsearch just acknowledged as indexed sometimes missing from an immediate search?

level: juniorimportance: must knowfreq 78%

answer

  1. acknowledged is not the same as searchable
  2. search is near-real-time, not real-time
  3. a new segment must be opened first
  4. refresh, roughly one second by default
  5. get by id reads the translog

basics

~20 s

Elasticsearch search is near-real-time, not real-time. Indexing puts the document in an in-memory buffer and the translog; it only becomes searchable when a refresh turns that buffer into a new searchable Lucene segment, by default about once a second.

solid answer

~40 s

When an index request returns success, the document is safely recorded — it is in the shard's in-memory buffer and has been written to the translog on the primary and its in-sync replicas. But it is not yet *searchable*. Search runs against Lucene segments, and a new document only enters a segment when the shard performs a **refresh**, which opens a new searchable segment. `index.refresh_interval` defaults to `1s`, so there is normally a sub-second window where the document exists but no query will match it. Note the asymmetry: a `GET /index/_doc/<id>` by document id *is* real-time, because it can read the operation out of the translog, while a search is not. If a specific request must be visible before you return, use `refresh=wait_for` on that request rather than shortening the interval globally.

go deeper

for a junior

Recall the phrase near-real-time and the reason: search reads Lucene segments, and a refresh has to build one first, about once a second by default.

for a middle

Explain the buffer, translog and segment mechanics, and separate visibility (refresh) from durability (translog and flush). Know that a get by id is real-time while a search is not.

for a senior

Be ready to diagnose a read-your-writes complaint in production: choose refresh=wait_for on the one path that needs it, and explain why forcing a refresh per write wrecks segment count and merge load.

for a principal

Own the policy: which indices get a tuned refresh_interval, whether the product genuinely needs read-your-writes, and what visibility latency you are willing to publish as a contract to consuming teams.

## The claim being made Elasticsearch describes itself as a **near-real-time** search engine. The precise meaning is: there is a bounded, usually sub-second delay between a write being accepted and that write being visible to queries. Accepting the write and making it visible are two different events, driven by two different mechanisms, and confusing them is the single most common source of "my document disappeared" bug reports. ## What happens when you index a document A write request is routed to the primary shard that owns the document. The primary: 1. Writes the operation into an **in-memory buffer** (Lucene's indexing buffer for the shard). 2. Appends the operation to the shard's **translog** (transaction log), a sequential append-only file on disk. 3. Forwards the operation to the in-sync replica copies, which do the same. 4. Returns success to the client once the replicas have acknowledged. At that point the document is durable in the sense that a crash will not lose it — the translog can be replayed on restart. What it is *not* is queryable. ## Why the buffer is not queryable Lucene, the library underneath each shard, searches **segments**: immutable, self-contained mini-indexes containing an inverted index, doc values, stored fields and so on. Segments are built, not edited. Documents sitting in the in-memory buffer have not yet been turned into a segment structure, so there is nothing for a query to match against. ## Refresh A **refresh** is the operation that closes the current in-memory buffer, writes it out as a new segment (into the filesystem cache — no `fsync` is involved), and reopens the shard's searcher so that queries can see the new segment. After a refresh, the documents that were in the buffer are searchable. Refresh is scheduled per index by `index.refresh_interval`, whose default is `1s`. So the normal visibility delay is up to about one second. There is an important extra rule: an index whose `refresh_interval` has not been set explicitly and which has received no search requests recently becomes **search idle** and stops refreshing on a timer altogether. The next search against that shard triggers a refresh and waits for it. This saves a lot of pointless work on write-heavy, rarely-queried indices such as log data, but it means a manual test that indexes and immediately searches may see a longer or a shorter delay than one second depending on traffic. ## Refresh is not durability Refresh makes data *visible*. It does not make it *durable* — the new segment lives in the filesystem cache and has not been fsynced. Durability comes from the translog (fsynced per request by default) and from a **flush**, which performs a Lucene commit, fsyncs segments and trims the translog. Two independent axes: refresh controls visibility, translog and flush control durability. ## Getting a document without waiting A real-time `GET` by id sidesteps the whole issue. Elasticsearch's get API is real-time: if the requested document has not been refreshed into a segment yet, the engine can serve it out of the translog. This is why a workflow of "index then fetch by id" appears to work while "index then search" appears broken. ## Forcing visibility deliberately Write requests take a `refresh` parameter: - `refresh=false` (default) — do nothing extra. - `refresh=wait_for` — hold the response until the next scheduled refresh has made the change visible. - `refresh=true` — force an immediate refresh of the affected shards, then respond. There is also a standalone `POST /index/_refresh`. In tests and in read-your-writes flows, `wait_for` is usually the right choice; `refresh=true` on a high-throughput write path creates a flood of tiny segments and heavy merge pressure. ## What a good answer sounds like "Acknowledged means recorded and replicated, not searchable. Search hits Lucene segments; the document is in a buffer until a refresh builds a segment, which happens about every second by default. I'd use `refresh=wait_for` on the one request that needs read-your-writes rather than lowering the global interval."

  • Does a GET by document id have the same delay as a search in Elasticsearch?
    No. The get API is real-time: if the document has not been refreshed into a segment yet, the engine serves it from the translog. Only search, which runs against Lucene segments, has to wait for a refresh. That is why "index then get by id" works while "index then search" appears to lose the document.
  • How would you make an integration test reliably see the document it just indexed?
    Send the write with `refresh=wait_for`, or call `POST /index/_refresh` between the write and the assertion. Both are explicit and local to the test. Do not fix it by lowering `index.refresh_interval` globally or by sleeping — the first hurts production indexing throughput and the second is flaky.
  • What is search idle, and how does it change refresh behaviour?
    If an index has no explicit `refresh_interval` and has received no search requests for a while, its shards stop refreshing on a timer. The next search triggers a refresh and waits for it. It saves work on write-heavy, rarely-searched indices such as logs, but it makes the visibility delay depend on query traffic.

saying these in an interview costs you the question

  • Says an index response of 201 means the document is searchable
  • Thinks a flush, not a refresh, is what makes documents searchable
  • Claims refresh fsyncs data to disk for durability
  • Fixes it by setting refresh=true on every production write
  • Concludes the write was lost because a search missed it

context

open as a page

In Elasticsearch, what is the difference between a refresh and a flush?

level: middleimportance: must knowfreq 68%

basics

~20 s

A refresh opens the in-memory buffer as a new Lucene segment so recent documents become searchable, without fsyncing anything. A flush performs a Lucene commit that fsyncs segments to disk and trims the translog, which is about durability and recovery, not visibility.

open as a page

How do _seq_no and _primary_term give Elasticsearch optimistic concurrency control on document updates?

level: middleimportance: must knowfreq 58%

basics

~20 s

Every write stamps the document with the primary's sequence number and the primary term. A client re-sends both as if_seq_no and if_primary_term; the shard applies the write only if they still match the current document, otherwise it returns a 409 version conflict.

open as a page

How does refresh=wait_for on an Elasticsearch index request differ from refresh=true?

level: middleimportance: should knowfreq 52%

basics

~20 s

refresh=wait_for holds the response until the next scheduled refresh makes the change visible, forcing no extra work. refresh=true forces an immediate refresh of the affected shards, creating an extra segment per request and much more merge pressure.

open as a page

What happens on the primary and its in-sync replicas when Elasticsearch indexes one document?

level: seniorimportance: should knowfreq 62%

basics

~20 s

The coordinating node routes the request to the primary, which validates it, applies it locally with a sequence number and primary term, then forwards it in parallel to every in-sync replica. Once they respond the primary acknowledges; a failed replica is reported to the master and removed from the in-sync set.

open as a page

What does setting index.translog.durability to async cost you when an Elasticsearch node crashes?

level: seniorimportance: should knowfreq 44%

basics

~20 s

With async durability the translog is fsynced on a timer rather than per request, so a node crash or power loss can lose every acknowledged write since the last sync interval. The gain is far fewer fsyncs and higher indexing throughput.

open as a page