In Elasticsearch, what is the difference between a refresh and a flush?
answer
- one is about visibility, the other about durability
- which one fsyncs, which one doesn't
- one produces a new segment each time
- the other lets the translog be trimmed
- Lucene commit versus reopened searcher
basics
~20 sA refresh opens the in-memory buffer as a new Lucene segment so recent documents become searchable, without fsyncing anything. A flush performs a Lucene commit that fsyncs segments to disk and trims the translog, which is about durability and recovery, not visibility.
solid answer
~50 sThey sit on two different axes. **Refresh** is about *visibility*: it turns the shard's in-memory indexing buffer into a new Lucene segment and reopens the searcher, so newly indexed documents can be matched. It writes into the filesystem cache and does not fsync, so it is comparatively cheap — the default `index.refresh_interval` is `1s`. **Flush** is about *durability and recovery cost*: it performs a Lucene commit, fsyncing the segments so they survive a crash, and then rolls over and trims the translog because the operations it held are now safely in committed segments. Elasticsearch flushes automatically based on translog size and age; the `_flush` API exists but is rarely needed manually. Crucially, per-request durability does not depend on flush at all — that comes from the translog fsync governed by `index.translog.durability`. Flush mainly bounds how much translog has to be replayed after a restart.
go deeper
Hold on to the one-line split: refresh equals searchable, flush equals committed to disk plus a trimmed translog. Do not swap the two.
Explain the mechanics both ways: which one writes a segment, which one fsyncs, what triggers each, and why refreshing too often creates merge pressure.
Show you can tune this per workload — refresh interval for ingestion-heavy indices, and an argument for why durability lives in the translog rather than in flush frequency.
Frame it as a cost model across the fleet: visibility latency, segment and merge overhead, and restart recovery time are three budgets, and refresh, translog durability and flush thresholds are the knobs that spend them.
## Two independent axes Every Elasticsearch shard has to answer two separate questions about a freshly indexed document: - *Can a query see it?* — controlled by **refresh**. - *Does it survive a crash?* — controlled by the **translog** and, eventually, **flush**. Candidates who merge these two get the follow-ups wrong, because the operations have different costs, different triggers and different consequences. ## What a refresh does Indexed documents accumulate in an in-memory buffer belonging to the shard's Lucene index writer. A refresh: - closes that buffer and writes it out as a new **segment** — an immutable mini-index with its own inverted index, doc values and stored fields; - reopens the shard's searcher over the new set of segments. The segment is written to the filesystem cache. There is **no fsync**. That is what makes refresh cheap enough to run every second by default (`index.refresh_interval: 1s`), and it is also why refresh gives you no durability guarantee: an unflushed segment in page cache dies with the machine. The cost of refreshing is not zero. Each refresh creates a segment, and segments must eventually be merged. Refreshing very frequently on a write-heavy index produces many tiny segments, which increases merge work, per-search segment overhead and file handles. This is exactly why forcing `refresh=true` on every write is an anti-pattern, and why log-style indices are often tuned to `30s` or higher. An index with no explicit `refresh_interval` that has not been searched for a while goes **search idle** and stops refreshing on a timer until the next search arrives, which then triggers and waits for a refresh. ## What a flush does A flush performs a **Lucene commit**: the segment files are fsynced, a new commit point is written, and the index on disk is now self-consistently recoverable without any external log. Having committed, Elasticsearch can start a new translog generation and discard the old one, because everything it described is now durably in segments. Flush is triggered automatically. The main triggers are translog size (`index.translog.flush_threshold_size`) and translog age. You can call `POST /index/_flush`, and it is occasionally useful before a planned node restart or shutdown to shorten recovery, but routine manual flushing is not something a healthy system needs. ## Where durability actually comes from This is the part interviewers push on. Per-request durability is **not** provided by flush. It is provided by the **translog**: with the default `index.translog.durability: request`, the translog is fsynced on the primary and every in-sync replica before the write is acknowledged. So a crash one millisecond after the acknowledgement loses nothing — on restart the shard recovers its last Lucene commit and replays the translog forward. What flush buys you is a *bound on recovery time*. Without flushes, the translog would grow without limit and a restart would have to replay every operation ever written. Flushing periodically caps that replay. ## Putting the sequence together For a single document: 1. Write reaches the primary; document enters the in-memory buffer; operation appended to translog and fsynced (default durability); replicated to in-sync copies; acknowledged. 2. Within about a second, a refresh converts the buffer into a searchable segment — visible now, still not fsynced. 3. Some time later, a flush commits the segments to disk and the translog entries covering them are trimmed. 4. In the background, merges combine small segments into larger ones and purge deleted documents. ## Common confusions worth naming - "Flush makes documents searchable" — no, refresh does. A flush does not change what queries can see. - "Refresh makes data durable" — no, the translog does, and the flush is what lets the translog be trimmed. - "`_flush` is the fix for a slow node" — flushing manually rarely helps and can add I/O; look at merges, refresh rate and thread-pool rejections instead. - Confusing flush with **force merge** (`POST /index/_forcemerge`), which rewrites segments down to a target count. Force merge is an expensive rewrite intended for indices that are no longer being written to, not a durability operation.
- If a flush is what fsyncs segments, what protects a write acknowledged one millisecond before a crash?The translog. With the default `index.translog.durability: request`, the operation is appended and fsynced to the translog on the primary and every in-sync replica before the response is sent. On restart the shard opens its last Lucene commit and replays the translog forward, so the acknowledged write reappears.
- Why is raising index.refresh_interval a standard tuning step for log ingestion?Each refresh creates a segment. At one second, a heavy write load produces a stream of tiny segments that must then be merged, consuming CPU and I/O and inflating segment count. Raising the interval to 30s or higher amortises the buffer into fewer, larger segments, at the cost of a longer visibility delay that log search rarely cares about.
- How does a force merge differ from a flush?A force merge rewrites existing segments into fewer, larger ones, physically removing documents marked as deleted. It is an expensive full rewrite of the shard's data, appropriate only for indices that will receive no further writes. A flush merely commits what is already there and trims the translog.
saying these in an interview costs you the question
- Says a flush is what makes documents searchable
- Claims a refresh fsyncs segments to disk
- Thinks durability requires waiting for the next flush
- Confuses flush with force merge
- Recommends calling the flush API on a schedule to keep data safe