How does a Lucene near-real-time reader make new documents searchable without calling IndexWriter.commit()?
answer
- Visibility and durability are separate concerns
- One reader sees the commit, another sees the writer
- Flushing is not fsyncing
- segments_N is written last, on purpose
- Reopen shares what has not changed
basics
~20 sOpening a reader with DirectoryReader.open(IndexWriter) flushes buffered documents into new segments and lets the reader see them without any fsync or new commit point. The documents become searchable in milliseconds but are not yet crash-durable.
solid answer
~50 sThere are two ways to get a reader. `DirectoryReader.open(Directory)` reads the latest **commit point** — the newest `segments_N` file — so it only ever sees data that has been committed. `DirectoryReader.open(IndexWriter)` opens a **near-real-time reader**: the writer flushes whatever it has buffered into new segment files, applies pending deletes, and hands the reader the current in-memory view of the segment list, including segments no commit point references yet. That makes new documents visible in milliseconds, but nothing has been fsynced and no `segments_N` was written, so a crash loses those segments. `IndexWriter.commit()` is the separate, expensive step: it fsyncs the segment files and writes a new `segments_N` in a two-phase sequence, which is what makes the data survive a restart. Production code reopens with `DirectoryReader.openIfChanged` behind a `SearcherManager` rather than opening from scratch, so unchanged segments are reused.
code
java · 5 lines// durable but stale: reads the newest segments_N
DirectoryReader committed = DirectoryReader.open(directory);
// fresh but not yet durable: sees uncommitted segments
DirectoryReader nrt = DirectoryReader.open(writer);go deeper
Remember that new documents are not searchable to an already-open reader: a reader must be opened or reopened, and doing so does not require a commit.
Explain the split cleanly — DirectoryReader.open(IndexWriter) for freshness, commit for durability — and describe what commit does with fsync and the segments_N pointer.
Show judgment about reopen cadence: aggressive refreshing creates tiny segments, merge pressure and cold caches, so freshness has to be budgeted against query latency and durability handled by a log beneath Lucene.
Own the contract you are offering the application: eventual visibility with a bounded lag, read-your-own-writes on request, and a durability story that names what is lost on a crash and who replays it.
## Two different questions: visible, and durable Lucene deliberately separates *when new documents can be searched* from *when they are safe against a crash*. Confusing the two is the single most common mistake in this area. **Visibility** is controlled by which set of segments a reader is looking at. **Durability** is controlled by commits. ## Readers from a Directory see only commits `DirectoryReader.open(Directory)` finds the newest `segments_N` file in the index directory and opens exactly the segments listed there. A commit point is an immutable snapshot: the file names it references are guaranteed to exist and to be fsynced. A reader opened this way therefore sees a consistent, durable state, and nothing that has happened since the last commit. ## Near-real-time readers `DirectoryReader.open(IndexWriter writer)` returns a **near-real-time (NRT) reader**. Internally the writer flushes its in-memory document buffer to new segment files, resolves any buffered deletes so the live-docs bitsets are current, and gives the reader the writer's own current view of the segment list — including segments that are not named by any `segments_N` on disk. The result is a searchable view of everything indexed so far, typically available in milliseconds to tens of milliseconds rather than the hundreds of milliseconds an fsync-bearing commit costs. The catch is precisely the missing fsync: if the process dies, those uncommitted segments are not part of any commit point and are discarded on restart. Any system that promises write durability therefore layers its own write-ahead log underneath — this is exactly the role Elasticsearch's translog plays while its NRT reopen provides visibility. ## Reopening rather than opening Opening a reader from scratch re-reads every segment's metadata and rebuilds per-segment state. Because segments are immutable, that is nearly all wasted work: only the newly added segments and changed live-docs generations are actually new. `DirectoryReader.openIfChanged(oldReader, writer)` returns a new reader that shares the underlying per-segment readers with the old one and only opens what changed, or returns null if nothing changed at all. In practice you do not call this by hand on the search path. `SearcherManager` holds the current `IndexSearcher`, hands it out under reference counting through `acquire()`/`release()`, and swaps in a fresh one when `maybeRefresh()` is called. `ControlledRealTimeReopenThread` drives those refreshes on a schedule with two targets — a relaxed interval for ordinary refreshes and a tighter one when a caller is waiting for a specific write to become visible — which is how you build read-your-own-writes on top of near-real-time semantics without reopening on every request. ## What commit actually does `IndexWriter.commit()` flushes any buffered documents, fsyncs every file belonging to the segments that will be in the new commit, and then writes a new `segments_N` file naming them, where N is an incrementing generation. Writing that pointer file last is what makes the commit atomic: either the new `segments_N` exists and the commit took effect, or it does not and the previous commit is still the truth. `prepareCommit()` plus `commit()` splits this into the two phases explicitly, which lets a caller coordinate a Lucene commit with an external transaction. Old commit points are cleaned up by an `IndexDeletionPolicy`; the default keeps only the most recent one. Substituting `SnapshotDeletionPolicy` pins a commit so its files are not deleted, which is how a consistent backup of a live index is taken. Note that `IndexWriter.flush()` is not `commit()`: flushing pushes buffered documents into segment files but writes no commit point and provides no durability guarantee. ## The cost of reopening too often Near-real-time reopening is cheap but not free, and its real cost is downstream. Every reopen forces the writer to cut whatever is buffered into a segment, so a very aggressive reopen cadence produces a stream of tiny segments. Tiny segments mean more per-segment work on every query, more merge pressure, and colder caches, because each new segment arrives with nothing warmed. Reopening a hundred times a second on a heavy indexing stream is a reliable way to make search slower while chasing freshness nobody asked for. The usual answer is a reopen interval measured in hundreds of milliseconds to seconds, with a mechanism to wait for a specific write when a particular caller genuinely needs to see it.
- What exactly makes an IndexWriter.commit() atomic?The commit writes and fsyncs all the segment files first, then writes the new segments_N file that names them. Until that pointer file exists, the index still resolves to the previous commit generation; once it exists, the whole new set is in effect. There is no intermediate state where a commit is half applied.
- Why is reopening a reader far cheaper than opening one from scratch?Segments are immutable, so a reopened reader can reuse the existing per-segment readers unchanged and only open the segments that appeared since, plus any updated live-docs generations. openIfChanged does exactly that, and returns null when nothing changed at all so the caller keeps the reader it has.
- How would you let a client read its own write without reopening on every request?Track a sequence number for the write and have the caller wait until a reader that includes it becomes current. ControlledRealTimeReopenThread supports precisely this: it refreshes on a relaxed interval normally, but honours a tighter deadline when a caller is blocked waiting for a specific generation to become searchable.
saying these in an interview costs you the question
- Says a commit is required before new documents can be searched
- Treats IndexWriter.flush() as providing durability
- Thinks a near-real-time reader survives a process crash
- Opens a brand-new reader per query instead of reopening
- Claims reopening more often is always better for users