skip to content

Why are Lucene index segments immutable, and what does that design buy the search engine?

level: middleimportance: must knowfreq 68%

answer

  1. Written once, never edited
  2. Think about locking and caching
  3. Compression assumes no in-place update
  4. Deletes and space are the price
  5. Merging exists to pay that price

basics

~20 s

Lucene writes each segment once and never modifies it. Immutability lets readers share segments lock-free, enables aggressive write-once compression and per-segment caching, and makes indexing append-only. The cost is that deletes become tombstones and space returns only when segments merge.

solid answer

~50 s

A Lucene index is a set of **segments**, each a small self-contained inverted index with its own term dictionary, postings, doc values and stored fields. `IndexWriter` buffers incoming documents in RAM and flushes them into a brand-new segment; once written, a segment's files are never edited again. That buys a lot: a reader can hold a segment open with no locking or copy-on-write, structures can be built once in their most compact form (FST term dictionaries, bit-packed postings, block-compressed stored fields), per-segment caches and warmed structures stay valid forever, and replication or backup can copy files that will never change. The cost is the other half of the story: an update is a delete plus an add, deletes are recorded as tombstones outside the segment data, and neither the old terms nor the disk space go away until a background merge rewrites the survivors. Merging is therefore a permanent, unavoidable cost of the design.

code

bash · 5 lines
bash
# a Lucene index directory: two segments plus the commit pointer
_0.cfs  _0.cfe  _0.si
_1.cfs  _1.cfe  _1.si  _1_1.liv   # _1 has a deletes generation
segments_3
write.lock

go deeper

for a junior

Be able to say that a Lucene index is made of segments, that segments are written once and never modified, and that an update is really a delete followed by an add.

for a middle

Explain the mechanics: buffer, flush to a new segment, per-segment search, background merges. Name at least two concrete benefits of immutability, such as lock-free readers and write-once compression.

for a senior

Show that you have felt the consequences in production — disk headroom for merges, deleted documents lingering, segment count driving query latency, and I/O spikes when merges land during peak traffic.

for a principal

Frame immutability as the log-structured tradeoff it is: cheap appends and lock-free reads paid for with write amplification and background compaction, and reason about which workloads that bargain suits.

## What a segment actually is A Lucene index is a directory of files, not a single file, and those files are grouped into **segments**. A segment is a complete, standalone inverted index over some subset of the documents: it has its own term dictionary, its own postings lists, its own doc values, stored fields and norms, and its own internal document numbering starting at 0. All of a segment's files share a name prefix (`_0.si`, `_0.fdt`, `_0.dvd`, or a single `_0.cfs` compound file for small segments). A search executes against every segment in the index independently, and the per-segment results are merged into one result list at the end. ## The write path `IndexWriter` accumulates added documents in an in-memory buffer. When that buffer fills (or something asks for a flush), the buffered documents are encoded and written out as one **new** segment. Nothing in the existing segments is touched. The only files that change after a segment is created are tiny side files carrying a new generation number — the live-documents bitset for deletes and, in later Lucene versions, updated numeric or binary doc values — plus the `segments_N` commit file that lists which segments belong to the index. ## Why immutability is worth it **Lock-free concurrency.** A reader captures the set of segment files it opened and reads them for as long as it lives. Because a writer can only add new files, never rewrite one in place, readers and the writer need no coordination at all beyond reference-counting files so they are not deleted while in use. That is what makes a search engine able to serve heavy read traffic during heavy indexing. **Write-once encoding.** When you know a data structure will never be modified, you can compress it as hard as you like. Lucene stores the term dictionary as a finite state transducer, delta-encodes and bit-packs document IDs in postings, and block-compresses stored fields. None of these formats can absorb an insert in the middle; all of them are dramatically smaller and faster to scan than an update-friendly equivalent. **Cache validity.** Filter caches, warmed data structures and OS page cache entries are keyed to a segment. Because a segment's bytes never change, nothing ever needs invalidating. Reopening a reader after new documents arrive means only the newly created segments require any work; everything already cached is reused verbatim. **Cheap durability and cheap replication.** A commit is essentially fsyncing files that were written once and then writing a pointer file naming them. Backup and replication are incremental for free: a file that has already been copied will never differ, so only new segment files need shipping. ## What immutability costs **Updates are delete-plus-add.** There is no in-place edit of a document. `IndexWriter.updateDocument(term, doc)` is defined as deleting everything matching the term and then adding the new document, which lands in a different (newer) segment. **Deletes do not free space.** A delete flips a bit in the segment's live-documents set; the postings entries and the stored fields for that document stay on disk and are simply filtered from results. The space returns only when a merge rewrites that segment without the dead documents. **Write amplification.** Because the only way to reclaim space or reduce segment count is to rewrite data, every document is typically written several times over its life as it is merged up through successive segment sizes. That background I/O competes with indexing and search on the same disks. **Segment count is a search cost.** Every query pays a per-segment fixed cost — a term dictionary lookup per segment, a per-segment collector, a merge of the per-segment top-N. A thousand tiny segments is much slower to search than a handful of large ones, which is exactly why merging exists. **Stale statistics.** Deleted documents still contribute to a segment's term statistics until the segment is merged, so document frequencies (and therefore scores) reflect documents that no longer match anything. ## The lifecycle in one line Documents are buffered in RAM, flushed into a new immutable segment, made visible to readers, merged in the background into progressively larger segments, and finally recorded durably by a commit that fsyncs the files and writes a new `segments_N` pointing at them. ## Why interviewers care Almost every operational surprise in Lucene-based systems is a consequence of immutability: why deleting rows does not shrink the index, why an index needs spare disk headroom, why there is a delay before a new document is searchable, why merges spike I/O at inconvenient moments, and why refresh, flush and merge are three separate concepts in the engines built on top of Lucene rather than one.

  • If segments are immutable, how does a searcher ever see a newly indexed document?
    A new document is buffered in RAM and then flushed into a brand-new segment. A reader opened afterwards sees a larger list of segments, including the new one. Existing readers keep their old snapshot and never change; freshness comes entirely from reopening a reader over the newer segment list, not from any segment mutating underneath a reader.
  • Does immutability mean nothing in an existing segment ever changes on disk?
    Almost. The main data files are never rewritten, but a segment can gain small generation files: a new live-documents bitset when documents are deleted, and updated doc-values files when numeric or binary doc values are changed. These are written as new files with an incremented generation, so the original bytes still stay untouched.
  • Why does a Lucene index need free disk headroom beyond its current size?
    Merging writes the merged output before the input segments can be deleted, so a merge temporarily holds both copies. Deleted documents also occupy space until merged away. A commonly used rule of thumb is to keep enough free space for the largest merge you allow, plus room for tombstoned documents that have not been reclaimed yet.

It is like a photo album where you never erase a page: to correct a picture you paste in a new page and cross out the old one, and the album only gets thinner when someone reprints it without the crossed-out pages.

saying these in an interview costs you the question

  • Claims Lucene updates a document in place inside its segment
  • Thinks deleting documents immediately shrinks the index on disk
  • Believes more segments always means faster searches
  • Says merging is optional tuning rather than structural to the design
  • Confuses a segment with a shard or an index

context