How does Lucene delete a document from an immutable segment, and when is the disk space actually reclaimed?
answer
- No in-place edit means no true delete
- A bit per document, kept beside the segment
- Searches filter, they do not remove
- Statistics still count the dead
- Space comes back at merge time
basics
~20 sLucene never edits a segment, so a delete only marks the document in that segment's live-documents bitset and searches filter the hit out. The terms, stored fields and disk space go away only when a merge rewrites the segment without the dead documents.
solid answer
~50 s`IndexWriter.deleteDocuments` resolves the term or query to internal document numbers and clears their bits in a per-segment **live docs** bitset, written as a small new generation file (`_1_1.liv`) alongside the untouched segment data. Searches consult that bitset and skip dead documents, so they vanish from results immediately, but their postings, stored fields and doc values are still on disk and still counted in the segment's term statistics — which is why document frequencies, and therefore scores, can drift slightly on a heavily-deleted index. `maxDoc` keeps counting them while `numDocs` does not. The space comes back only when a merge rewrites the surviving documents into a fresh segment and drops the old files. Lucene also supports **soft deletes**: instead of clearing a live-docs bit, the document is marked via a doc-values field and retained by a retention-aware merge policy, so it can still be read for operation history — this is what Elasticsearch uses to support replica recovery.
code
java · 3 lineswriter.deleteDocuments(new Term("id", "42"));
// document is filtered from results, but still on disk:
// reader.maxDoc() unchanged, reader.numDocs() decreased by 1go deeper
Know that a Lucene delete does not remove data: the document is marked and filtered out of results, and the space is only reclaimed later during a merge.
Explain the live-documents bitset written as a new generation file, why updates are delete-plus-add, and the maxDoc versus numDocs distinction that exposes tombstones.
Be ready to diagnose an index far larger than its live data — an update-heavy workload, starved merges, or retained soft deletes — and to choose between letting merging catch up and a targeted delete-expunging merge.
Own the design question: an update-per-document-per-minute workload on an immutable-segment engine is a compaction problem, and the answer is usually to change the data model or the index lifecycle, not to tune deletes.
## Why a delete cannot be a delete A Lucene segment is written once and never modified, so there is no way to cut a document out of a postings list or a stored-fields block. The delete has to be recorded somewhere else and applied as a filter at read time. ## The live-documents bitset Each segment can carry a **live docs** bitset — one bit per internal document number in that segment, set when the document is alive. When you call `IndexWriter.deleteDocuments(Term)` or `deleteDocuments(Query)`, Lucene buffers the deletion, resolves it against each affected segment to a set of internal document numbers, and clears those bits. The result is written as a new small file with an incremented deletes generation, for example `_1_1.liv`, then `_1_2.liv`, and so on. The segment's real data files are not rewritten. Every query applies that bitset. `IndexSearcher` and the collectors it drives skip document numbers whose live bit is clear, so a deleted document disappears from results as soon as a reader is opened that sees the new deletes generation. ## What is still there The important consequence is that the document is only invisible, not gone: - Its postings entries remain in the term dictionary and postings lists, so scanning a term's postings still walks over them and pays the cost of skipping. - Its stored fields, doc values and term vectors are still on disk, occupying space. - It still contributes to the segment's collection and term statistics: document frequency for its terms, total document count, average field length. Scoring uses those statistics, so on an index where a large fraction of documents are deleted, BM25 scores are computed against a document population that includes the dead. - `IndexReader.maxDoc()` counts it; `IndexReader.numDocs()` does not. The gap between them is exactly the number of tombstoned documents, and it is the number to watch operationally. ## Updates are deletes `IndexWriter.updateDocument(term, doc)` is not an update. It deletes everything matching the term and adds a fresh document, which lands in whichever segment is currently being written. A workload that rewrites the same documents repeatedly therefore generates tombstones at exactly the rate it generates updates, and its index grows far beyond the live data size until merging catches up. ## When space actually returns Only a **merge** reclaims the space. A merge reads a set of segments, writes the still-live documents of all of them into one new segment with fresh internal numbering, and then deletes the input files once no reader is using them. Deleted documents are simply not copied. Lucene's `TieredMergePolicy` weighs deleted-document percentage when selecting merges and will schedule merges specifically to bring the index-wide deleted fraction back under its allowed threshold, so on a normally-configured index the space does come back on its own — eventually and asynchronously, not at delete time. If you need it back sooner on an index that has stopped changing, `IndexWriter.forceMergeDeletes()` rewrites segments whose deleted percentage is above a threshold. That is a targeted, cheaper alternative to a full `forceMerge(1)`, but it is still a rewrite of real data and should not be a routine background job. ## Soft deletes Lucene also supports **soft deletes**, configured with `IndexWriterConfig.setSoftDeletesField(...)` and applied with `IndexWriter.softUpdateDocument(...)`. Instead of clearing a live-docs bit, the delete is recorded by writing a marker doc-values field on the document. A retention-aware merge policy (`SoftDeletesRetentionMergePolicy`) then decides when such a document may finally be dropped by a merge, and readers can be wrapped so that soft-deleted documents are hidden from ordinary searches. The point of soft deletes is that the deleted documents remain physically readable for a while. That gives the layer above Lucene an ordered history of operations it can replay — which is how Elasticsearch supports bringing a lagging replica back into sync without copying whole segment files. The tradeoff is that retained soft-deleted documents keep occupying space and keep being carried through merges until the retention policy releases them. ## What this means in practice Deleting a large fraction of an index does not shrink it, and does not immediately speed anything up. Index size on disk must be read as live data plus tombstones plus merge headroom. If the deleted-document ratio is chronically high, the cause is usually an update-heavy workload or a merge policy that is not keeping up because merges are starved of I/O — not something you fix by deleting more carefully.
- How do you tell how much of an index is tombstoned rather than live?Compare the reader's maxDoc with numDocs: maxDoc counts every document slot ever written in the currently-live segments, numDocs excludes the deleted ones, and the difference is the tombstone count. Per-segment tooling exposes the same figures alongside each segment's size, so you can see which segments are carrying the dead weight.
- Why can scores shift slightly after a big merge even though no documents changed?Deleted documents keep contributing to term statistics — document frequency and average field length — until the segment holding them is merged. A merge drops them, so the collection statistics that BM25 uses change, and computed scores move a little. Statistics are also per-segment, so relevance can differ across segments before merging evens things out.
- Is deleting by query more expensive than deleting by term?Yes. A delete by term can be resolved with a single term-dictionary lookup per segment, while a delete by query must execute that query against every segment to enumerate matching document numbers. Both then only clear bits, but the resolution step for a broad query can be as expensive as running the search itself.
saying these in an interview costs you the question
- Says deleting documents frees disk space right away
- Thinks the postings entries for a deleted document are removed
- Ignores that deleted documents still affect term statistics
- Confuses numDocs and maxDoc, or never mentions the gap
- Treats an update as an in-place modification of the document