skip to content

What does forceMerge(1) on a Lucene index actually cost, and when is it justified?

level: seniorimportance: should knowfreq 52%

answer

  1. Cost scales with total size, not staleness
  2. Old files stay until the new one is complete
  3. Everything afterwards must be re-warmed
  4. Output exceeds the normal size ceiling
  5. Right for finished indexes, wrong for live ones

basics

~20 s

forceMerge(1) rewrites every live document into a single segment: a full read and re-encode of the whole index, saturated CPU and I/O, up to roughly double the disk space while it runs, and a cold page cache afterwards. It is justified only on an index that has stopped receiving writes.

solid answer

~50 s

`IndexWriter.forceMerge(1)` ignores the normal merge budget and rewrites the entire index into one segment. That means reading and re-encoding every live document, so the work is proportional to index size, not to how much has changed. While it runs, the input segments cannot be deleted, so the index temporarily needs roughly its own size again in free disk, and the merge saturates I/O and CPU in competition with live traffic. Afterwards, the page cache holds nothing useful because every file is new. Worse, the output is far above the policy's maximum merged segment size, so ordinary tiered merging will largely leave it alone; modern Lucene can rewrite such a segment when its deleted fraction exceeds the allowed threshold, but that is another full-size merge. So: run it on an index that is finished being written — a completed time-based index, a search index built offline — and never on a schedule against a live one. To reclaim tombstones instead, use `forceMergeDeletes()`.

code

java · 5 lines
java
// wrong: routine maintenance on a live index
writer.forceMerge(1);

// better: reclaim tombstones only, in segments above the threshold
writer.forceMergeDeletes();

go deeper

for a junior

Know that force-merging a Lucene index compacts it into fewer segments and that it is an expensive, deliberate operation rather than something to run casually.

for a middle

Explain the concrete costs — a full rewrite of every live document, temporary double disk usage, cold caches — and that the output segment sits above the merge policy's normal size ceiling.

for a senior

Draw the operational line: force-merge only indexes that have stopped taking writes, reach for forceMergeDeletes when the real goal is reclaiming tombstones, and treat a chronic segment problem as a merge-throughput problem.

for a principal

Own the lifecycle policy: decide where in an index's life compaction happens, budget the I/O and disk headroom for it, and resist compaction-as-ritual in favour of measured segment and deletion metrics.

## What the call does `IndexWriter.forceMerge(int maxNumSegments)` merges the index down to at most that many segments, bypassing the ordinary selection logic and the `maxMergedSegmentBytes` cap that `TieredMergePolicy` normally enforces. `forceMerge(1)` is the extreme case: one segment for the whole index. It is a blocking operation by default — the call returns when the merges are done — and the merges themselves run on the merge scheduler's threads. ## Why it is genuinely attractive On a finished index the result is the best possible shape: - One segment means one term-dictionary lookup per term per query instead of one per segment, and one collector instead of a per-segment collect-and-merge. On indexes with many segments this is a real, measurable latency win. - Every deleted document is gone, so no query wastes time skipping tombstones and term statistics reflect only live documents. - The single segment is as compact as Lucene can make it, which improves cache residency and shrinks backups. This is why the operation exists and why it is the right move for read-only indexes: an index built offline and shipped to search nodes, or a time-based index that has rolled over and will never be written again. ## What it costs **Full-index I/O and CPU.** The cost scales with total index size, not with how much data is stale. A 500 GB index means reading 500 GB and writing something close to it, decoding and re-encoding every postings list, doc-values column and stored-fields block on the way. On a busy node this competes directly with query serving and indexing. **Disk headroom.** Merge output is written before the input segments can be dropped, and inputs cannot be dropped while any open reader still references them. Peak usage during a full force-merge is therefore roughly twice the index size, sometimes more if long-lived readers pin old commits. Running out of disk mid-merge is a genuinely bad failure mode. **Cache destruction.** Every file afterwards is new, so the operating system page cache and every per-segment cache the layer above maintains start empty. Query latency is elevated until things warm again — often the most user-visible part of the cost. **A segment that does not merge normally.** The output far exceeds `maxMergedSegmentBytes`, so routine tiered merging will not fold it into anything. As deletes accumulate in that giant segment, the index-wide deleted percentage climbs. Since Lucene 8 `TieredMergePolicy` can rewrite a single oversized segment on its own when the deleted fraction exceeds `deletesPctAllowed`, so it is no longer true that such a segment is permanently stranded — but the cleanup is another whole-segment rewrite, scheduled by the engine at a time you did not choose. On an index that keeps taking writes, force-merging therefore trades a controlled, incremental merge cost for occasional enormous ones. **Long non-interruptible work.** A multi-hour merge cannot be politely cancelled midway without discarding the work; aborting an in-flight merge throws away everything it has written. ## The narrower alternative If the actual goal is reclaiming space from deleted documents rather than reducing segment count, `IndexWriter.forceMergeDeletes()` is the targeted tool: it rewrites only those segments whose deleted percentage exceeds `forceMergeDeletesPctAllowed`, leaving well-packed segments alone. That is substantially cheaper than a full force-merge and does not create an oversized segment. Even so, it rewrites real data and should be a deliberate maintenance action, not a cron job. Most of the time the correct answer is to do nothing and let `TieredMergePolicy` work. It already schedules merges when the deleted fraction gets too high and when segment counts drift above budget. If it appears not to be keeping up, the cause is usually starved merge I/O, an unusually high update rate, or long-lived readers pinning old commits — none of which a manual force-merge fixes for longer than a few hours. ## The interview answer State the rule crisply: force-merging is a build-time optimisation for indexes that have stopped changing, not a maintenance routine for live ones. Then show you know why — full-index rewrite cost, double disk headroom, cold caches, and an oversized segment whose eventual cleanup is a merge just as large as the one you ran deliberately.

  • Why does the index need roughly twice its size in free disk during a full force-merge?
    The merged segment is written in full before any input segment can be deleted, and inputs stay on disk while open readers still reference the commit that names them. So both the old and new copies of the data coexist at the peak. Long-lived readers or a pinned commit can extend that overlap well past the end of the merge.
  • When is force-merging clearly the right call?
    When an index has stopped receiving writes and will be searched for a long time afterwards: a rolled-over time-based index, a catalogue rebuilt offline and shipped read-only, or a snapshot taken for archival. In those cases the one-off cost is paid once and every subsequent query benefits, and the oversized-segment drawback never bites because no deletes arrive.
  • A team schedules forceMerge(1) nightly to keep search fast. What do you tell them?
    That it rewrites the whole index every night regardless of how much changed, empties every cache during the following morning's traffic, needs double disk headroom each run, and produces a segment the merge policy will later have to rewrite anyway. Measure segment count and deleted percentage instead, and fix merge throughput if those are genuinely out of line.

saying these in an interview costs you the question

  • Recommends forceMerge as routine maintenance on a live index
  • Assumes the cost is proportional to deleted documents, not index size
  • Forgets the temporary doubling of disk usage
  • Ignores cache invalidation and the latency spike afterwards
  • Cannot distinguish forceMerge from forceMergeDeletes

context