skip to content

How would you set a Lucene index's merge and commit strategy for continuous heavy ingest with low search latency?

level: principalimportance: should knowfreq 30%

answer

  1. Three dials, not one
  2. Freshness, durability and compaction are separate costs
  3. Every small segment is future merge work
  4. Ingest stalls can be a merge symptom
  5. Stop making one index do two jobs

basics

~20 s

Treat merge throughput as provisioned capacity, not a tuning afterthought. Size the RAM buffer so flushes produce reasonably large segments, reopen readers on a bounded interval rather than per write, commit on a cadence matched to an external durability log, and give merges enough I/O headroom to avoid indexing stalls.

solid answer

~50 s

Start by naming the three independent dials people conflate: **flush** (how big new segments are), **reopen** (how fresh search results are), and **commit** (how much you lose in a crash). Then set each against a stated requirement. A larger `IndexWriter` RAM buffer produces fewer, larger segments and less merge work, so raise it as far as heap allows. Reopen on a fixed interval — hundreds of milliseconds to seconds — behind a `SearcherManager`, and offer a wait-for-generation path for the rare caller needing read-your-own-writes, rather than reopening per request. Commit rarely, and let a write-ahead log above Lucene provide per-operation durability. For merging, keep `TieredMergePolicy` defaults and instead provision I/O: `ConcurrentMergeScheduler` stalls indexing when pending merges pile up, so a merge backlog surfaces as ingest latency. Finally, separate lifecycles — write into a hot index, compact it once when it stops taking writes.

code

java · 6 lines
java
IndexWriterConfig cfg = new IndexWriterConfig(analyzer);
cfg.setRAMBufferSizeMB(512.0);          // fewer, larger segments at flush
ConcurrentMergeScheduler cms = new ConcurrentMergeScheduler();
cms.setMaxMergesAndThreads(12, 4);      // headroom before indexing stalls
cfg.setMergeScheduler(cms);
IndexWriter writer = new IndexWriter(dir, cfg);

go deeper

for a junior

Know the three distinct events — flushing buffered documents, reopening a reader for freshness, and committing for durability — and that they are configured separately.

for a middle

Explain the mechanics linking them: a reopen forces a flush, small segments create merge work, and a commit is the fsync that makes data survive a crash.

for a senior

Diagnose the steady state — merge backlog presenting as ingest stalls, segment count driving query latency — and choose a reopen and commit cadence from stated freshness and durability requirements.

for a principal

Own the capacity and lifecycle strategy: merge throughput and disk headroom are provisioned resources, freshness is a negotiated requirement with a measurable price, and separating write-hot from read-optimised indexes is the structural lever.

## Frame the problem as three dials, not one The most common failure in this conversation is treating "make it fast" as a single knob. In a Lucene-based system there are three separable decisions, each with its own requirement and its own cost: - **Flush size** — how much is buffered in RAM before becoming a segment. Governs segment count at birth and therefore total merge work. - **Reopen interval** — how often a new reader is opened. Governs search freshness. - **Commit interval** — how often files are fsynced and a new `segments_N` is written. Governs crash durability. Stating them separately, and asking what the business actually requires for each, is most of the answer. ## Flush: make segments as large as you can afford Every segment born small must be merged upward, and each byte gets rewritten once per tier it climbs. The cheapest merge is the one never scheduled, so give `IndexWriter` as large a RAM buffer as heap safely permits. Be aware that anything forcing a flush — an NRT reopen, a commit, hitting a document-count limit — cuts the buffer short, so an aggressive reopen cadence silently defeats a generous buffer setting. If ingest is bursty, batching documents into larger `addDocuments` calls also helps by amortising per-call overhead. ## Reopen: bound the lag, do not chase zero Freshness has a price paid by everyone. Each reopen forces a flush, creating a segment; a hundred reopens a second creates a hundred segments a second, all of which must be searched until merged and all of which arrive with cold caches. Choose an interval that satisfies the actual product requirement — usually somewhere between a few hundred milliseconds and a few seconds — and implement it with `SearcherManager` plus a reopen thread, so all queries share one searcher generation. When a specific caller genuinely needs to see its own write, do not lower the global interval. Track the write's sequence number and let that caller wait for a searcher that includes it; `ControlledRealTimeReopenThread` supports exactly this with a relaxed interval for background refreshes and a tighter one for waiters. That way a rare requirement is paid for by the requests that have it. ## Commit: rarely, with a log underneath A commit fsyncs every file in the new commit point. On a large index that is expensive and it does nothing for visibility, which the NRT reader already handles. So commit on a slow cadence — measured in seconds to minutes depending on how much replay you can tolerate — and get per-operation durability from a write-ahead log maintained above Lucene, replaying anything after the last commit on restart. This is exactly the architecture the engines built on Lucene use, and it is the right answer to give: durability is not Lucene's commit's job alone. If you take backups, remember that pinning a commit with a snapshot deletion policy keeps its files alive; a long-held snapshot inflates disk usage because merged-away inputs cannot be deleted. ## Merges: provision, do not micro-tune The instinct is to tune `TieredMergePolicy`. Usually the defaults are fine and the real constraint is I/O. The signals to watch are segment count relative to the policy's budget, the deleted-document percentage, and merge-thread saturation. The failure mode to recognise is back-pressure: `ConcurrentMergeScheduler` stalls indexing threads once pending merges exceed its limit, so a storage device that cannot sustain merge writes presents as mysterious ingest latency, not as a merge alarm. When you do adjust: lowering `segmentsPerTier` buys query speed with more merge I/O and only makes sense if you have that I/O spare; raising `maxMergedSegmentBytes` yields fewer, larger segments at the cost of very long individual merges and more disk headroom; raising merge-thread count helps on SSD or NVMe and hurts on spinning disks where seeks dominate. Every one of these is a trade, and saying so explicitly is what distinguishes a considered answer. ## Separate the lifecycles The strongest structural move is to stop asking one index to be both write-optimised and read-optimised. Write into a current index; when it rolls over and stops taking writes, compact it once and leave it alone. That converts an unbounded steady-state merge cost into a one-off, schedulable cost, makes retention a file deletion instead of a mass delete-and-merge, and means the heavy compaction never competes with peak ingest. ## Capacity planning Budget three resources explicitly: I/O bandwidth for merges as a first-class line item alongside ingest and query; free disk equal to at least the largest permitted merge output, plus room for tombstones between merges; and heap for RAM buffers and per-segment structures. Set alerts on segment count, deleted percentage and merge backlog. Then verify the whole thing under sustained load rather than a burst — merge pressure is a steady-state phenomenon, and a benchmark that ends before the first large merge lands has measured nothing.

  • Ingest latency spikes periodically with no change in traffic. How does merging enter your hypothesis list?
    ConcurrentMergeScheduler deliberately stalls indexing threads when pending merges exceed its limit, so a large merge saturating disk shows up as ingest latency rather than as a merge alert. Correlate the spikes with merge activity, segment counts and device utilisation. If they line up, the fix is I/O headroom or a lower maximum merged segment size, not indexing-side tuning.
  • Product asks for sub-second freshness on a stream of 50,000 documents per second. How do you respond?
    Ask what the requirement really is, since near-real-time reopening can meet sub-second visibility but each reopen cuts a segment and multiplies merge work. I would set a bounded interval that meets the stated need, give the few flows needing read-your-own-writes a wait-for-generation path, and show the query-latency cost of going lower so the tradeoff is chosen rather than assumed.
  • How does separating a hot write index from compacted read-only indexes change the economics?
    It turns a permanent steady-state merge cost into a one-off compaction you schedule off-peak. The read-only indexes reach their best possible shape once and stay there, retention becomes deleting whole indexes rather than mass deletes plus merges, and heavy compaction never competes with peak ingest for the same I/O.

saying these in an interview costs you the question

  • Treats refresh, commit and merge as one setting
  • Chases zero-latency freshness without costing the merge work
  • Tunes merge policy before checking I/O saturation
  • Plans no free-disk headroom for merges in flight
  • Assumes durability comes from Lucene commits alone

context