How does Lucene compress stored fields, and what does choosing BEST_COMPRESSION trade away?
answer
- It is the row store, read after ranking
- Documents are not compressed one at a time
- Fetching one document costs more than its own bytes
- Two modes: fast versus small
- The choice is recorded per segment, not per index
basics
~20 sLucene groups consecutive documents into chunks and compresses each chunk as a unit, so fetching one document decompresses its whole chunk. BEST_SPEED uses fast LZ4; BEST_COMPRESSION uses a higher-ratio DEFLATE-family mode that shrinks the index but makes retrieval and merging measurably slower.
solid answer
~50 sStored fields are Lucene's row store: `.fdt` holds the compressed document data, `.fdx` indexes chunk offsets, `.fdm` holds metadata. Documents are **not** compressed individually — they are batched into chunks and compressed together, because tiny documents compress badly on their own and share a lot of structure with their neighbours. The cost is that retrieving a single document means decompressing the chunk it lives in. `Lucene90StoredFieldsFormat` (Lucene 9.x) offers two modes. `BEST_SPEED` uses **LZ4** with small chunks: fast decompression, modest ratio. `BEST_COMPRESSION` uses **DEFLATE** with larger chunks and preset dictionaries: substantially smaller on disk, but slower to decompress on every fetch and slower to merge, since merging re-encodes the data. The mode is part of the **Codec**, and a codec is recorded per segment. Changing it therefore affects only segments written afterwards — existing segments keep their format until they are merged or the data is reindexed.
code
java · 4 linesIndexWriterConfig cfg = new IndexWriterConfig(analyzer);
// smaller on disk, slower to fetch and to merge
cfg.setCodec(new Lucene912Codec(Lucene912Codec.Mode.BEST_COMPRESSION));
IndexWriter writer = new IndexWriter(dir, cfg);go deeper
Recall that the values returned in search results are stored separately from the searchable index and are compressed on disk.
Be ready to explain chunk-level compression and the read amplification it implies, and to name the fast versus high-ratio modes and what each costs.
Demonstrate the diagnosis: a fetch phase that scales with documents returned, and the codec-is-per-segment rule that explains why a settings change does not shrink the index right away.
Own the tiering policy — which indexes are cold enough to pay decompression latency for disk savings, how merge cost changes, and how you migrate an existing corpus without a disruptive full rewrite.
## What stored fields are for Stored fields hold the verbatim values you asked Lucene to keep so you can return them. They are not searchable, not sortable and not aggregatable — those jobs belong to postings and doc values. Their entire access pattern is: rank first, then fetch the handful of documents on the current page. That pattern is what justifies aggressive compression; you decompress for ten documents, not for ten million. ## Chunked layout Three files make up the format in Lucene 9.x: - `.fdt` — the field data, written as a sequence of compressed **chunks**, each covering a run of consecutive document IDs. - `.fdx` — the index from document ID to the chunk containing it, itself compactly encoded (monotonic values compress well). - `.fdm` — format metadata. The decision to compress *chunks* rather than *documents* is the design's centre of gravity. Individually, a small JSON document is close to incompressible: the compressor has no history to work with and the dictionary overhead swamps the payload. Batched with neighbouring documents, the same field names, punctuation and repeated values appear over and over, and the ratio improves dramatically. The price is read amplification. Fetching document 1,000,417 means locating its chunk and decompressing the entire chunk, including the neighbours you did not ask for. For a top-10 page that is trivially acceptable. For a query that retrieves stored fields of tens of thousands of documents — a naive export, a deep page, a script that reads source per hit — it becomes the dominant cost, and this is a classic production diagnosis: the search is fast, the fetch phase is not. ## The two modes `Lucene90StoredFieldsFormat` takes a `Mode`: - **`BEST_SPEED`** — LZ4 over relatively small chunks. LZ4 decompresses extremely fast, at a modest ratio. This is the default for general-purpose indexes because it keeps the fetch phase cheap. - **`BEST_COMPRESSION`** — DEFLATE over larger chunks, with preset dictionaries built from the data to help the encoder. Substantially smaller on disk. The costs are real and threefold: slower decompression on every retrieval, larger decompression units so more wasted work per single-document fetch, and heavier merges, because merging stored fields means decoding and re-encoding them. A good rule of thumb: choose the high-compression mode when the data is large, cold and retrieved rarely relative to its volume — log and archival indexes are the canonical case. Keep the fast mode where the fetch phase is on the user-facing path and every millisecond counts. ## Where the mode lives: the Codec A Lucene **`Codec`** is the bundle of formats that decides how a segment is encoded: `PostingsFormat`, `DocValuesFormat`, `StoredFieldsFormat`, `TermVectorsFormat`, `NormsFormat`, `PointsFormat`, `KnnVectorsFormat`, plus the metadata formats. Codecs are named, discovered through the Java service-provider mechanism, and — critically — **the codec name is recorded in the segment's metadata when the segment is written**. That has a direct operational consequence. Changing the stored-fields mode does not rewrite anything. Segments already on disk keep their existing encoding and are read back through the codec they name; only newly flushed segments, and merged segments (which are new segments), use the new setting. So the index converges on the new mode as merging proceeds, and a full conversion means forcing a rewrite or reindexing. It also means an index can legitimately contain segments in several codecs at once, which is what makes reading indexes written by an older major version possible through the backwards-codecs module. Per-field variation is possible too: the per-field wrappers let different fields use different postings or doc-values formats within one segment. Custom codecs are a real extension point, but writing one is rare — the usual lever is picking a mode on the standard format. ## Term vectors sit next door Term vectors (`.tvd`/`.tvx`) are a separate, and much more expensive, per-document structure: a miniature inverted index for one document, listing its terms with optional frequencies, positions and offsets. They exist for highlighting and for term-based similarity. They are chunk-compressed like stored fields, and they are frequently enabled by reflex and never used — a large, silent tax. Modern offset-based highlighting can usually read positions and offsets from the postings instead, if the field indexed offsets, which is generally cheaper than carrying full term vectors. ## How to reason about it in an interview Say the shape first: chunk-compressed row store, read only for returned documents, so ratio is worth more than latency — up to the point where you start fetching many documents. Then name the two modes and what each trades. Then close with the codec point, because it is the one people miss: the setting is recorded per segment, so a change is gradual and retroactive only through merging or a reindex.
- Why does Lucene compress chunks of documents rather than each document separately?A small document offers a compressor almost no history to exploit, so per-document compression yields poor ratios. Batching consecutive documents lets repeated field names, punctuation and values be encoded once across the chunk. The tradeoff is read amplification: fetching one document decompresses the whole chunk it belongs to.
- You switch an index to the high-compression stored-fields mode. When does the index actually get smaller?Gradually. The codec is recorded per segment, so existing segments keep their old encoding and only newly written segments use the new mode. The index converges as merging rewrites old segments; converting everything immediately means forcing a rewrite or reindexing.
- A search is fast but the response is slow, and the query returns thousands of documents with their stored fields. What is happening?The fetch phase dominates. Every returned document requires decompressing its stored-fields chunk, including neighbouring documents in that chunk, so retrieval cost scales with hits returned rather than with hits matched. The fixes are returning fewer documents per request, returning fewer fields, or serving the needed values from doc values instead.
- When are term vectors worth their cost?Rarely. They are a per-document inverted index used for highlighting and term-based similarity, and they are one of the most expensive optional structures. If the field indexes offsets, an offset-based highlighter can work from the postings instead. Enable term vectors only when a specific highlighter or similarity feature measurably needs them.
saying these in an interview costs you the question
- Thinking each document is compressed independently
- Believing a codec change rewrites existing segments
- Assuming high compression only costs disk, not query time
- Confusing stored fields with doc values as a retrieval path
- Enabling term vectors by default for highlighting