skip to content

In Lucene, what is the difference between postings, doc values, and stored fields for the same field?

level: middleimportance: must knowfreq 68%

answer

  1. The same field is written more than once
  2. Three different questions, three structures
  3. One maps term to docs, one docs to value
  4. Stored fields are only read for the top hits
  5. Sorting and faceting read the columnar copy

basics

~20 s

Postings are the inverted index that maps a term to the documents containing it. Doc values are a columnar per-document copy used for sorting, faceting and scripts. Stored fields are the row-wise original bytes returned for the documents you display.

solid answer

~50 s

Lucene writes the same field through several independent paths, and each answers a different question. - **Postings** (`.tim`/`.tip` term dictionary plus `.doc`/`.pos`/`.pay`) answer *which documents contain this term*. The mapping is one-way: you cannot ask postings what value document 42 holds without scanning every term. - **Doc values** (`.dvd`/`.dvm`) are the column-oriented inverse: for each document, its value(s) for that field. That is what sorting, faceting, grouping and script access read. They are written once per segment and read off-heap through the filesystem cache. - **Stored fields** (`.fdt`/`.fdx`) hold the verbatim bytes you passed in, chunk-compressed, row-oriented. They are only fetched for the handful of documents you actually return. So a typical query touches postings to match, doc values to sort or aggregate, and stored fields for the top *N* hits only. Enabling all three for a field you only ever filter on is pure waste.

code

java · 7 lines
java
Document doc = new Document();
// searchable, but not sortable and not returnable
doc.add(new StringField("status", "open", Field.Store.NO));
// sortable / facetable column
doc.add(new SortedDocValuesField("status", new BytesRef("open")));
// returnable copy for the top hits
doc.add(new StoredField("status", "open"));

go deeper

for a junior

Recall that a field can be searchable, sortable and retrievable, and that those are separate switches. Knowing that returning a value and searching for it use different structures is enough at this stage.

for a middle

Be ready to name the three structures, say which operation reads each, and explain why the inverted index alone cannot answer a per-document value lookup. Expect a follow-up on why sorting needs doc values.

for a senior

Show that you use this model to size indexes and diagnose failures: an unexpected doc-values-type error, an index three times larger than the source data, or a sort that suddenly costs heap. Say plainly that these choices need a reindex to change.

for a principal

Own the tradeoff at fleet scale: which fields pay for which structures is a contract between the query workload and the storage budget, and it is irreversible without a reindex pipeline. Argue for measuring per-structure file sizes rather than guessing.

## The same value, written several ways Lucene never stores a field just once. When `IndexWriter` processes a document, each field is routed through as many independent write paths as its `FieldType` requests: an inverted path (postings), a columnar path (doc values), a row path (stored fields), plus optional norms, points and term vectors. Those paths land in different files inside the segment, are read by different APIs, and are tuned for opposite access patterns. Understanding which structure serves which operation is the single most useful mental model for reasoning about index size and about errors like *"unexpected docvalues type NONE"*. ## Postings: term to documents An analyzed field is broken into terms; the terms are sorted and written to a term dictionary (`.tim`), indexed by a compact automaton (`.tip`), and each term points at a postings list of document IDs (`.doc`), optionally with positions (`.pos`) and offsets/payloads (`.pay`) depending on the field's `IndexOptions`. This is the structure that makes search fast: for the term `kotlin`, Lucene jumps straight to a delta-compressed, skip-list-accelerated list of matching document IDs and their term frequencies, and never looks at non-matching documents. The crucial limitation is direction. Postings map term to documents. There is no cheap way to invert them at query time — asking "what is document 42's value for `status`?" from postings alone means enumerating terms until you find one whose list contains 42. That is why the second structure exists. ## Doc values: document to value Doc values are Lucene's columnar store, written per segment at flush time and never mutated afterwards. The types are `NUMERIC`, `SORTED_NUMERIC`, `BINARY`, `SORTED` and `SORTED_SET`. The two `SORTED*` variants do not store the raw bytes per document at all; they store a sorted set of distinct values once, and per document only an *ordinal* — a small integer pointing into that set. Comparing ordinals is what makes sorting and faceting on a string field cheap, because most of the work happens on packed integers rather than on byte strings. Doc values are consumed by sorting (`SortField`), faceting and aggregation, grouping, function/score computation, script field access, and join-style lookups. Since Lucene 7 they are exposed as a forward-only iterator (`DocValuesIterator` extends `DocIdSetIterator`), which cooperates well with the document-at-a-time query execution model but means code that needs random access has to re-acquire the iterator. Because doc values are packed and bit-compressed (delta, GCD and table encodings are applied where they fit), they are usually far smaller than the equivalent stored-field copy — but they are still a full extra column on disk. ## Stored fields: document to original bytes Stored fields are the row store. Values handed to a `StoredField` are written verbatim into chunks in `.fdt`, with an offset index in `.fdx` and metadata in `.fdm`, and the chunks are compressed. They are not searchable, not sortable, and not aggregatable — they exist purely so that after ranking you can hand the caller back what was indexed. Because reads happen only for the documents you return, stored fields tolerate high compression ratios in a way that doc values (read for every matching document) do not. ## Why not just one structure Consider a query that filters on `status`, sorts by `price`, and displays `title`. Postings on `status` cut the candidate set. Doc values on `price` feed the comparator for every surviving document. Stored fields on `title` are decompressed for the ten documents on page one. A single structure cannot be simultaneously good at set intersection, at dense sequential per-document reads, and at bulk-compressed random retrieval — so Lucene keeps three, and lets the schema decide which ones a field pays for. ## Practical consequences - A field indexed but without doc values will match queries and then fail at sort or facet time with an error naming the expected doc-values type. The fix is a reindex; doc values cannot be added to existing segments. - A field with doc values but not indexed can be sorted and aggregated but not efficiently searched. - A field neither indexed nor doc-valued but stored is display-only. - Enabling everything everywhere is the most common cause of a Lucene index that is several times larger than the source data. Every path you enable is a separate copy with its own encoding. The search engines built on Lucene expose these as per-field schema flags, which is why flipping one flag can change index size by tens of percent.

  • Why can't Lucene just build doc values on the fly from the postings when a sort needs them?
    It can, but only by un-inverting the whole field into memory — walking every term and every posting to build a per-document array. That is what the old FieldCache did, and it was slow, unbounded in heap, and rebuilt per segment reader. Doc values move that work to index time, write it compactly to disk, and let the OS page cache manage it.
  • Why do SORTED and SORTED_SET doc values store ordinals instead of the values themselves?
    An ordinal is a small integer index into the segment's sorted set of distinct values. Storing ordinals shrinks the per-document column dramatically for low-cardinality fields, and makes comparisons and facet counting integer operations rather than byte-string comparisons. The raw terms are stored once per segment and resolved only when a value must be displayed.
  • Which structure does a highlighter read from?
    It depends on the highlighter. Re-analysing highlighters read the stored field and run the analyzer over it again. Offset-based highlighters read positions and offsets from the postings if the field indexed them, or from term vectors if those were stored. None of them use doc values.

saying these in an interview costs you the question

  • Thinking stored fields are searchable because they are in the index
  • Believing doc values are just a heap cache of the postings
  • Assuming enabling one structure implies the others
  • Claiming a field must be stored to be sortable
  • Saying doc values can be added later without reindexing

context