skip to content

In a multi-terabyte Lucene index, how would you budget which per-field structures each field keeps?

level: principalimportance: should knowfreq 26%

answer

  1. Each capability implies exactly one structure
  2. Cache residency matters more than disk cost
  3. Positions and term vectors are the usual fat
  4. Measure by file extension on a sample segment
  5. None of it can be changed in place

basics

~20 s

Start from a written query contract: for each field, which of match, phrase, score, sort, aggregate, highlight and return it must support. Enable only the structures those verbs require, measure the per-structure file sizes on a representative sample, and treat every choice as reindex-only.

solid answer

~60 s

Each field can pay for up to six independent structures, and each maps to a capability: - **postings** — matching; the `IndexOptions` level decides whether frequencies, positions and offsets are written, and positions alone are often the largest single line item on a text field. - **norms** — length normalization for scoring; useless on fields only ever filtered. - **doc values** — sorting, faceting, grouping, scripts. - **points** — numeric, date and geo range matching. - **stored fields** — returning the value. - **term vectors** — highlighting and term-based similarity; almost always the first thing to cut. The method is: write the query contract per field, delete the structures no verb needs, then *measure* — build a representative sample index, force it into one segment, and compare on-disk sizes by file extension so the argument rests on bytes rather than intuition. The strategic constraint dominates the tactical one: none of this can be changed in place. Segments are immutable, so every decision is only revisable through a reindex. The lead's real job is to keep the reindex pipeline cheap and routine, so a wrong storage bet costs a rerun rather than a migration project.

code

java · 5 lines
java
FieldType body = new FieldType(TextField.TYPE_NOT_STORED);
body.setIndexOptions(IndexOptions.DOCS_AND_FREQS); // no positions: no phrase queries
body.setOmitNorms(false);        // still scored, keep length normalization
body.setStoreTermVectors(false); // highlight from postings offsets if ever needed
body.freeze();

go deeper

for a junior

Recall that a searchable field, a sortable field and a returnable field are separate switches, and that turning all of them on for every field makes the index much larger than the data.

for a middle

Be ready to map each capability to the structure it needs and to name the two usual sources of waste: positions on fields nobody phrase-searches, and term vectors nobody uses.

for a senior

Show the measurement discipline — a sample index, one segment, sizes compared by file extension — and the diagnosis that unnecessary structures push the working set out of the file-system cache.

for a principal

Own the strategy: derive the schema from an explicit query contract, tier it by data temperature, and argue that because every choice is reindex-only, the reindex pipeline is the real investment that makes schema optimisation safe to attempt.

## Framing the problem At multi-terabyte scale, index size is not a disk-cost question — disk is the cheapest part. It is a **cache-residency** question. Search latency is dominated by whether the structures a query touches are in the file-system cache. Every unnecessary structure competes for that cache with the structures that matter, so trimming the schema is a latency intervention as much as a storage one. It also compounds: bigger segments make merging slower, snapshots larger, recovery longer and replica rebuilds more expensive. ## The per-field cost model For every field, ask which of these capabilities are genuinely required, and note which structure each one implies: | Capability the product needs | Structure it requires | |---|---| | match a term or a filter | postings | | phrase / proximity / span matching | postings **with positions** | | offset-based highlighting | postings with offsets, or term vectors | | relevance scoring by length | norms | | sort, facet, group, script access | doc values | | numeric, date or geo range | points | | return the value to the caller | stored fields | | term-based similarity, some highlighters | term vectors | The two most common wins are at the top and the bottom of that table. **Positions** are frequently paid for on fields nobody phrase-searches; dropping to a frequencies-only or docs-only `IndexOptions` removes an entire postings file's worth of data on a large text field. **Term vectors** are frequently enabled once, for a highlighting experiment, and never removed. The cheap wins are smaller but free: norms on filter-only fields, doc values on fields never sorted or aggregated, stored copies of fields the API never returns. ## Deciding rather than guessing A principal-level answer has a method, not a list of tips. **1. Write the query contract.** For each field, enumerate the verbs the product actually performs against it. Not "someone might want to sort on this one day" — what is on the roadmap and what is in the API. This document is the artefact the schema is derived from, and the thing you point at when someone asks why a field is not sortable. **2. Derive the structures mechanically.** With the contract in hand the schema stops being a matter of taste; each verb turns on exactly one structure. **3. Measure on a sample.** Build a representative sample — real data, real analysis chain, enough documents that ratios stabilise — force it into a single segment, and list file sizes by extension. Postings, term dictionary, doc values, points, stored fields, norms and term vectors each live in their own extensions, so the breakdown is directly readable. Toggle one structure, rebuild, compare. Tools that inspect a Lucene index directly, or the engine's own disk-usage reporting, do this without hand-rolling anything. **4. Model the whole lifecycle, not the steady state.** A structure's cost is paid again on every merge, in every snapshot, and during every replica recovery. A 20% index reduction shortens all of those. **5. Decide the reversibility budget.** Which decisions do you want to be able to undo, and what would undoing them cost? ## The irreversibility constraint This is what separates a senior answer from a principal one. Every structure listed above is written at index time into an immutable segment. There is no operation that adds doc values to existing documents, turns norms back on, or starts writing positions for text already indexed. Merging only rewrites what is already there; it does not synthesise missing structures. So the real leadership decision is not the schema itself — it is **how expensive a wrong schema is to correct**. If reindexing the corpus is a routine, automated, well-rehearsed operation, you can trim aggressively and reverse a bad call in a day. If reindexing is a multi-week project nobody has run recently, every trim is a gamble and the rational move is to keep more than you need. Investing in the reindex pipeline is what buys the freedom to optimise the schema at all, and that inversion is the point worth making out loud. ## Tiering instead of a single answer One schema for all data is usually the wrong shape at this scale. Data ages, and its query contract ages with it: - **Hot data** — full capability, fast stored-fields compression, everything the product might need on the interactive path. - **Warm/cold data** — capabilities the product actually still exercises on old data, high-ratio stored-fields compression, positions and term vectors dropped if nothing phrase-searches or highlights history. Because the codec and the schema are fixed per segment, this is naturally implemented by writing new time-based indexes with different settings rather than by mutating anything. ## Failure modes to name - Enabling everything "to be safe", then discovering the index is several times the source data and no longer cache-resident. - Trimming a structure with no measurement, and discovering the saving was two percent while the lost capability was material. - Dropping positions on a field that a phrase query silently depends on, so a feature quietly stops matching rather than erroring. - Making the decision once, at project start, and never revisiting it as the query mix changes. ## The one-sentence version Derive the schema from an explicit query contract, prove each choice with measured per-extension sizes on a representative sample, tier the settings by data temperature, and — because every one of these choices is reindex-only — invest in making reindex cheap so the schema stays a decision you can revise rather than one you have to get right forever.

  • Which single change usually yields the biggest index-size reduction on a large text corpus?
    Lowering the `IndexOptions` level on large analysed fields that nobody phrase-searches. Positions are typically the biggest line item in the postings, and dropping them removes an entire file's worth of data. Term vectors, when enabled and unused, are the other frequent offender. Both are wasted only if you can prove no query depends on them.
  • How would you prove a proposed schema trim is worth it before shipping it?
    Build a sample index from real data with the real analysis chain, force it to a single segment, and record on-disk size per file extension. Rebuild with the structure disabled and compare. Then replay a representative query mix against both to confirm no capability regressed and that latency improved rather than merely disk usage.
  • Why is investing in the reindex pipeline the precondition for optimising the schema?
    Every per-field structure is written into immutable segments and cannot be added or removed in place. If reindexing is cheap and rehearsed, a bad storage bet costs one rerun and you can trim aggressively. If it is a rare, risky project, the rational response is to over-provision the schema — so pipeline investment is what actually unlocks the savings.
  • How does data temperature change the answer?
    Old data usually has a narrower query contract than fresh data: still filtered and aggregated, rarely highlighted or phrase-searched. Because the schema and codec are fixed per segment, you express that by writing new time-based indexes with trimmed structures and high-ratio stored-field compression, rather than trying to change settings on existing data.

saying these in an interview costs you the question

  • Enabling every structure by default to stay flexible
  • Assuming a schema flag can be flipped without reindexing
  • Trimming structures without measuring the actual saving
  • Treating index size as a disk-cost rather than cache problem
  • Applying one schema to hot and cold data alike

context