What are norms in a Lucene index, and how does BM25Similarity use them at query time?
answer
- Something about field length is written at index time
- One value per document per field
- Compressed very aggressively — think one byte
- It feeds BM25's denominator
- The b parameter decides how much it counts
basics
~20 sNorms are a per-document, per-field length value written at index time, one byte per document under Lucene's default similarity. BM25Similarity reads that byte to normalize scores by field length, so a term match in a short field outranks the same match in a long one.
solid answer
~50 sNorms are the index-time half of length normalization. While inverting a document, Lucene calls `Similarity.computeNorm(FieldInvertState)`; `BM25Similarity` returns the field's token count, lossily compressed into a **single byte** via `SmallFloat`, and the `NormsFormat` writes it to `.nvd`/`.nvm`. By default overlapping tokens — synonyms injected at position increment zero — are discounted from that length. At query time, BM25's denominator is `k1 * (1 - b + b * dl / avgdl)`, where `dl` is the decoded norm and `avgdl` the average field length in the segment. That is how a match in a five-word title beats the same match in a thousand-word body. The operational consequence: `k1` and `b` are query-time parameters and can be changed freely, but the *norm itself* is baked into the segment. Turning norms off saves roughly a byte per document per field and removes length normalization entirely — and turning them back on requires reindexing.
code
java · 7 lines// index time: one byte per doc per field, written by NormsFormat
public long computeNorm(FieldInvertState state) {
int numTerms = discountOverlaps
? state.getLength() - state.getNumOverlap()
: state.getLength();
return SmallFloat.intToByte4(numTerms);
}go deeper
Recall that Lucene scoring prefers a match in a short field over the same match in a long one, and that the length information for this is recorded when the document is indexed.
Be ready to say what computeNorm writes, that it is a single lossy byte, and where the value lands in the BM25 formula. Know which BM25 parameter turns length normalization up or down.
Demonstrate the index-time versus query-time split in a production argument: which relevance changes are a config edit and which force a reindex, and how you would spot inconsistent scoring across segments written under different settings.
Frame it as a schema policy question: which fields are scored at all, what the norm byte buys across a fleet of billion-document indexes, and how you keep relevance-affecting index-time decisions reversible through a cheap reindex path.
## What a norm actually is A norm in Lucene is a single value stored per document, per field, describing how "long" that field was when it was indexed. It exists so that scoring can discount matches in long fields. Without it, a query term appearing once in a book-length body would score the same as the same term appearing once in a short title, and relevance would collapse toward whichever documents happen to contain the most text. Norms are produced during inversion, not during search. As `IndexWriter` analyzes a field it accumulates a `FieldInvertState` — token count, position increments, overlap count, unique term count — and then calls `Similarity.computeNorm(FieldInvertState)`, which returns a `long`. That value goes through the `NormsFormat` into the segment's `.nvd` (data) and `.nvm` (metadata) files, alongside doc values in spirit: they are read per matching document as a forward iterator. ## The one-byte encoding `BM25Similarity.computeNorm` takes the field length — by default `state.getLength() - state.getNumOverlap()`, i.e. tokens minus those injected at the same position, which keeps synonym filters from inflating the apparent length — and compresses it with `SmallFloat` into a **single byte**. That is a deliberately lossy, roughly logarithmic bucketing: short lengths are represented precisely, long ones coarsely. A 900-token field and a 1000-token field very often land in the same bucket and receive an identical length penalty. Candidates are frequently surprised by this. They expect a precise integer length and try to explain a scoring difference by a few tokens. The right answer is that below a certain granularity the norm cannot distinguish them at all. The payoff of the lossy encoding is size: one byte per document per field, decoded through a small lookup table, cheap enough to read for every scored hit. ## How BM25 consumes it Lucene's `BM25Similarity` scores a term as roughly: ``` idf * tf / (tf + k1 * (1 - b + b * dl / avgdl)) ``` with `idf = ln(1 + (docCount - docFreq + 0.5) / (docFreq + 0.5))`. Here `tf` is the term frequency from the postings, `dl` is the decoded norm, and `avgdl` is the average field length derived from segment-level statistics. The defaults are `k1 = 1.2` and `b = 0.75`. - `k1` controls **term-frequency saturation**: how quickly repeating a term stops helping. - `b` controls **how strongly length normalization applies**. At `b = 0` the norm is ignored entirely; at `b = 1` the penalty is fully proportional to length over average length. Both are read at query time from the `Similarity` instance, so they can be changed without touching the index. The norm itself cannot. ## Index-time versus query-time, and why it matters This split is the interview point. If you replace the similarity with one whose `computeNorm` encodes something different — a different length definition, a different discretization — documents already in the index still carry norms written under the old rule, and their scores will be wrong until those segments are rewritten. Changing `k1`/`b` on `BM25Similarity`, by contrast, takes effect on the next query for every document. ## Omitting norms A field can be indexed with norms disabled (`FieldType.setOmitNorms(true)`, exposed as a schema flag by the engines built on Lucene). Three things follow: 1. Roughly one byte per document per field is saved. On a billion-document index across a dozen analyzed fields that is a meaningful, though not dramatic, saving. 2. Length normalization disappears for that field: BM25 behaves as if every document's field were of average length. 3. Fields that were never scored — a status code, a tenant ID, anything used purely as a filter — lose nothing, because filter clauses do not score. Because norms are written at index time, flipping the flag only affects newly written segments; existing documents keep whatever they were indexed with, and mixed segments will score inconsistently until a reindex. ## Common failure modes - A field that is only filtered on still carries norms because nobody turned them off; pure waste, but harmless to relevance. - A synonym filter is applied and someone expects the field to look longer; by default overlapping tokens are discounted, so it does not. - Someone concludes from a scoring difference of a few points that norms track exact token counts. They do not — the encoding is bucketed. - Someone disables norms on the main body field to save space and is then surprised that long documents dominate results.
- Which BM25 parameters can you change without reindexing, and which cannot?`k1` and `b` live on the Similarity instance and are applied at query time, so changing them affects every document immediately. The norm value is written into the segment at index time by `computeNorm`, so a similarity that encodes length differently only applies to newly written segments — existing documents keep their old norms until reindexed.
- Why does Lucene discount overlapping tokens when computing the norm?Tokens emitted at position increment zero — synonyms, stem variants, injected forms — do not make the text longer for a human reader; they are alternative surface forms at the same position. Counting them would penalize documents merely for having synonym expansion applied, so by default `BM25Similarity` subtracts the overlap count from the length.
- When is it right to disable norms on a field?When the field is never scored: identifiers, status codes, tenant keys and anything used only inside filter clauses. Filters contribute no score, so length normalization is irrelevant and the byte per document per field is pure overhead. Never disable them on the fields users actually search as free text.
saying these in an interview costs you the question
- Thinking norms store the exact token count as an integer
- Believing norms are computed at query time from postings
- Claiming disabling norms only affects index size, not scoring
- Saying k1 controls length normalization and b controls saturation
- Assuming toggling norms rescores documents already indexed