Beyond document IDs, what does a postings list store, and what does each part enable?
answer
- More than just which documents matched
- One number lets you rank the matches
- Another lets you match words in order
- Character coordinates support the highlighter
- Document frequency is not stored per posting
basics
~10 sA posting carries the document ID plus, optionally, the term frequency in that document for scoring, the token positions for phrase and proximity matching, character offsets for highlighting, and sometimes per-position payloads.
solid answer
~50 sA postings list is layered, and engines let you choose how many layers to pay for. The minimum is the **document ID** — enough for pure filtering. Add the **term frequency**, the count of occurrences in that document, and the index can score with a model like BM25. Add **positions**, the token ordinals of each occurrence, and phrase and proximity queries become possible, because matching a phrase means checking that consecutive terms occupy consecutive positions. Add **character offsets**, the start and end character of each occurrence, and the engine can highlight matched text without re-analysing the stored content. Some engines also support **payloads**: arbitrary bytes attached to a position, used for things like per-occurrence weights. Note what is *not* in the postings list: the term's document frequency lives once in the term dictionary entry, and the document's field length, needed for length normalisation, is stored per document separately.
code
json · 10 lines{
"quick": {
"doc_freq": 3,
"postings": [
{ "doc": 1, "tf": 2, "positions": [4, 17], "offsets": [[22, 27], [95, 100]] },
{ "doc": 7, "tf": 1, "positions": [0], "offsets": [[0, 5]] },
{ "doc": 9, "tf": 1, "positions": [31], "offsets": [[188, 193]] }
]
}
}go deeper
Recall the layers by the capability each unlocks: document ID for matching, term frequency for ranking, positions for phrases, offsets for highlighting. Naming them in that order is a solid answer at this level.
Explain the mechanics: how a phrase match is a position-list intersection with an expected offset, and why token positions and character offsets are different coordinate systems that analysis pulls apart.
Be ready to argue per-field choices on a real schema — which fields drop positions, which get offsets versus term vectors versus query-time re-analysis for highlighting, and what reindexing that decision later would cost.
Own the sizing story: positions scale with token occurrences and dominate index footprint, so the postings layout you standardise across teams is a long-term storage and latency commitment, not a per-index detail.
## What a posting is A postings list is the value side of the inverted index: for one term, the ascending sequence of documents containing it. Each element of that sequence is a *posting*. The interesting design question is how much a posting carries beyond the document ID, because every extra field multiplies across every occurrence of every term in the corpus. ## Layer 1 — the document ID Always present, always ascending. Ascending order is what makes intersections and unions linear merges, and it is what lets the list be delta-compressed. A field indexed at this level only can answer "does this document contain the term?" — enough for filters and boolean matching, and nothing else. ## Layer 2 — term frequency The number of times the term occurs in that document. This is the *tf* that every classical ranking model consumes. Without it, the engine can tell you a document matched but cannot rank matches against one another on textual evidence; scoring collapses to something constant per hit. Storing it costs one small integer per posting, and it compresses well because most values are 1. ## Layer 3 — positions The token ordinal of each occurrence: if the analysed document is `[the, quick, brown, fox]`, then `brown` occurs at position 2. There are *tf* positions per posting, so this layer scales with total token occurrences in the corpus, not with document count — which is why it is usually the largest part of a positional index. Positions buy the queries that depend on word order and nearness: - **Phrase queries.** Matching `"machine learning"` means finding a document where `machine` occurs at some position *p* and `learning` at *p+1*. The engine intersects the two documents' position lists with the expected offset. - **Proximity / slop queries.** Same mechanism with a tolerance on the distance, and often with a scoring bonus for closer occurrences. - **Span and ordered-window constructs**, and some phrase-aware relevance features. Remove positions and a phrase query cannot be answered correctly at all; the best the engine could do is match the words independently, which is a different query. ## Layer 4 — character offsets The start and end character index of each occurrence in the original field text. Positions are *token* coordinates; offsets are *character* coordinates, and they differ whenever analysis removes, merges or rewrites tokens — a stemmer maps `running` to `run`, but the offsets still point at the original six characters. Offsets exist to support **highlighting**: rendering the snippet with the matched words marked. There are three ways to get them and they trade storage for query cost: store offsets in the postings (biggest index, fastest highlight), store a per-document term vector (a forward structure, moderate cost), or re-analyse the stored field at query time for just the top few hits (no index cost, more CPU per result page). Highlighting only ever runs on a page of results, which is why re-analysis is often the right default. ## Layer 5 — payloads Some engines allow arbitrary bytes attached to a specific position, typically written by the analysis chain. Uses include per-occurrence weights (a term appearing in bold or in a title-cased span), part-of-speech tags, or entity IDs, consumed by a custom scoring function. This is a specialist feature and the one you should mention last. ## What lives elsewhere Two numbers that ranking needs are deliberately *not* in the postings list: - **Document frequency** — how many documents contain the term. It is a property of the term, so it is stored once in the term dictionary entry, not repeated in every posting. It drives IDF. - **Field length** — the number of tokens in the field, needed by length normalisation. It is a property of the document and is kept in a per-document structure, so scoring reads it by document ID. Getting this split right in an interview signals that you have actually reasoned about the layout rather than recited a list. ## The engineering consequence Because the layers are additive and independent, most engines let you choose them per field. A status code, a tag or an identifier is a filter field: document IDs only. A product title takes phrase queries and is short: positions on. A large body field that is searched only as a bag of words but never phrased can drop positions and shrink dramatically. The decision is per field and reversible only by reindexing, so it belongs in the mapping design conversation, not in a later tuning pass.
- Why are positions usually the largest part of a positional inverted index?Because positions scale with token occurrences rather than documents. A posting contributes one document ID but *tf* position values, so a term appearing ten times in a document costs ten position entries. Summed over a corpus, that is roughly one entry per indexed token, which dwarfs the one-entry-per-matching-document cost of the document ID layer.
- If you have positions, why would you ever also store character offsets?Because positions are token coordinates and highlighting needs character coordinates in the original text. Analysis breaks the correspondence: stemming, stopword removal, synonym injection and character filters all shift or merge tokens. Offsets record where each occurrence really started and ended, so the highlighter can mark the source text without re-running analysis.
- What breaks if a field is indexed with document IDs but no term frequencies?Textual relevance ranking. Without tf, a scoring model cannot distinguish a document that mentions the term once from one that mentions it twenty times, so matches are effectively unordered on that field. Boolean filtering, existence checks and set operations still work perfectly — which is exactly why filter-only fields are indexed this way on purpose.
saying these in an interview costs you the question
- Thinks a postings list holds only document IDs, always
- Confuses token positions with character offsets
- Says document frequency is repeated in every posting
- Claims phrase queries work without positional data
- Assumes highlighting always requires offsets in the index