skip to content

Which parts of an inverted index dominate its size, and what would you stop storing first at scale?

level: principalimportance: should knowfreq 28%

answer

  1. Different components grow by different laws
  2. One entry per document versus one per token
  3. Vocabulary size is not corpus size
  4. Highlighting only runs on one page
  5. Every cut here means a reindex

basics

~20 s

Positional data usually dominates, because it scales with token occurrences rather than documents; the term dictionary grows with distinct terms. Cut positions on fields that never take phrase queries, drop unqueried fields entirely, and reconsider n-gram fields before anything else.

solid answer

~50 s

Break the index into components and size each one. **Postings** scale with matching documents per term; **positions** scale with token *occurrences* and are normally the largest single component when enabled; **offsets** add two more numbers per occurrence; the **term dictionary** scales with distinct terms, so identifiers, hashes, n-grams and misspellings inflate it; and the forward structures — per-document columnar values for sorting and faceting, term vectors, stored content — often rival the inverted side. Measure the per-field breakdown before cutting anything; intuition about which field is heavy is usually wrong. Then cut in order of least capability lost: fields that are indexed but never queried, positions on long-body fields that only take bag-of-words queries, offsets replaced by query-time re-analysis for highlighting (which runs only on one page of results), and n-gram fields swapped for a cheaper prefix strategy. Each cut is a reindex, so decide it as a schema policy rather than a tuning knob.

go deeper

for a junior

Know that an index stores more than document lists — frequencies, positions, offsets and forward structures — and that each of those costs storage you can choose not to pay.

for a middle

Explain the growth law of each component, especially that positions scale with token occurrences while postings scale with matching documents, and what capability disappears when each is dropped.

for a senior

Demonstrate the diagnostic path: measure per-field component sizes, cross-reference against query logs, and stage the cuts by capability lost, knowing each one requires a reindex and a cutover plan.

for a principal

Own it as governance — a schema standard where declared query patterns determine indexing options, continuous measurement of unused capability, and a clear position on when to shrink versus when to shard.

## Decompose before you cut "The index is too big" is not actionable. An inverted index is several structures with different growth laws, and until you know the per-field, per-component breakdown you will cut the wrong thing. The components and what each scales with: - **Term dictionary** — number of *distinct* terms. Independent of how many documents contain each one. - **Postings (document IDs and frequencies)** — total (term, document) pairs; roughly documents times distinct terms per document. - **Positions** — total token *occurrences*; roughly one entry per indexed token in the corpus. - **Offsets** — two more values per occurrence, when stored in the index. - **Skip / block metadata** — a small fraction of postings. - **Forward structures** — per-document columnar values for sorting, faceting and aggregation; term vectors; the stored original content used to render results. ## Why positions usually dominate A posting contributes one document ID per matching document, but *tf* positions. Over a corpus this means the positional stream carries approximately one entry per indexed token, while the document-ID stream carries one entry per (term, document) pair. For long documents with repeated vocabulary, the ratio is severe. Positions also compress less well than document gaps in dense lists, since within-document positions are less predictable. The practical implication: the single largest lever on a text-heavy index is usually whether long body fields carry positions at all. ## Why the term dictionary can surprise you Dictionary growth is driven by vocabulary, not volume. Natural language vocabulary grows sublinearly with corpus size, so an honest prose corpus has a manageable dictionary. What breaks it is anything that manufactures distinct terms: n-gram and edge-n-gram fields (which multiply terms per token), identifiers, hashes, URLs, log lines with embedded numbers, OCR noise and misspellings, and indexing the same content once per language analyser. A large dictionary hurts more than storage: it slows merges, enlarges the in-memory term index, and makes wildcard and fuzzy enumeration expensive. ## Order the cuts by capability lost Rank candidate cuts by what the product actually stops being able to do: 1. **Fields indexed but never queried.** Pure waste. Find them from query logs, not from opinion. This is the only free cut, and there is usually more of it than anyone expects. 2. **Positions on fields that never take phrase or proximity queries.** Large saving, and the capability lost is precise and testable. If a body field is only ever queried as a bag of words with a relevance model, positions buy nothing there — but confirm no proximity-based relevance feature depends on them first. 3. **Offsets in the index, replaced by query-time re-analysis or per-document term vectors for highlighting.** Highlighting runs on one page of results, so recomputing it costs a small amount of CPU on ten documents instead of storage on all of them. This trade is favourable far more often than it is taken. 4. **N-gram fields.** They multiply both dictionary and postings. Substitutes exist for most of what they are used for: prefix-oriented structures for autocomplete, a reversed field for leading wildcards, dedicated suggester structures. 5. **Term frequencies on genuine filter fields.** Status codes, tags, tenant IDs and enumerations never contribute textual relevance. Indexing them as membership only is small but free. 6. **Forward structures nobody sorts or facets on.** Columnar per-document values are silently enabled far more widely than they are used. ## The cuts that are not really cuts Beware the ones that move cost rather than remove it. Aggressive stopword removal shrinks the biggest postings lists but damages phrase matching and can make some queries unanswerable. Shorter retention shrinks the index but is a product decision about what users can find. Compression changes cost the same information at different CPU. Sharding does not shrink an index at all; it distributes it, which helps memory pressure per node and hurts per-query fan-out. ## Make it a policy, not an incident Every one of these decisions is fixed at index time and reversible only by reindexing the corpus. At scale that is a multi-hour to multi-day operation with an aliasing or dual-write strategy behind it. So the leadership move is to set a schema standard — new fields declare their query patterns, and the indexing options follow from the declared patterns — rather than to run a size-reduction project every year. Pair it with continuous measurement: per-field component sizes, query-log coverage of each field, and an alert when a field is indexed with capabilities no query has used in a quarter. ## Framing the answer A strong answer names the growth law of each component, identifies positions as the usual dominator with a reason, ranks the cuts by capability lost rather than by bytes saved, and closes on the governance point: because these choices are baked in at index time, they belong in schema review.

  • How would you actually measure which field is responsible for the size?
    Get a per-field, per-component breakdown from the engine's index statistics rather than reasoning from document counts — most engines can report bytes by field and by structure. Cross-reference it with query logs showing which fields are searched, phrased, highlighted, sorted or faceted. The intersection of large and unused is where the free wins are.
  • Why is dropping positions a decision that must be made before indexing rather than tuned later?
    Because positional data is written during indexing and cannot be added to existing index units. Turning it off shrinks only what is written afterwards, and turning it back on requires a full reindex of the corpus. At scale that means building a parallel index and cutting over behind an alias, so the choice belongs in schema design with a documented rationale.
  • What is wrong with reaching for aggressive stopword removal as a size lever?
    It shrinks the largest postings lists, which looks attractive, but it destroys queries that depend on those words — phrases and titles where the common word is meaningful, and any query consisting mostly of stopwords. Modern ranking already discounts common terms through IDF, and dynamic pruning already limits how much of a long list is read, so the size win rarely justifies the recall loss.
  • When is sharding the wrong answer to an index that is too large?
    Sharding redistributes bytes; it does not remove them. It helps when per-node memory or merge throughput is the constraint. It hurts when the problem is total cost or query latency, because every query fans out to more shards, term statistics are computed per shard, and the coordination overhead grows. Shrink what you store first, then decide how to spread it.

saying these in an interview costs you the question

  • Assumes the term dictionary is the largest component
  • Treats sharding as a way to reduce index size
  • Cuts bytes without asking which queries stop working
  • Believes positions can be disabled without reindexing
  • Reaches for stopword removal as the first size lever

context