skip to content

Search & Information Retrieval Concepts

The engine-agnostic ideas every search system rests on: inverted indexes, text analysis, ranking functions like TF-IDF and BM25, relevance tuning, and how you measure whether search is actually any good. Learn these once and Elasticsearch, Solr, and Lucene stop looking like magic.

on this pageshow

questions

page 2 of 2

Should synonym expansion happen at index time or query time in a search system?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Index-time expansion makes queries simple and fast but requires reindexing to change and distorts term statistics. Query-time expansion is editable immediately and keeps statistics honest, at the cost of larger queries and hard multi-word cases. Most teams expand at query time.

open as a page

How do you decide whether hybrid retrieval is worth its cost over keyword search alone?

level: principalimportance: should knowfreq 40%

basics

~20 s

Size the addressable query segment first, not the technology. Hybrid pays when a meaningful share of real traffic is descriptive or vocabulary-mismatched; it does not when traffic is navigational, the corpus is small, or the latency and re-embedding budget is tight.

open as a page

Which parts of an inverted index dominate its size, and what would you stop storing first at scale?

level: principalimportance: should knowfreq 28%

basics

~20 s

Positional data usually dominates, because it scales with token occurrences rather than documents; the term dictionary grows with distinct terms. Cut positions on fields that never take phrase queries, drop unqueried fields entirely, and reconsider n-gram fields before anything else.

open as a page

How would you run relevance tuning so ranking changes stay measurable and reversible?

level: principalimportance: should knowfreq 38%

basics

~20 s

Version the ranking configuration like code, change one signal at a time, gate it offline on a fixed query set, then expose a small traffic slice behind a runtime kill switch. Review results per segment: an average win often hides a tail regression.

open as a page

Why can a learned sparse retriever like SPLADE make inverted-index queries slower than BM25?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Learned sparse models expand a query into tens or hundreds of weighted terms, many of them common with long postings lists, and their flatter weights defeat the dynamic-pruning tricks that make keyword scoring fast. More lists to traverse, far less skipping.

open as a page

Why can BM25's IDF term go negative, and what do implementations do about it?

level: seniorimportance: nice to knowfreq 25%

basics

~20 s

The classic probabilistic IDF is a log-odds ratio that falls below zero once a term appears in more than half the collection, meaning a matching document is penalised for containing a very common word. Implementations add one inside the logarithm or floor the value.

open as a page

showing 31–36 of 36