Search & Information Retrieval Concepts
The engine-agnostic ideas every search system rests on: inverted indexes, text analysis, ranking functions like TF-IDF and BM25, relevance tuning, and how you measure whether search is actually any good. Learn these once and Elasticsearch, Solr, and Lucene stop looking like magic.
on this pageshowhide
explore
- Inverted Index Structure6 questions
- Text Analysis & Tokenization6 questions
- Ranking Models: TF-IDF & BM256 questions
- Relevance Tuning6 questions
- Search Evaluation6 questions
- Vector & Hybrid Search in Search Engines6 questions
- Backend Developerroleanchors this topic
- Data Engineerroleanchors this topic
- Elasticsearchskillanchors this topic
- Full Stack Developerroleanchors this topic
- Java Backend Developerroleanchors this topic
- Kotlin Backend Developerroleanchors this topic
- Software Architectroleanchors this topic
- Forward Deployed Engineerrole
- Server-Side Game Developerrole
questions
page 2 of 2Should synonym expansion happen at index time or query time in a search system?
basics
~20 sIndex-time expansion makes queries simple and fast but requires reindexing to change and distorts term statistics. Query-time expansion is editable immediately and keeps statistics honest, at the cost of larger queries and hard multi-word cases. Most teams expand at query time.
How do you decide whether hybrid retrieval is worth its cost over keyword search alone?
basics
~20 sSize the addressable query segment first, not the technology. Hybrid pays when a meaningful share of real traffic is descriptive or vocabulary-mismatched; it does not when traffic is navigational, the corpus is small, or the latency and re-embedding budget is tight.
Which parts of an inverted index dominate its size, and what would you stop storing first at scale?
basics
~20 sPositional data usually dominates, because it scales with token occurrences rather than documents; the term dictionary grows with distinct terms. Cut positions on fields that never take phrase queries, drop unqueried fields entirely, and reconsider n-gram fields before anything else.
How would you run relevance tuning so ranking changes stay measurable and reversible?
basics
~20 sVersion the ranking configuration like code, change one signal at a time, gate it offline on a fixed query set, then expose a small traffic slice behind a runtime kill switch. Review results per segment: an average win often hides a tail regression.
Why can a learned sparse retriever like SPLADE make inverted-index queries slower than BM25?
basics
~20 sLearned sparse models expand a query into tens or hundreds of weighted terms, many of them common with long postings lists, and their flatter weights defeat the dynamic-pruning tricks that make keyword scoring fast. More lists to traverse, far less skipping.
Why can BM25's IDF term go negative, and what do implementations do about it?
basics
~20 sThe classic probabilistic IDF is a log-odds ratio that falls below zero once a term appears in more than half the collection, meaning a matching document is penalised for containing a very common word. Implementations add one inside the logarithm or floor the value.
showing 31–36 of 36