Search
The engines and libraries that answer "find me the documents matching this text", built on an inverted index rather than a B-tree. Interviewers use this area to check you know when a search engine is the right store and when it is just an expensive copy of your database.
on this pageshowhide
explore
- Elasticsearch (has its own guide)144 questions
- Indices and Mappings30 questions
- Analyzers and Tokenization18 questions
- Query DSL30 questions
- Aggregations24 questions
- Relevance Scoring18 questions
- Sharding and Replication24 questions
- Solr12 questions
- Cores, Collections & SolrCloud6 questions
- Schema & Query Configuration6 questions
- Apache Lucene12 questions
- Segments, Merges & Commits6 questions
- Postings, Doc Values & Codecs6 questions
- Elastic Cloud5 questions
- Search & Information Retrieval Concepts36 questions
- Inverted Index Structure6 questions
- Text Analysis & Tokenization6 questions
- Ranking Models: TF-IDF & BM256 questions
- Relevance Tuning6 questions
- Search Evaluation6 questions
- Vector & Hybrid Search in Search Engines6 questions
→ has its own guide
questions
209 · 5 sectionsIn Elasticsearch, what does a terms aggregation return and what does its size parameter control?
basics
~20 sA terms aggregation groups matching documents by the distinct values of a field and returns one bucket per value with a doc_count. size caps how many buckets come back, defaulting to 10, ordered by descending count.
How do you use the _analyze API to see the exact tokens Elasticsearch produced for a field?
basics
~20 sCall POST /_analyze with an analyzer name and text to test any built-in analyzer, or POST /<index>/_analyze with a field and text to run that field's configured analyzer. Add explain: true to see each pipeline stage's output.
How do Elasticsearch's standard, simple, whitespace, and keyword analyzers differ on the same input?
basics
~20 sThe standard analyzer splits on Unicode word boundaries and lowercases; simple splits at any non-letter and lowercases, so digits disappear; whitespace splits only on spaces and keeps case and punctuation; keyword emits the whole input as one token.
How is a custom analyzer composed in Elasticsearch, and in what order do its parts run?
basics
~20 sA custom analyzer is zero or more character filters, exactly one tokenizer, then zero or more token filters. Character filters rewrite the raw string, the tokenizer splits it into tokens, and token filters transform that token stream in listed order.
When is a text field analyzed in Elasticsearch — at index time, at search time, or both?
basics
~20 sBoth. Elasticsearch analyzes a text field's value when the document is indexed and stores the resulting terms; it analyzes the query string of a full-text query at search time. A hit requires the two term sets to overlap.
In Solr, what is the difference between a core and a collection?
basics
~20 sA Solr core is one physical Lucene index living on one node, with its own configuration and data directory. A collection is a SolrCloud-level logical index spread across shards and replicas; every replica is physically a core on some node.
In a Solr schema, what is the difference between the solr.StrField and solr.TextField field types?
basics
~20 ssolr.StrField stores the value verbatim with no analysis, so it only matches, sorts and facets on the whole string. solr.TextField runs an analysis chain that tokenizes and normalizes text, enabling full-text matching on individual words.
How do SolrCloud's NRT, TLOG and PULL replica types differ?
basics
~20 sNRT replicas index documents locally and keep a transaction log, so they support near-real-time search and can become leader. TLOG replicas keep a transaction log but copy the leader's index instead of indexing, and can become leader. PULL replicas only copy the index and can never lead.
Why do Solr applications send user queries through the edismax parser instead of the default lucene parser?
basics
~20 sThe default lucene parser expects strict query syntax and errors on stray characters, and it searches one default field. edismax accepts raw user text, searches many weighted fields via qf, and adds relevance controls such as mm, pf, tie and boost.
How does SolrCloud's compositeId router decide which shard a document lands on?
basics
~20 sIt hashes the document's uniqueKey into a 32-bit value and sends the document to whichever shard owns that value's hash range. If the key contains a routing prefix such as tenant!docId, the high bits come from the prefix, so every document sharing that prefix lands on the same shard.
How does Lucene delete a document from an immutable segment, and when is the disk space actually reclaimed?
basics
~20 sLucene never edits a segment, so a delete only marks the document in that segment's live-documents bitset and searches filter the hit out. The terms, stored fields and disk space go away only when a merge rewrites the segment without the dead documents.
Why are Lucene index segments immutable, and what does that design buy the search engine?
basics
~20 sLucene writes each segment once and never modifies it. Immutability lets readers share segments lock-free, enables aggressive write-once compression and per-segment caching, and makes indexing append-only. The cost is that deletes become tombstones and space returns only when segments merge.
What are norms in a Lucene index, and how does BM25Similarity use them at query time?
basics
~20 sNorms are a per-document, per-field length value written at index time, one byte per document under Lucene's default similarity. BM25Similarity reads that byte to normalize scores by field length, so a term match in a short field outranks the same match in a long one.
In Lucene, what is the difference between postings, doc values, and stored fields for the same field?
basics
~20 sPostings are the inverted index that maps a term to the documents containing it. Doc values are a columnar per-document copy used for sorting, faceting and scripts. Stored fields are the row-wise original bytes returned for the documents you display.
How does a Lucene near-real-time reader make new documents searchable without calling IndexWriter.commit()?
basics
~20 sOpening a reader with DirectoryReader.open(IndexWriter) flushes buffered documents into new segments and lets the reader see them without any fsync or new commit point. The documents become searchable in milliseconds but are not yet crash-durable.
How does Elastic Cloud handle snapshots and version upgrades, and what still needs your action?
basics
~20 sHosted deployments get a managed snapshot repository and an automatic snapshot policy, plus one-click orchestrated rolling upgrades. You still resolve deprecations, reindex indices from two majors back, verify client and extension compatibility, and own long-term backup retention.
In an Elastic Cloud deployment, how do availability zones and replicas provide high availability?
basics
~20 sZone count is infrastructure placement; replicas are what actually survive a zone loss. Elastic Cloud spreads instances across the zones you choose and keeps a master quorum possible, but an index with zero replicas still goes red when its zone fails.
When would you choose self-managed Elasticsearch over Elastic Cloud, and why?
basics
~20 sSelf-manage when you need control the hosted service withholds — custom native plugins, JVM or kernel tuning, specific hardware, an air-gapped or unsupported region — or when steady large-scale spend clearly exceeds the cost of the operations team you already have.
What is the Cloud ID in Elastic Cloud, and how do clients use it to connect?
basics
~20 sA Cloud ID is a base64-encoded string shown in the Elastic Cloud console that encodes a deployment's Elasticsearch and Kibana endpoints. Official clients, Beats and Logstash accept it plus credentials instead of a full URL.
What is an inverted index in a search engine, and how does it differ from a forward index?
basics
~20 sAn inverted index maps each term to the list of documents containing it, so a query is a dictionary lookup instead of a scan over documents. A forward index maps each document to its terms — the natural storage direction.
In search relevance tuning, when should a requirement be a hard filter rather than a boost?
basics
~20 sFilter when a non-matching document must never be shown: permissions, region, availability. Boost when the signal is only a preference — recent or in-stock items should rank higher, but a strong match elsewhere may still outrank them.
In search relevance evaluation, what do precision@k and recall@k measure, and how does k change each?
basics
~20 sPrecision@k is the fraction of the top k results that are relevant. Recall@k is the fraction of all relevant documents that appear in the top k. Raising k can only raise recall, and usually lowers precision.
What steps turn a raw text field into indexed terms in a search engine's analysis pipeline?
basics
~20 sText analysis runs in three stages: character filtering (strip markup, map characters), tokenization (cut the stream into tokens), then token filtering (lowercase, fold accents, drop stopwords, stem, add synonyms). The surviving tokens become the index terms.
What does the k constant in reciprocal rank fusion control?
basics
~20 sReciprocal rank fusion gives each document 1/(k + rank) from every list it appears in, summed. The constant k controls how sharply top ranks dominate: small k makes rank one overwhelming, large k flattens the curve so agreement across lists matters more.