skip to content

Search

The engines and libraries that answer "find me the documents matching this text", built on an inverted index rather than a B-tree. Interviewers use this area to check you know when a search engine is the right store and when it is just an expensive copy of your database.

on this pageshow

explore

→ has its own guide

questions

209 · 5 sections

In Elasticsearch, what does a terms aggregation return and what does its size parameter control?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A terms aggregation groups matching documents by the distinct values of a field and returns one bucket per value with a doc_count. size caps how many buckets come back, defaulting to 10, ordered by descending count.

open as a page

How do you use the _analyze API to see the exact tokens Elasticsearch produced for a field?

level: juniorimportance: must knowfreq 66%
basics
~20 s

Call POST /_analyze with an analyzer name and text to test any built-in analyzer, or POST /<index>/_analyze with a field and text to run that field's configured analyzer. Add explain: true to see each pipeline stage's output.

open as a page

How do Elasticsearch's standard, simple, whitespace, and keyword analyzers differ on the same input?

level: juniorimportance: must knowfreq 74%
basics
~20 s

The standard analyzer splits on Unicode word boundaries and lowercases; simple splits at any non-letter and lowercases, so digits disappear; whitespace splits only on spaces and keeps case and punctuation; keyword emits the whole input as one token.

open as a page

How is a custom analyzer composed in Elasticsearch, and in what order do its parts run?

level: juniorimportance: must knowfreq 78%
basics
~20 s

A custom analyzer is zero or more character filters, exactly one tokenizer, then zero or more token filters. Character filters rewrite the raw string, the tokenizer splits it into tokens, and token filters transform that token stream in listed order.

open as a page

When is a text field analyzed in Elasticsearch — at index time, at search time, or both?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Both. Elasticsearch analyzes a text field's value when the document is indexed and stores the resulting terms; it analyzes the query string of a full-text query at search time. A hit requires the two term sets to overlap.

open as a page

In Solr, what is the difference between a core and a collection?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A Solr core is one physical Lucene index living on one node, with its own configuration and data directory. A collection is a SolrCloud-level logical index spread across shards and replicas; every replica is physically a core on some node.

open as a page

In a Solr schema, what is the difference between the solr.StrField and solr.TextField field types?

level: juniorimportance: must knowfreq 72%
basics
~20 s

solr.StrField stores the value verbatim with no analysis, so it only matches, sorts and facets on the whole string. solr.TextField runs an analysis chain that tokenizes and normalizes text, enabling full-text matching on individual words.

open as a page

How do SolrCloud's NRT, TLOG and PULL replica types differ?

level: middleimportance: must knowfreq 60%
basics
~20 s

NRT replicas index documents locally and keep a transaction log, so they support near-real-time search and can become leader. TLOG replicas keep a transaction log but copy the leader's index instead of indexing, and can become leader. PULL replicas only copy the index and can never lead.

open as a page

Why do Solr applications send user queries through the edismax parser instead of the default lucene parser?

level: middleimportance: must knowfreq 74%
basics
~20 s

The default lucene parser expects strict query syntax and errors on stray characters, and it searches one default field. edismax accepts raw user text, searches many weighted fields via qf, and adds relevance controls such as mm, pf, tie and boost.

open as a page

How does SolrCloud's compositeId router decide which shard a document lands on?

level: middleimportance: should knowfreq 50%
basics
~20 s

It hashes the document's uniqueKey into a 32-bit value and sends the document to whichever shard owns that value's hash range. If the key contains a routing prefix such as tenant!docId, the high bits come from the prefix, so every document sharing that prefix lands on the same shard.

open as a page

How does Lucene delete a document from an immutable segment, and when is the disk space actually reclaimed?

level: middleimportance: must knowfreq 58%
basics
~20 s

Lucene never edits a segment, so a delete only marks the document in that segment's live-documents bitset and searches filter the hit out. The terms, stored fields and disk space go away only when a merge rewrites the segment without the dead documents.

open as a page

Why are Lucene index segments immutable, and what does that design buy the search engine?

level: middleimportance: must knowfreq 68%
basics
~20 s

Lucene writes each segment once and never modifies it. Immutability lets readers share segments lock-free, enables aggressive write-once compression and per-segment caching, and makes indexing append-only. The cost is that deletes become tombstones and space returns only when segments merge.

open as a page

What are norms in a Lucene index, and how does BM25Similarity use them at query time?

level: middleimportance: must knowfreq 48%
basics
~20 s

Norms are a per-document, per-field length value written at index time, one byte per document under Lucene's default similarity. BM25Similarity reads that byte to normalize scores by field length, so a term match in a short field outranks the same match in a long one.

open as a page

In Lucene, what is the difference between postings, doc values, and stored fields for the same field?

level: middleimportance: must knowfreq 68%
basics
~20 s

Postings are the inverted index that maps a term to the documents containing it. Doc values are a columnar per-document copy used for sorting, faceting and scripts. Stored fields are the row-wise original bytes returned for the documents you display.

open as a page

How does a Lucene near-real-time reader make new documents searchable without calling IndexWriter.commit()?

level: middleimportance: should knowfreq 45%
basics
~20 s

Opening a reader with DirectoryReader.open(IndexWriter) flushes buffered documents into new segments and lets the reader see them without any fsync or new commit point. The documents become searchable in milliseconds but are not yet crash-durable.

open as a page

What does Elastic Cloud manage for you, and what stays your responsibility?

level: middleimportance: must knowfreq 62%
basics
~10 s

Elastic Cloud handles provisioning, node placement, TLS endpoints, backups, hardware replacement and orchestrated upgrades. Data modelling stays yours: mappings and analysis, shard counts, ILM policies, queries and relevance, role design, and cost control.

open as a page

How does Elastic Cloud handle snapshots and version upgrades, and what still needs your action?

level: middleimportance: should knowfreq 44%
basics
~20 s

Hosted deployments get a managed snapshot repository and an automatic snapshot policy, plus one-click orchestrated rolling upgrades. You still resolve deprecations, reindex indices from two majors back, verify client and extension compatibility, and own long-term backup retention.

open as a page

In an Elastic Cloud deployment, how do availability zones and replicas provide high availability?

level: seniorimportance: should knowfreq 51%
basics
~20 s

Zone count is infrastructure placement; replicas are what actually survive a zone loss. Elastic Cloud spreads instances across the zones you choose and keeps a master quorum possible, but an index with zero replicas still goes red when its zone fails.

open as a page

When would you choose self-managed Elasticsearch over Elastic Cloud, and why?

level: principalimportance: should knowfreq 38%
basics
~20 s

Self-manage when you need control the hosted service withholds — custom native plugins, JVM or kernel tuning, specific hardware, an air-gapped or unsupported region — or when steady large-scale spend clearly exceeds the cost of the operations team you already have.

open as a page

What is the Cloud ID in Elastic Cloud, and how do clients use it to connect?

level: juniorimportance: nice to knowfreq 26%
basics
~20 s

A Cloud ID is a base64-encoded string shown in the Elastic Cloud console that encodes a deployment's Elasticsearch and Kibana endpoints. Official clients, Beats and Logstash accept it plus credentials instead of a full URL.

open as a page

What is an inverted index in a search engine, and how does it differ from a forward index?

level: juniorimportance: must knowfreq 85%
basics
~20 s

An inverted index maps each term to the list of documents containing it, so a query is a dictionary lookup instead of a scan over documents. A forward index maps each document to its terms — the natural storage direction.

open as a page

In search relevance tuning, when should a requirement be a hard filter rather than a boost?

level: juniorimportance: must knowfreq 58%
basics
~20 s

Filter when a non-matching document must never be shown: permissions, region, availability. Boost when the signal is only a preference — recent or in-stock items should rank higher, but a strong match elsewhere may still outrank them.

open as a page

In search relevance evaluation, what do precision@k and recall@k measure, and how does k change each?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Precision@k is the fraction of the top k results that are relevant. Recall@k is the fraction of all relevant documents that appear in the top k. Raising k can only raise recall, and usually lowers precision.

open as a page

What steps turn a raw text field into indexed terms in a search engine's analysis pipeline?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Text analysis runs in three stages: character filtering (strip markup, map characters), tokenization (cut the stream into tokens), then token filtering (lowercase, fold accents, drop stopwords, stem, add synonyms). The surviving tokens become the index terms.

open as a page

What does the k constant in reciprocal rank fusion control?

level: middleimportance: must knowfreq 62%
basics
~20 s

Reciprocal rank fusion gives each document 1/(k + rank) from every list it appears in, summed. The constant k controls how sharply top ranks dominate: small k makes rank one overwhelming, large k flattens the curve so agreement across lists matters more.

open as a page