What do stopwords cost in a search index, and why do modern engines usually keep them?
answer
- A few word types dominate the token count
- The loss is permanent for indexed documents
- Some band names consist only of these words
- Ranking already down-weights ubiquitous terms
- Skipping at query time beats deleting at index time
basics
~20 sStopword removal shrinks postings for very common words and speeds queries, but it destroys phrases and titles built from them. Modern engines usually keep stopwords because ranking already gives near-zero weight to ubiquitous terms and compression makes them cheap.
solid answer
~50 sStopwords are the highest-frequency function words — *the*, *of*, *and*, *to*. Because word frequency follows a Zipf-like distribution, a small handful of them account for a large share of all tokens, so dropping them at index time historically cut index size and query work substantially. The problem is that removal is **lossy and permanent**: `The Who`, `Take That`, `to be or not to be` and `vitamin a` become empty or unrecognisable, and any phrase query crossing a removed word depends on both sides having dropped it identically. Modern practice usually keeps them because the original justifications weakened: IDF already assigns a term appearing in nearly every document almost no discriminating weight, postings compression and skipping make common lists cheap to store and traverse, and top-k query optimisations can skip low-impact terms dynamically rather than deleting them from the vocabulary forever. Removal still earns its place in narrow cases — huge corpora with hard latency budgets, or domain-specific stopwords that genuinely carry no signal.
code
text · 7 linesquery: "The Who"
with english stopwords removed -> [ ] => no terms, no match
with stopwords kept -> [the] [who] => phrase match possible
document: "king of denmark"
removal keeps a position gap: king@1 ____@2 denmark@3
phrase "king of denmark" matches ... and so does "king in denmark"go deeper
Know what a stopword is, that removal happens during analysis, and that dropping words like the and of can break searches for titles made only of such words.
Explain the size and latency motivation, the phrase and position consequences, and why inverse document frequency already de-emphasises ubiquitous terms without deleting them.
Bring the operational view: removal is irreversible without reindexing, dynamic top-k skipping gives the performance benefit without the loss, and corpus-specific stopwords may matter more than a generic list.
Frame it as a policy decision across locales and fields — where phrase fidelity is a product requirement, what latency budget actually binds, and how a vocabulary change is rolled out and measured rather than argued from anecdotes.
## What a stopword is A stopword is a term deemed to carry so little discriminating information that it is dropped from the token stream during analysis. Standard English lists run to a few dozen or a few hundred function words: articles, prepositions, conjunctions, auxiliary verbs, pronouns. Removal happens as a token filter, after tokenization, and — critically — the removal at index time is *permanent* for those documents. The word is not stored anywhere; it is simply gone. ## Why removal was standard Natural language word frequencies follow a Zipf-like law: a very small number of word types account for a very large fraction of all tokens. That means the postings lists for stopwords are enormous — often close to "every document" — while conveying almost nothing about which document you want. In early systems where disk was scarce and postings were traversed with little skipping machinery, deleting those lists produced a big win in both index size and query latency, with little apparent quality loss. ## What removal breaks The damage falls into three buckets. **Phrases and titles.** Entire meaningful strings consist of stopwords: *The Who*, *Take That*, *Let It Be*, *The The*, *To Be or Not to Be*. After removal these produce zero terms, so the query has nothing to match on and the user gets an empty result or an unrelated one. Short queries like *vitamin a* or *the who band* lose their most discriminating token. **Meaning-bearing function words.** *Not*, *without*, *no* invert meaning, and many stock lists include them. `flights not to Paris` and `flights to Paris` become the same query. **Phrase and proximity queries generally.** Phrase matching relies on token positions. A well-behaved removal filter leaves a position gap where the token was, so `king of denmark` indexed as `king` (pos 1) and `denmark` (pos 3) can still be matched by a query analyzed the same way. But this only works if both sides remove identically — and it also means a document containing `king in denmark` matches the phrase `king of denmark`, a precision loss that surprises people. **Irreversibility.** Because index-time removal deletes information, changing your mind requires a full reindex. There is no query-time knob that brings the words back. ## Why modern engines usually keep them Three things changed. First, **ranking got better at ignoring them by itself**. Inverse document frequency is defined so that a term appearing in nearly every document receives a weight near zero. A term present in 98% of documents contributes almost nothing to a BM25-style score whether or not you deleted it from the vocabulary. The historical fear that stopwords would swamp relevance was really a fear about naïve term-frequency ranking. Second, **storage and traversal got cheap**. Postings lists are delta-encoded and compressed, and skip structures let a conjunctive query jump over long stretches of a common term's list rather than reading it. A high-frequency term costs far less per query than it once did. Third, **dynamic top-k pruning replaced static deletion**. Techniques in the WAND and block-max family let a query computing the top *k* results skip whole blocks of a postings list whose maximum possible contribution cannot change the result set. A term with negligible weight is effectively skipped *at query time*, without ever having been removed at index time — which gives you the performance benefit while keeping the word available for phrase matching. This is the key modern argument, and it is the one interviewers are hoping to hear. ## When removal is still right - **Domain stopwords.** In a corpus of legal filings, words like *court*, *plaintiff* or *hereby* appear everywhere and carry as little signal as *the*. In a set of resumes, *experience* and *responsibilities* qualify. A tuned, corpus-specific list can be worth more than any generic English list, though IDF again handles much of it automatically. - **Very large corpora with hard latency budgets**, where index size or tail latency is the binding constraint and phrase quality on stopword-only queries is not a product requirement. - **Fields where phrases will never be searched** — tag soup, aggregation-only text. ## Middle grounds Rather than an all-or-nothing choice, teams commonly: keep stopwords in a phrase-capable field while removing them in a bag-of-words field used for recall; require a specified fraction of query terms to match so a stopword-heavy query is not over-constrained; or treat high-frequency terms as optional refinements that only influence ranking among documents already matched by the rare terms. Whatever you choose, remember the same list must apply on both sides, and any change on the index side means a reindex.
- Why does IDF alone largely solve the problem stopword lists were invented for?IDF weights a term by how rare it is across the corpus, so a word appearing in nearly every document gets a weight close to zero and contributes almost nothing to the final score. The relevance danger stopword lists guarded against was really a property of naïve term-frequency ranking, not of the words themselves.
- How can a phrase query still work when stopwords are removed at index time?A well-behaved removal filter leaves a position gap where the token was, so "king of denmark" indexes as king at position 1 and denmark at position 3. A query analyzed identically preserves the same gap and matches. The catch is that "king in denmark" now also matches the phrase, and any asymmetry between the two sides breaks it entirely.
- What is a domain stopword and how would you find one?A word that is ubiquitous in your corpus specifically — court in legal filings, patient in clinical notes — so it discriminates nothing even though a general English list would never include it. You find them by ranking terms by document frequency in your own index and inspecting the head of that list, then judging whether users ever search them meaningfully.
Deleting stopwords is like throwing away every common brick before building: cheaper to store, but you can no longer build the walls where those bricks were load-bearing.
saying these in an interview costs you the question
- Believes stopword removal is always a best practice
- Cannot name a query that stopword removal destroys
- Thinks removed words can be restored without reindexing
- Assumes common words would otherwise dominate relevance scores
- Applies an English stopword list to non-English fields