What is the difference between stemming and lemmatization in search text analysis?
answer
- One uses rules, the other uses a word list
- Output of one need not be a real word
- Irregular plurals separate them sharply
- Over-merging versus failing to merge
- Cost per token decides most production choices
basics
~20 sStemming strips affixes with language rules, fast and dictionary-free, producing stems that may not be real words. Lemmatization uses a dictionary and word context to return the true base form. Search engines usually choose stemming for speed and recall.
solid answer
~50 sBoth reduce inflected words so that *running*, *runs* and *ran* can meet, but they work differently. **Stemming** applies ordered suffix-rewriting rules — the Porter and Snowball families are the standard ones — with no dictionary and no grammar. It is fast and language-specific but crude: it happily produces non-words (*studies* to *studi*), over-stems (*universe* and *university* colliding on *univers*), and under-stems irregulars (*ran* stays *ran*). **Lemmatization** looks the word up in a morphological dictionary, often using part-of-speech context, and returns the real dictionary form: *better* to *good*, *mice* to *mouse*, *ran* to *run*. It is more precise and much more expensive, needs curated language resources, and is ambiguous without context. Search engines default to stemming because it costs almost nothing per token and recall usually matters more than the occasional false match. Whichever you pick, it must run on both index and query text.
code
text · 8 linesword aggressive stemmer lemmatizer
---------------------------------------------
running run run
ran ran (missed) run
mice mice (missed) mouse
studies studi study
university univers (collides) university
universe univers (collides) universego deeper
Recall that both collapse word variants so a search for one form finds the others, that stemming is rule-based and may output non-words, and that lemmatization returns the real base form.
Explain the mechanisms, name over- and under-stemming with examples, and articulate why production systems usually pick stemming. Know that the choice is index-time and needs matching query-side treatment.
Show you can tune the aggressiveness against a corpus, protect brand and identifier tokens, and use a dual exact-plus-stemmed field with boosting when a single setting cannot satisfy both precision and recall.
Own the language strategy: which locales get which normalisation, what the maintenance burden of lexicons is, how a stemmer change is rolled out across a populated corpus, and how you measure whether it helped rather than argue about examples.
## The problem both solve An inverted index matches terms exactly, so *running*, *runs*, *ran* and *runner* are four unrelated terms. A user searching one of them expects documents containing the others. Reducing morphological variants to a shared form is the fix, and there are two families of technique. ## Stemming: rules, no dictionary A stemmer is a small program of ordered rewrite rules over word endings. The classical English stemmer (Porter, and its maintained successor family Snowball) runs steps such as "replace a final *sses* with *ss*", "drop a final *ing* if what remains contains a vowel", each guarded by conditions on syllable structure. Nothing consults a word list, and the output need not be a word — the only requirement is that variants of the same word converge on the same string. `argue`, `argued`, `argues` and `arguing` all reduce to `argu`; that string is meaningless to a human and perfectly serviceable as an index term. Properties: microseconds per token, tiny memory, deterministic, easy to ship per language. Stemmers exist for most European languages plus several others, each hand-written against that language's morphology — an English stemmer applied to Spanish text produces garbage. Stemmers come in aggressiveness grades. Light or *minimal* stemmers only handle plurals and the most regular inflections; aggressive ones also strip derivational suffixes (*-ation*, *-ness*, *-ment*) and collapse far more words together. ## Lemmatization: dictionary and morphology A lemmatizer maps a surface form to its **lemma**, the canonical dictionary headword. It needs a morphological lexicon of the language, and typically the word's part of speech, because the mapping is genuinely ambiguous: *saw* is the past tense of *see* (lemma *see*) or a noun for a cutting tool (lemma *saw*); *left* is a verb form or a direction. Because it consults real word knowledge, it handles irregulars a rule-based stemmer cannot: *mice* to *mouse*, *went* to *go*, *better* to *good*. The costs: you need a maintained lexicon per language, the lookup is far slower than rule application, and getting the part of speech right means running at least a lightweight tagger — which itself needs sentence context that a short query string may not have. Dictionary-driven approaches based on spellchecker lexicons sit between the two extremes, offering real word forms without full grammatical analysis. ## Two error modes **Over-stemming** merges words that should stay apart: *universe*, *university* and *universal* colliding, or *organization* and *organ*. It inflates recall and injects irrelevant results, and it also corrupts term statistics, because the merged term now looks far more common than either original word. **Under-stemming** fails to merge words that belong together: *ran* versus *run*, *mouse* versus *mice*, *children* versus *child*. Recall silently suffers and users never see the documents they wanted. Aggressive stemmers over-stem; light stemmers and lemmatizers under-merge less but leave more variants distinct. The choice is a straightforward precision/recall dial, and the right setting depends on the corpus. A legal or medical corpus where terms of art must not blur tolerates far less aggression than a general product catalogue. ## Which to choose In practice most production search uses stemming, because the throughput cost of lemmatization on ingest and on every query is real, the language coverage is thinner, and search users generally prefer a couple of extra results over missing the one they wanted. Lemmatization earns its keep where linguistic precision drives the product: question answering, legal and biomedical retrieval, highly inflected languages such as Finnish, Turkish or Russian where rule-based stemming is weak and morphology carries much of the meaning. A hybrid pattern is common and worth naming in an interview: index the field **twice**, once with an exact/lightly-normalised analysis and once stemmed, then query both and weight the exact field higher. Users who type the precise form get precision-ranked results at the top, while the stemmed field supplies recall underneath. This costs index size but sidesteps the all-or-nothing choice. ## Operational notes Two details bite teams repeatedly. First, **symmetry**: the same reduction must run on query text, or a stemmed index becomes unreachable. Second, **ordering and protection**: stemming interacts with synonyms (a synonym list written in surface forms will not fire after stemming) and with brand names, product codes and proper nouns you never want reduced — most stacks provide a protected-words or keyword-marker mechanism to exempt those tokens from the stemmer. Finally, changing the stemmer changes indexed terms, so it is an index-time change requiring a reindex, not a knob you flip on a live index.
- Give a concrete example of over-stemming and explain the damage it does.An aggressive English stemmer reduces *universe*, *university* and *universal* to *univers*, so a search for astronomy content returns campus pages. Beyond the irrelevant hits, the merged term's document frequency is now the sum of all three words, so its IDF drops and every query using it is ranked as if the word were more common than it is.
- When is lemmatization worth its cost over stemming?When linguistic precision drives the product or the language resists rule-based stemming: biomedical and legal retrieval where terms of art must not blur, question answering, and highly inflected languages such as Finnish, Turkish or Russian where surface forms diverge far from stems. It is also easier to justify offline, on ingest, than on every query.
- How do you keep a stemmer from mangling brand names and product codes?Exempt them. Analysis chains provide a protected-words or keyword-marker filter that flags listed tokens so downstream stemmers skip them, and identifier-shaped fields should not be stemmed at all. The complementary tactic is indexing an unstemmed copy of the field and boosting it, so an exact brand match outranks a stemmed one.
A stemmer is a pair of scissors that cuts off word endings by rule; a lemmatizer is a dictionary that looks the word up and tells you its headword.
saying these in an interview costs you the question
- Says a stemmer always returns a valid dictionary word
- Treats the two as interchangeable with no cost difference
- Thinks one stemmer works across all languages
- Ignores that stemming must also run on the query
- Assumes more aggressive stemming is always better recall-wise