What does Elasticsearch's english analyzer do to text that the standard analyzer does not?
answer
- Language-aware, not language-neutral
- Something disappears before stemming even starts
- Apostrophes get special treatment
- Recall goes up, precision can go down
- A parameter exists to protect brand names
basics
~10 sThe english analyzer adds English-specific processing on top of standard tokenization: it strips possessives, removes English stopwords, and stems words to a common root, so "running" and "runs" both index as "run".
solid answer
~40 sBoth start from the same tokenization, but the `english` analyzer then applies a language-specific chain: a possessive stemmer that turns `Foxes'` into `Foxes`, lowercasing, removal of English stopwords (the `_english_` list, configurable via `stopwords`/`stopwords_path`), protection of terms listed in `stem_exclusion`, and finally an English stemmer. The effect is recall: a query for `run` finds documents containing `running` and `runs`, because both sides are stemmed to the same term. The costs are real too — stems are not dictionary words (`shoes` becomes `shoe`, some distinct words collapse together), aggressive stemming can conflate meanings and hurt precision, and stopword removal makes phrases built entirely from stopwords unmatchable. It is also monolingual: applying it to mixed-language content damages the other languages.
code
json · 6 linesPOST /_analyze
{
"analyzer": "english",
"text": "The Foxes' running shoes"
}
// tokens: fox, run, shoego deeper
Recall that the english analyzer reduces words to a stem and drops common English words, so a search for one inflection finds the others. Being able to predict the tokens for a short sentence is enough at this level.
Describe the stages in order — possessive stripping, lowercasing, stopword removal, stemming — and explain the recall/precision trade-off, including why stems are not real words and why irregular verbs are missed.
Show that you have handled the fallout: over-stemmed brand names fixed with stem_exclusion, an unstemmed companion field so exact matches outrank stemmed ones, and a plan for multilingual content that a single language analyzer would damage.
Own the policy question: whether stemming and stopword removal are defaults across the platform, how relevance regressions from analysis changes are measured before rollout, and what the reindex cost of revisiting that decision across the estate would be.
## The two analyzers side by side The `standard` analyzer is language-neutral: Unicode word segmentation plus lowercasing. It preserves every word it finds, in a normalized surface form. The `english` analyzer is one of the per-language analyzers Elasticsearch ships, and it layers English knowledge on top of that tokenization. For the text `The Foxes' running shoes`: - **standard** → `the`, `foxes`, `running`, `shoes` - **english** → `fox`, `run`, `shoe` Three differences are visible: the stopword `the` is gone, the possessive apostrophe-s handling has removed the trailing `'`, and every remaining word has been reduced to a stem. ## What the chain contains The language analyzers are defined as pre-assembled chains. For `english` the stages are, in order: standard tokenization; an English possessive stemmer that strips trailing `'s`; lowercasing; a keyword-marker step that protects any terms you list in `stem_exclusion` from being stemmed; removal of stopwords from the `_english_` list; and an English stemmer. Two parameters make it configurable without building your own chain: - **`stopwords`** (or `stopwords_path`) replaces the default list. Setting it to `_none_` keeps stopwords in the index — useful when phrase queries over common words matter. - **`stem_exclusion`** lists terms that must survive stemming intact, which is how you protect brand and product names from being mangled. ## Why stemming helps Search users type one inflection and expect all of them. Without stemming, a query for `optimize` misses `optimizing`, `optimized` and `optimization`. Stemming solves that by mapping every inflection of a word to a single index term — crucially at *both* index time and query time, so the two sides meet on the same stem. This is a recall mechanism: it increases the number of documents a query can reach. ## Why stemming hurts Algorithmic stemmers are rule-based suffix strippers, not dictionaries. Three consequences follow. First, **stems are not words**. You cannot show them to users, and you cannot assume a stem is the singular form. Second, **over-stemming conflates meanings**. Two words with different senses can reduce to the same stem, so a search for one starts returning the other. Precision drops in a way that looks, from the outside, like the relevance model got worse. Third, **under-stemming misses relations** the stemmer's rules do not cover. Irregular forms are the classic case: a suffix-stripping stemmer has no way to connect `ran` to `run`, or `better` to `good` — that would require lemmatization with a dictionary, which is a different (and more expensive) technique. Stopword removal carries its own trap. Dropping `the`, `to`, `be` shrinks the index and speeds up scoring, but it makes a phrase composed of stopwords unmatchable, and it destroys meaning in cases where the small word is the point (`the who`, `to be or not to be`, `vitamin a`). Modern hardware makes the storage argument for stopword removal much weaker than it was, so many teams keep stopwords and let the ranking model down-weight them. ## Symmetry matters A stemmed field only works if the query is stemmed the same way. If documents are indexed with `english` and searched with `standard`, a query for `running` produces the term `running`, which does not exist in the index — the field silently returns nothing for inflected queries. Elasticsearch avoids this by default: the field's analyzer is used for both sides unless you deliberately configure otherwise. ## Practical patterns The usual production shape is not "stemmed or not" but *both*. Index the same content twice — once with a language analyzer for recall, once with a minimally-processed analyzer for precision — and query them together so exact matches outrank stemmed ones. That gives users the forgiving behaviour they expect without letting an aggressive stemmer decide the top result. For multilingual corpora, a single language analyzer is the wrong tool: it will stem the wrong language's words with English rules. The realistic options are detecting language at ingest and routing to per-language fields, or falling back to language-neutral analysis. ## Verifying, not guessing Every claim above is checkable in seconds with the `_analyze` API: run the same text through `standard` and `english` and read the token lists. When a relevance complaint arrives — "searching for our product name returns competitors" — analyzing the product name through the field's analyzer is usually where the answer is, and `stem_exclusion` is usually the fix.
- How do you stop the english analyzer from stemming specific terms such as brand names?Configure a copy of the analyzer with `stem_exclusion`, listing the terms to protect; they are marked as keywords before the stemmer runs and pass through unchanged. The alternative, often used alongside it, is to index the field a second time with minimal analysis and query both, so the unstemmed form can win on exact matches.
- How would you keep stopwords in an english-analyzed field?Define an analyzer of type `english` with `stopwords` set to `_none_`, which disables removal while keeping possessive handling and stemming. You do this when phrase queries over common words matter — song and film titles, quotations, or any corpus where "the" carries meaning. The index grows slightly and scoring must down-weight the common terms instead.
- Why can't a suffix-stripping stemmer connect "ran" to "run"?Algorithmic stemmers remove suffixes by rule; irregular English verbs change the stem itself, so no suffix rule reaches them. Connecting them requires lemmatization — a dictionary lookup with part-of-speech context — which is slower and not what the built-in english analyzer does. In practice teams cover the important irregulars with synonyms instead.
saying these in an interview costs you the question
- Assumes stemming always produces valid dictionary words
- Thinks the english analyzer also expands synonyms
- Treats stopword removal as unconditionally beneficial
- Says language analyzers only change tokenization
- Believes stemming can be toggled per query at search time