skip to content

An Elasticsearch analyzer lists filters as ["stop", "lowercase"]; why do capitalized stopwords survive?

level: middleimportance: should knowfreq 50%

answer

  1. Filters run in the order listed
  2. The default list is all lowercase
  3. One filter compares, the other normalizes
  4. Case sensitivity has its own parameter
  5. Sentence-initial words behave differently

basics

~10 s

The stop filter's default English stopword list is lowercase and it does not ignore case, so a token like The does not match and is kept, then lowercased afterwards. Putting lowercase first removes it.

solid answer

~40 s

Token filters form a pipeline, and `stop` runs before `lowercase` here. Its default stopword list is `_english_`, whose entries are lowercase, and `ignore_case` defaults to `false` — so the token `The` fails to match any entry and passes through, only to be lowercased into `the` and indexed. Sentence-initial stopwords therefore survive while mid-sentence ones vanish, which shows up as inconsistent phrase matching. The fix is `["lowercase", "stop"]`, or setting `ignore_case: true` on a named `stop` filter. The same ordering logic governs the rest of a chain: English stemmers expect lowercase input, `asciifolding` should come before a stemmer so folded forms stem consistently, `keyword_marker` must precede `stemmer` to protect terms from it, and synonym rules are parsed using the filters that precede them, so moving a stemmer across a synonym filter changes which rules fire.

code

json · 10 lines
json
// wrong: The survives stop, then becomes the
"bad": {
  "tokenizer": "standard",
  "filter": ["stop", "lowercase"]
},
// right: normalize first, then compare against the list
"good": {
  "tokenizer": "standard",
  "filter": ["lowercase", "stop"]
}

go deeper

for a junior

Remember that token filters run in the order you list them and that lowercase almost always goes first, before anything that compares tokens against a word list.

for a middle

Trace the token stream filter by filter and name the parameter involved — the default stopword list is lowercase and ignore_case is false — rather than describing the symptom.

for a senior

Generalize the rule across the chain: normalize before matching, fold before stemming, protect before stemming, and know how ordering interacts with synonym rule parsing and shingle filler tokens.

for a principal

Treat analysis chains as reviewable, testable assets: a shared chain definition with sample-input tests catches this class of defect once, instead of every team rediscovering it in production relevance.

## The pipeline, not the set The `filter` array in an Elasticsearch analyzer is executed in order. Each filter consumes the token stream the previous one produced. Nothing reorders them for you, and there is no dependency resolution — which means a chain that looks correct as a bag of components can behave badly as a sequence. ## The concrete bug Given `"filter": ["stop", "lowercase"]` and the sentence `The cat sat on the mat`: 1. The `standard` tokenizer emits `The`, `cat`, `sat`, `on`, `the`, `mat`. 2. `stop` uses its default `stopwords` value of `_english_`. That list is written in lowercase and `ignore_case` defaults to `false`, so `The` does not match, while `on` and the lowercase `the` do. 3. Survivors are `The`, `cat`, `sat`, `mat`. 4. `lowercase` turns `The` into `the`, which is now indexed. The index therefore contains `the` for documents where the word happened to start a sentence and not for others. Term frequencies for the most common word in the language become effectively random, phrase queries behave inconsistently, and the stopword list appears to be "half working". ## The two fixes Either reorder to `["lowercase", "stop"]`, which is the conventional chain, or define a named stop filter with `ignore_case: true`. Reordering is usually better: everything downstream — stemmers, synonyms, n-grams — also expects normalized input, so lowercasing early makes the whole chain predictable. ## The general ordering rules **Normalize before you match.** Any filter that compares tokens against a list — `stop`, `synonym_graph`, `keyword_marker`, `stemmer_override` — is only as good as the normalization that ran before it. Put `lowercase` first in almost every chain. **Fold before you stem.** `asciifolding` maps `café` to `cafe`. If it runs after an English stemmer, the stemmer sees the accented form and may treat `café` and `cafe` as different words, producing two distinct stems. Folding first gives one consistent input. **Protect before you stem.** `keyword_marker` flags tokens as keywords so stemmers skip them; a `stemmer` earlier in the list has already done its damage. The same applies to `stemmer_override`, which must precede the stemmer it overrides. **Mind synonyms and stemming together.** Synonym filters parse their rules using the analysis chain that precedes them, so a rule written as `laptop, notebook` is itself analyzed by whatever filters came earlier. Place the synonym filter before aggressive stemming when your rules are written in natural word forms, and be aware that reordering silently changes which rules match. **Stopwords before shingles, usually.** The `shingle` filter inserts a `filler_token` (default `_`) where a token was removed, so a chain that strips stopwords and then shingles produces grams like `cat _ mat`. That may be exactly what you want for phrase-ish matching, or it may be surprising; either way it is a consequence of ordering. **Cheap before expensive.** Filters that remove tokens — `stop`, `length` — shrink the stream for everything downstream. If a heavy filter such as `synonym_graph` or `ngram` is in the chain, removing tokens first reduces its work, provided that does not change semantics. ## How to catch this class of bug Run the chain against a representative sentence and read the output tokens. Comparing what you expected with what came out finds ordering mistakes in seconds — far faster than inferring them from search results, where the symptom is usually "some documents match and some do not" with no obvious pattern. Worth checking specifically: sentence-initial words, accented words, words you expect to be protected from stemming, and any multi-word phrase that a synonym rule is supposed to cover. ## Why interviewers use it The question is a small, verifiable trap. A candidate who has only copied analysis snippets will say the stopword list is broken or that the language is wrong. A candidate who has actually debugged a chain immediately asks what order the filters are in and what `ignore_case` is set to. It separates recipe-following from understanding, without requiring anything exotic.

  • Where does asciifolding belong relative to a stemmer, and why?
    Before it. `asciifolding` maps accented characters to ASCII, so folding first means the stemmer always sees one spelling of a word instead of two. Run it after the stemmer and `café` and `cafe` can stem to different terms, splitting the postings for what users regard as the same word.
  • How do you stop a stemmer from mangling a specific brand name in the chain?
    Put a `keyword_marker` filter before the stemmer with that word in its `keywords` list; it flags the token so stemmers leave it alone. `stemmer_override` is the alternative when you want a specific stem rather than no stemming, and it too must precede the stemmer it overrides.

saying these in an interview costs you the question

  • Says the English stopword list is simply wrong
  • Treats the filter array as order-independent
  • Thinks stop is case-insensitive by default
  • Places keyword_marker after the stemmer
  • Adds a second stop filter instead of reordering

context