How do Elasticsearch's standard, simple, whitespace, and keyword analyzers differ on the same input?
answer
- Four different pairs of scissors
- One of them never cuts at all
- One preserves the original casing
- One quietly throws away every digit
- Unicode word boundaries are the default
basics
~20 sThe standard analyzer splits on Unicode word boundaries and lowercases; simple splits at any non-letter and lowercases, so digits disappear; whitespace splits only on spaces and keeps case and punctuation; keyword emits the whole input as one token.
solid answer
~40 sAll four are pre-assembled analysis pipelines, and they differ in how aggressively they cut the input. **standard** uses the standard tokenizer (Unicode text segmentation), then lowercases — punctuation is dropped, digits are kept. **simple** uses the lowercase tokenizer: it breaks at every non-letter, so `iphone15` becomes `iphone` and the digits vanish entirely. **whitespace** only splits on whitespace and does nothing else, so `JUMPED!` stays `JUMPED!` — case-sensitive and punctuation-bearing. **keyword** is a no-op: the entire input becomes a single token. Elasticsearch also ships `stop` (like simple, plus English stopword removal), `pattern` (regex-based), and the per-language analyzers. In practice you index prose with `standard` or a language analyzer, and exact identifiers go into a `keyword` field rather than through the keyword analyzer.
code
json · 6 linesPOST /_analyze
{
"analyzer": "standard",
"text": "The Quick-Brown Foxes JUMPED! 2 times"
}
// tokens: the, quick, brown, foxes, jumped, 2, timesgo deeper
Be ready to name the four analyzers and say in one sentence what each does to a sample string. Knowing that standard lowercases and drops punctuation, and that keyword does nothing, covers most screening questions.
Explain the mechanics: which tokenizer each built-in wraps, why the simple analyzer loses digits, and why the standard analyzer's stopword list is empty by default. Expect to be asked to predict tokens for a given input.
Show judgment about consequences in production: picking whitespace makes a field case-sensitive, picking simple destroys numeric tokens in a catalogue, and any change requires reindexing before old documents behave the new way.
Own the standardization argument: which analyzer families the organization defaults to per content type, how those choices are shipped through index templates, and what the migration cost is when a default turns out wrong across hundreds of indices.
## What an analyzer is An analyzer is the pipeline Elasticsearch runs over the value of a `text` field before anything is written to the inverted index, and over the query string at search time. It has three stages: zero or more **character filters** (rewrite the raw string), exactly one **tokenizer** (cut the string into tokens), and zero or more **token filters** (modify, add, or drop tokens). A *built-in* analyzer is simply a named, pre-assembled combination of those stages, so you can write `"analyzer": "standard"` in a mapping instead of wiring the pieces yourself. Only `text`-family fields are analyzed. A `keyword` field is indexed verbatim; it accepts a `normalizer`, not an analyzer. ## The built-ins, one at a time **standard** — the default. Tokenizer: the standard tokenizer, which implements Unicode text segmentation (UAX #29 word boundaries). That means punctuation and symbols are treated as separators and thrown away, while letters and digits survive. Then a `lowercase` token filter runs. It also contains a stop filter, but it is configured with an **empty** stopword list by default, so `the` is indexed unless you set the `stopwords` (or `stopwords_path`) parameter. It also exposes `max_token_length`. **simple** — tokenizer only: the `lowercase` tokenizer, which splits at every character that is not a letter and lowercases as it goes. The consequence people forget is that **digits are separators and are discarded**: `iphone 15` yields just `iphone`. That makes it a poor fit for product catalogues and version strings. **whitespace** — the `whitespace` tokenizer and nothing else. Tokens are exactly what was between spaces, with original case and any attached punctuation. `JUMPED!` and `jumped` become two different terms, so matching is case-sensitive. **keyword** — a no-op analyzer: the whole input becomes one token, unchanged. It exists so you can keep a field in the `text` family while treating its value atomically; for genuinely exact values, mapping the field as the `keyword` **type** is the normal choice. Do not confuse the two — the keyword *analyzer* and the keyword *field type* are different things that happen to share a name. **stop** — like `simple` (lowercase tokenizer) plus a stop filter that removes English stopwords by default; the list is configurable. **pattern** — splits with a regular expression, defaulting to a non-word-character pattern, and lowercases by default. Useful for structured strings such as camelCase or delimiter-separated codes. **Language analyzers** (`english`, `french`, `german`, …) — tokenization plus language-specific stopwords and stemming. ## A worked example Input: `The Quick-Brown Foxes JUMPED! 2 times` - **standard** → `the`, `quick`, `brown`, `foxes`, `jumped`, `2`, `times` - **simple** → `the`, `quick`, `brown`, `foxes`, `jumped`, `times` (the `2` is gone) - **whitespace** → `The`, `Quick-Brown`, `Foxes`, `JUMPED!`, `2`, `times` - **keyword** → `The Quick-Brown Foxes JUMPED! 2 times` (one token) - **stop** → `quick`, `brown`, `foxes`, `jumped`, `times` You never have to memorize this: the `_analyze` API prints the tokens for any analyzer and any input, and checking it is the first move when a search misbehaves. ## How to choose For human-readable prose you want recall, so `standard` is the sane default and a language analyzer is the upgrade when the corpus is single-language. For identifiers, SKUs, enum values, hostnames and status codes you want exactness, so use a `keyword` field — optionally with a lowercase normalizer if you need case-insensitive filtering. `whitespace` is the tool when the punctuation itself is meaningful (`C++`, `AT&T`, log tokens) and you are willing to handle case yourself. `simple` and `stop` are rarely the right production choice on their own; they mostly show up as teaching examples or as the base for something custom. ## Common traps The analyzer applies to both indexing and searching, so changing it changes what queries mean, not only what is stored — and existing documents keep their old tokens until they are reindexed. Choosing `whitespace` silently makes the field case-sensitive, which usually surfaces as "search works for lowercase queries only". Choosing `simple` silently destroys every number in the corpus. And assuming `standard` removes English stopwords is a very common wrong answer in interviews: its stopword list is empty until you configure one.
- What does the stop analyzer add compared with the simple analyzer?Both use the lowercase tokenizer, so both split at non-letters and lowercase. The stop analyzer then removes stopwords, using an English list by default and accepting a custom `stopwords` array or `stopwords_path`. That raises precision on common words but makes phrases such as "to be or not to be" unmatchable, because nothing is left to index.
- When would you pick the pattern analyzer over the standard analyzer?When the field has machine-made structure rather than prose — camelCase identifiers, pipe- or colon-delimited codes, log keys. The pattern analyzer tokenizes with a regular expression (defaulting to splitting on non-word characters) and lowercases by default, so you can express exactly where the boundaries are instead of accepting Unicode word segmentation.
- Does the keyword analyzer do the same thing as mapping a field as keyword?They produce a single indexed term either way, but they are different mechanisms. The keyword *analyzer* is a no-op analyzer usable on a `text` field; the keyword *type* is a field type that skips analysis, supports doc_values for sorting and aggregations, and accepts a `normalizer`. For exact-match values the field type is the normal choice.
Think of the four as scissors with different blade widths: standard trims at every word boundary, simple hacks at anything that is not a letter, whitespace cuts only where you already left a gap, and keyword leaves the sheet of paper whole.
saying these in an interview costs you the question
- Says the whitespace analyzer lowercases its tokens
- Confuses the keyword analyzer with the keyword field type
- Claims the simple analyzer keeps numbers in the text
- Believes the standard analyzer removes English stopwords by default
- Thinks an analyzer also applies to keyword-typed fields