skip to content

How do Elasticsearch's standard, simple, whitespace, and keyword analyzers differ on the same input?

level: juniorimportance: must knowfreq 74%

answer

  1. Four different pairs of scissors
  2. One of them never cuts at all
  3. One preserves the original casing
  4. One quietly throws away every digit
  5. Unicode word boundaries are the default

basics

~20 s

The standard analyzer splits on Unicode word boundaries and lowercases; simple splits at any non-letter and lowercases, so digits disappear; whitespace splits only on spaces and keeps case and punctuation; keyword emits the whole input as one token.

solid answer

~40 s

All four are pre-assembled analysis pipelines, and they differ in how aggressively they cut the input. **standard** uses the standard tokenizer (Unicode text segmentation), then lowercases — punctuation is dropped, digits are kept. **simple** uses the lowercase tokenizer: it breaks at every non-letter, so `iphone15` becomes `iphone` and the digits vanish entirely. **whitespace** only splits on whitespace and does nothing else, so `JUMPED!` stays `JUMPED!` — case-sensitive and punctuation-bearing. **keyword** is a no-op: the entire input becomes a single token. Elasticsearch also ships `stop` (like simple, plus English stopword removal), `pattern` (regex-based), and the per-language analyzers. In practice you index prose with `standard` or a language analyzer, and exact identifiers go into a `keyword` field rather than through the keyword analyzer.

code

json · 6 lines
json
POST /_analyze
{
  "analyzer": "standard",
  "text": "The Quick-Brown Foxes JUMPED! 2 times"
}
// tokens: the, quick, brown, foxes, jumped, 2, times

go deeper

for a junior

Be ready to name the four analyzers and say in one sentence what each does to a sample string. Knowing that standard lowercases and drops punctuation, and that keyword does nothing, covers most screening questions.

for a middle

Explain the mechanics: which tokenizer each built-in wraps, why the simple analyzer loses digits, and why the standard analyzer's stopword list is empty by default. Expect to be asked to predict tokens for a given input.

for a senior

Show judgment about consequences in production: picking whitespace makes a field case-sensitive, picking simple destroys numeric tokens in a catalogue, and any change requires reindexing before old documents behave the new way.

for a principal

Own the standardization argument: which analyzer families the organization defaults to per content type, how those choices are shipped through index templates, and what the migration cost is when a default turns out wrong across hundreds of indices.

## What an analyzer is An analyzer is the pipeline Elasticsearch runs over the value of a `text` field before anything is written to the inverted index, and over the query string at search time. It has three stages: zero or more **character filters** (rewrite the raw string), exactly one **tokenizer** (cut the string into tokens), and zero or more **token filters** (modify, add, or drop tokens). A *built-in* analyzer is simply a named, pre-assembled combination of those stages, so you can write `"analyzer": "standard"` in a mapping instead of wiring the pieces yourself. Only `text`-family fields are analyzed. A `keyword` field is indexed verbatim; it accepts a `normalizer`, not an analyzer. ## The built-ins, one at a time **standard** — the default. Tokenizer: the standard tokenizer, which implements Unicode text segmentation (UAX #29 word boundaries). That means punctuation and symbols are treated as separators and thrown away, while letters and digits survive. Then a `lowercase` token filter runs. It also contains a stop filter, but it is configured with an **empty** stopword list by default, so `the` is indexed unless you set the `stopwords` (or `stopwords_path`) parameter. It also exposes `max_token_length`. **simple** — tokenizer only: the `lowercase` tokenizer, which splits at every character that is not a letter and lowercases as it goes. The consequence people forget is that **digits are separators and are discarded**: `iphone 15` yields just `iphone`. That makes it a poor fit for product catalogues and version strings. **whitespace** — the `whitespace` tokenizer and nothing else. Tokens are exactly what was between spaces, with original case and any attached punctuation. `JUMPED!` and `jumped` become two different terms, so matching is case-sensitive. **keyword** — a no-op analyzer: the whole input becomes one token, unchanged. It exists so you can keep a field in the `text` family while treating its value atomically; for genuinely exact values, mapping the field as the `keyword` **type** is the normal choice. Do not confuse the two — the keyword *analyzer* and the keyword *field type* are different things that happen to share a name. **stop** — like `simple` (lowercase tokenizer) plus a stop filter that removes English stopwords by default; the list is configurable. **pattern** — splits with a regular expression, defaulting to a non-word-character pattern, and lowercases by default. Useful for structured strings such as camelCase or delimiter-separated codes. **Language analyzers** (`english`, `french`, `german`, …) — tokenization plus language-specific stopwords and stemming. ## A worked example Input: `The Quick-Brown Foxes JUMPED! 2 times` - **standard** → `the`, `quick`, `brown`, `foxes`, `jumped`, `2`, `times` - **simple** → `the`, `quick`, `brown`, `foxes`, `jumped`, `times` (the `2` is gone) - **whitespace** → `The`, `Quick-Brown`, `Foxes`, `JUMPED!`, `2`, `times` - **keyword** → `The Quick-Brown Foxes JUMPED! 2 times` (one token) - **stop** → `quick`, `brown`, `foxes`, `jumped`, `times` You never have to memorize this: the `_analyze` API prints the tokens for any analyzer and any input, and checking it is the first move when a search misbehaves. ## How to choose For human-readable prose you want recall, so `standard` is the sane default and a language analyzer is the upgrade when the corpus is single-language. For identifiers, SKUs, enum values, hostnames and status codes you want exactness, so use a `keyword` field — optionally with a lowercase normalizer if you need case-insensitive filtering. `whitespace` is the tool when the punctuation itself is meaningful (`C++`, `AT&T`, log tokens) and you are willing to handle case yourself. `simple` and `stop` are rarely the right production choice on their own; they mostly show up as teaching examples or as the base for something custom. ## Common traps The analyzer applies to both indexing and searching, so changing it changes what queries mean, not only what is stored — and existing documents keep their old tokens until they are reindexed. Choosing `whitespace` silently makes the field case-sensitive, which usually surfaces as "search works for lowercase queries only". Choosing `simple` silently destroys every number in the corpus. And assuming `standard` removes English stopwords is a very common wrong answer in interviews: its stopword list is empty until you configure one.

  • What does the stop analyzer add compared with the simple analyzer?
    Both use the lowercase tokenizer, so both split at non-letters and lowercase. The stop analyzer then removes stopwords, using an English list by default and accepting a custom `stopwords` array or `stopwords_path`. That raises precision on common words but makes phrases such as "to be or not to be" unmatchable, because nothing is left to index.
  • When would you pick the pattern analyzer over the standard analyzer?
    When the field has machine-made structure rather than prose — camelCase identifiers, pipe- or colon-delimited codes, log keys. The pattern analyzer tokenizes with a regular expression (defaulting to splitting on non-word characters) and lowercases by default, so you can express exactly where the boundaries are instead of accepting Unicode word segmentation.
  • Does the keyword analyzer do the same thing as mapping a field as keyword?
    They produce a single indexed term either way, but they are different mechanisms. The keyword *analyzer* is a no-op analyzer usable on a `text` field; the keyword *type* is a field type that skips analysis, supports doc_values for sorting and aggregations, and accepts a `normalizer`. For exact-match values the field type is the normal choice.

Think of the four as scissors with different blade widths: standard trims at every word boundary, simple hacks at anything that is not a letter, whitespace cuts only where you already left a gap, and keyword leaves the sheet of paper whole.

saying these in an interview costs you the question

  • Says the whitespace analyzer lowercases its tokens
  • Confuses the keyword analyzer with the keyword field type
  • Claims the simple analyzer keeps numbers in the text
  • Believes the standard analyzer removes English stopwords by default
  • Thinks an analyzer also applies to keyword-typed fields

context