skip to content

Analyzers and Tokenization

The pipeline that turns text into the terms actually stored in the index: character filters, a tokenizer, then token filters. Interviewers probe it because search only matches what analysis produced, and index-time and query-time analysis have to agree or nothing matches at all.

part ofElasticsearchoverview, primer and where to startread it →
on this pageshow

explore

questions

18

How do you use the _analyze API to see the exact tokens Elasticsearch produced for a field?

level: juniorimportance: must knowfreq 66%

answer

  1. One endpoint answers "what got indexed?"
  2. It can run against a named analyzer or a mapped field
  3. Index-scoped form resolves the real mapping
  4. A boolean flag shows each stage separately
  5. It simulates; it does not read stored terms

basics

~20 s

Call POST /_analyze with an analyzer name and text to test any built-in analyzer, or POST /<index>/_analyze with a field and text to run that field's configured analyzer. Add explain: true to see each pipeline stage's output.

solid answer

~40 s

`_analyze` is the debugging endpoint for analysis. In its cluster-level form you pass `analyzer` (or a `tokenizer` plus `filter`/`char_filter` combination) and `text`, and it returns the token stream: each token's text, `start_offset`, `end_offset`, `type` and `position`. Scoped to an index — `POST /my-index/_analyze` — you can instead pass `"field": "title"`, and Elasticsearch runs whatever analyzer that field actually resolves to, including the index default. That is the version you want when debugging, because it removes guesswork about the mapping. Setting `"explain": true` returns a per-stage breakdown showing the output of the char filters, the tokenizer, and each token filter in turn, which is how you find which stage mangled a term. `_analyze` shows what the analyzer *would* produce; it does not read existing documents.

code

json · 5 lines
json
POST /_analyze
{
  "analyzer": "english",
  "text": "The Foxes were running"
}

go deeper

for a junior

Be able to write the request from memory: POST /_analyze with an analyzer and text, or POST /index/_analyze with a field and text. Knowing this endpoint is the expected answer to "how would you debug a search that matches nothing?".

for a middle

Explain what the output fields mean — offsets drive highlighting, positions drive phrase matching — and what explain: true adds. Be clear that the API simulates the analyzer rather than reading indexed terms.

for a senior

Demonstrate the diagnostic routine: analyze the indexed text and the query text, compare the token sets, then rule out unanalyzed term queries, wrong sub-fields, and documents that predate a mapping change.

for a principal

Frame it as tooling: analysis regressions should be caught before deploy, so encode expected token streams as tests against _analyze in CI rather than relying on engineers to run ad-hoc requests after an incident.

## Why this API exists Almost every "why doesn't my search match?" bug in Elasticsearch is an analysis bug: the term you are searching for is not the term that was indexed. `_analyze` is the microscope. It runs the analysis chain on text you supply and prints the resulting token stream, so you can compare what the index holds against what the query produces instead of guessing. ## The three ways to call it **1. By analyzer name, cluster-scoped.** ``` POST /_analyze { "analyzer": "standard", "text": "The Quick Brown Fox" } ``` This needs no index and is how you compare built-ins side by side. **2. By field, index-scoped.** ``` POST /my-index/_analyze { "field": "title", "text": "The Quick Brown Fox" } ``` Elasticsearch resolves the analyzer the way it would at index time — the field mapping's `analyzer`, otherwise the index default, otherwise `standard`. This is the form to reach for during debugging, because it tests the actual configuration rather than your recollection of it. The endpoint must be index-scoped; `field` on the cluster-level endpoint has no index to resolve against. A `normalizer` can be tested the same way on a keyword field. **3. By assembled parts.** You can pass `tokenizer` plus `filter` and `char_filter` arrays to prototype a chain before you commit it to index settings. ## Reading the output Each token comes back with: - `token` — the indexed term. - `start_offset` / `end_offset` — character offsets into the original text; these drive highlighting. - `type` — the tokenizer's classification, for example `<ALPHANUM>` or `<NUM>`. - `position` — the ordinal used for phrase and proximity matching. Gaps here explain why a phrase query fails after stopword removal, and `positionLength` shows tokens that span several positions. ## explain: true Adding `"explain": true` changes the shape of the response: instead of a flat token list you get the intermediate state after the character filters, after the tokenizer, and after each token filter, each labelled by name. This is how you localize the damage — for example seeing that the tokenizer produced `foxes` correctly and a stemmer collapsed it later. You can narrow the verbose output with the `attributes` parameter, which limits which token attributes are shown. ## What _analyze does not tell you It is a simulation of the analyzer, not a dump of the index. It cannot tell you what terms are actually stored — if the mapping changed after the documents were written, `_analyze` shows the *new* behaviour while the index still holds the *old* terms. To inspect what is really stored, use the `_termvectors` API for a specific document, or a `terms` aggregation on a keyword field. This distinction catches people out constantly: the analyzer looks right, the search still fails, and the answer is that the documents predate the mapping change and need reindexing. ## The debugging routine A productive sequence when a query misses: 1. Run `_analyze` with `field` on the *indexed* text. Note the tokens. 2. Run `_analyze` with `field` on the *query* text. Note those tokens. 3. Compare. If the two token sets do not intersect, the problem is analysis, not query type. 4. If they do intersect and there are still no hits, look elsewhere: the query might be a `term` query, which does not analyze its input at all and therefore must be given an indexed term verbatim; or you might be querying a multi-field such as `title.keyword` instead of `title`; or the document may not be searchable yet because the index has not refreshed. ## Practical notes `_analyze` accepts `text` as either a string or an array of strings — an array is analyzed as if the values were separate values of the same field, which is useful for checking position gaps between array entries. The API is cheap and read-only, so it is safe to run against production. And because it accepts arbitrary analyzer definitions inline, it is the right place to iterate on a chain before writing it into index settings, where changes are far more expensive to apply.

  • The _analyze output looks correct but your query still returns nothing — what do you check next?
    Check whether the query analyzes at all: a `term` query is not analyzed, so it must be handed an exact indexed term. Then confirm you are querying the field you think you are rather than a `.keyword` sub-field, and confirm the documents were indexed *after* the current mapping — `_analyze` shows today's analyzer, while old documents keep the terms they were written with until reindexed.
  • How can you inspect the terms actually stored for a document rather than what an analyzer would produce?
    Use the `_termvectors` API for a single document (or `_mtermvectors` for several), which returns the real terms, their frequencies and positions for the requested fields. For a field-wide view of stored values on a keyword field, a `terms` aggregation shows what is actually there. Both read the index; `_analyze` only simulates.
  • Why would you pass an array of strings to _analyze instead of one string?
    Because an array is analyzed as multiple values of the same field, which reproduces how Elasticsearch handles array-valued fields — including the position gap inserted between values. That is exactly what you need when a phrase query unexpectedly matches or fails across two array entries.

saying these in an interview costs you the question

  • Thinks _analyze dumps the terms stored in the index
  • Uses the cluster-level endpoint and expects field resolution to work
  • Confuses _analyze explain with the _explain scoring API
  • Believes _analyze needs a document to exist first
  • Never checks the query side, only the indexed side

context

open as a page

How do Elasticsearch's standard, simple, whitespace, and keyword analyzers differ on the same input?

level: juniorimportance: must knowfreq 74%

basics

~20 s

The standard analyzer splits on Unicode word boundaries and lowercases; simple splits at any non-letter and lowercases, so digits disappear; whitespace splits only on spaces and keeps case and punctuation; keyword emits the whole input as one token.

open as a page

How is a custom analyzer composed in Elasticsearch, and in what order do its parts run?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A custom analyzer is zero or more character filters, exactly one tokenizer, then zero or more token filters. Character filters rewrite the raw string, the tokenizer splits it into tokens, and token filters transform that token stream in listed order.

open as a page

When is a text field analyzed in Elasticsearch — at index time, at search time, or both?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Both. Elasticsearch analyzes a text field's value when the document is indexed and stores the resulting terms; it analyzes the query string of a full-text query at search time. A hit requires the two term sets to overlap.

open as a page

What tokens do the ngram and edge_ngram tokenizers emit, and how does that drive index size?

level: middleimportance: must knowfreq 64%

basics

~20 s

edge_ngram emits only prefixes anchored at the start, so Quick with grams 2 to 4 gives Qu, Qui, Quic. ngram emits substrings at every position, producing far more terms and a much larger index for infix matching.

open as a page

Why should an edge_ngram autocomplete field in Elasticsearch use a different analyzer at search time?

level: middleimportance: must knowfreq 60%

basics

~20 s

Because n-gramming the query too turns each typed word into many short prefixes, which match almost everything and destroy precision and scoring. Prefixes belong in the index; the query should contribute the whole typed token as one term.

open as a page

How does Elasticsearch decide which analyzer applies to a text field at index time?

level: middleimportance: should knowfreq 46%

basics

~10 s

Elasticsearch uses the field mapping's analyzer parameter first; if the field does not declare one it falls back to the index's analysis.analyzer.default setting, and if that is not defined it uses the standard analyzer.

open as a page

What does Elasticsearch's english analyzer do to text that the standard analyzer does not?

level: middleimportance: should knowfreq 56%

basics

~10 s

The english analyzer adds English-specific processing on top of standard tokenization: it strips possessives, removes English stopwords, and stems words to a common root, so "running" and "runs" both index as "run".

open as a page

In Elasticsearch, why does a keyword field need a normalizer instead of an analyzer?

level: middleimportance: should knowfreq 52%

basics

~20 s

keyword fields are indexed verbatim and reject the analyzer parameter. A normalizer applies character-level filters such as lowercase or asciifolding and must emit exactly one token, so exact matching, sorting and aggregations become case- or accent-insensitive without tokenizing the value.

open as a page

Why must html_strip and the mapping character filter run before the tokenizer rather than after it?

level: middleimportance: should knowfreq 44%

basics

~20 s

Character filters rewrite the raw text, so they change what the tokenizer can see and how it splits. A token filter runs after splitting and cannot recover characters the tokenizer already dropped, such as the hash in C#.

open as a page

An Elasticsearch analyzer lists filters as ["stop", "lowercase"]; why do capitalized stopwords survive?

level: middleimportance: should knowfreq 50%

basics

~10 s

The stop filter's default English stopword list is lowercase and it does not ignore case, so a token like The does not match and is kept, then lowercased afterwards. Putting lowercase first removes it.

open as a page

In an Elasticsearch mapping, what is the difference between the analyzer and search_analyzer parameters?

level: middleimportance: should knowfreq 70%

basics

~20 s

The analyzer parameter defines the pipeline that turns a field's value into indexed terms, and is also used on queries unless overridden. search_analyzer overrides only the query side, so the terms a search produces can differ from the terms stored.

open as a page

Why can't Elasticsearch's standard analyzer distinguish between "C++", "C#", and "C"?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The standard analyzer tokenizes with Unicode text segmentation, which treats punctuation and symbols as separators and discards them, so all three inputs reduce to the single token "c". The distinguishing characters never reach the index.

open as a page

Why does synonym_graph handle multi-word synonyms correctly where the older synonym filter does not?

level: seniorimportance: should knowfreq 40%

basics

~10 s

synonym_graph emits a graph token stream, so a multi-word replacement occupies the right span of positions. The plain synonym filter stacks the replacement at one position, corrupting positions and breaking phrase and proximity matching.

open as a page

What does it take to change the index-time analyzer of a field on an existing Elasticsearch index?

level: seniorimportance: should knowfreq 55%

basics

~20 s

You cannot update a field's analyzer parameter in place, and existing documents are never re-analyzed. Create a new index with the new analysis settings and mapping, reindex into it, and swap an alias — or add the analyzer to a closed index and reindex anyway.

open as a page

Should synonym expansion in Elasticsearch run at index time or at search time, and why?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Search time is the usual choice: rules can change without reindexing, the index stays clean, and synonym_graph handles multi-word synonyms correctly. Index-time expansion avoids per-query cost but bakes rules into segments and distorts term statistics.

open as a page

When would you add a shingle token filter to an Elasticsearch analysis chain, and what does it cost?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Shingles turn adjacent tokens into word n-grams, so word adjacency becomes an ordinary term you can match and boost. The cost is a much larger term dictionary, so they belong on a dedicated subfield rather than the main text field.

open as a page

What does the search_quote_analyzer mapping parameter do in Elasticsearch?

level: seniorimportance: nice to knowfreq 20%

basics

~10 s

It sets a separate analyzer for text inside quotes in query_string and simple_query_string queries. The usual purpose is to keep stopwords in quoted phrases while the ordinary search analyzer strips them.

open as a page