skip to content

How do you use the _analyze API to see the exact tokens Elasticsearch produced for a field?

level: juniorimportance: must knowfreq 66%

answer

  1. One endpoint answers "what got indexed?"
  2. It can run against a named analyzer or a mapped field
  3. Index-scoped form resolves the real mapping
  4. A boolean flag shows each stage separately
  5. It simulates; it does not read stored terms

basics

~20 s

Call POST /_analyze with an analyzer name and text to test any built-in analyzer, or POST /<index>/_analyze with a field and text to run that field's configured analyzer. Add explain: true to see each pipeline stage's output.

solid answer

~40 s

`_analyze` is the debugging endpoint for analysis. In its cluster-level form you pass `analyzer` (or a `tokenizer` plus `filter`/`char_filter` combination) and `text`, and it returns the token stream: each token's text, `start_offset`, `end_offset`, `type` and `position`. Scoped to an index — `POST /my-index/_analyze` — you can instead pass `"field": "title"`, and Elasticsearch runs whatever analyzer that field actually resolves to, including the index default. That is the version you want when debugging, because it removes guesswork about the mapping. Setting `"explain": true` returns a per-stage breakdown showing the output of the char filters, the tokenizer, and each token filter in turn, which is how you find which stage mangled a term. `_analyze` shows what the analyzer *would* produce; it does not read existing documents.

code

json · 5 lines
json
POST /_analyze
{
  "analyzer": "english",
  "text": "The Foxes were running"
}

go deeper

for a junior

Be able to write the request from memory: POST /_analyze with an analyzer and text, or POST /index/_analyze with a field and text. Knowing this endpoint is the expected answer to "how would you debug a search that matches nothing?".

for a middle

Explain what the output fields mean — offsets drive highlighting, positions drive phrase matching — and what explain: true adds. Be clear that the API simulates the analyzer rather than reading indexed terms.

for a senior

Demonstrate the diagnostic routine: analyze the indexed text and the query text, compare the token sets, then rule out unanalyzed term queries, wrong sub-fields, and documents that predate a mapping change.

for a principal

Frame it as tooling: analysis regressions should be caught before deploy, so encode expected token streams as tests against _analyze in CI rather than relying on engineers to run ad-hoc requests after an incident.

## Why this API exists Almost every "why doesn't my search match?" bug in Elasticsearch is an analysis bug: the term you are searching for is not the term that was indexed. `_analyze` is the microscope. It runs the analysis chain on text you supply and prints the resulting token stream, so you can compare what the index holds against what the query produces instead of guessing. ## The three ways to call it **1. By analyzer name, cluster-scoped.** ``` POST /_analyze { "analyzer": "standard", "text": "The Quick Brown Fox" } ``` This needs no index and is how you compare built-ins side by side. **2. By field, index-scoped.** ``` POST /my-index/_analyze { "field": "title", "text": "The Quick Brown Fox" } ``` Elasticsearch resolves the analyzer the way it would at index time — the field mapping's `analyzer`, otherwise the index default, otherwise `standard`. This is the form to reach for during debugging, because it tests the actual configuration rather than your recollection of it. The endpoint must be index-scoped; `field` on the cluster-level endpoint has no index to resolve against. A `normalizer` can be tested the same way on a keyword field. **3. By assembled parts.** You can pass `tokenizer` plus `filter` and `char_filter` arrays to prototype a chain before you commit it to index settings. ## Reading the output Each token comes back with: - `token` — the indexed term. - `start_offset` / `end_offset` — character offsets into the original text; these drive highlighting. - `type` — the tokenizer's classification, for example `<ALPHANUM>` or `<NUM>`. - `position` — the ordinal used for phrase and proximity matching. Gaps here explain why a phrase query fails after stopword removal, and `positionLength` shows tokens that span several positions. ## explain: true Adding `"explain": true` changes the shape of the response: instead of a flat token list you get the intermediate state after the character filters, after the tokenizer, and after each token filter, each labelled by name. This is how you localize the damage — for example seeing that the tokenizer produced `foxes` correctly and a stemmer collapsed it later. You can narrow the verbose output with the `attributes` parameter, which limits which token attributes are shown. ## What _analyze does not tell you It is a simulation of the analyzer, not a dump of the index. It cannot tell you what terms are actually stored — if the mapping changed after the documents were written, `_analyze` shows the *new* behaviour while the index still holds the *old* terms. To inspect what is really stored, use the `_termvectors` API for a specific document, or a `terms` aggregation on a keyword field. This distinction catches people out constantly: the analyzer looks right, the search still fails, and the answer is that the documents predate the mapping change and need reindexing. ## The debugging routine A productive sequence when a query misses: 1. Run `_analyze` with `field` on the *indexed* text. Note the tokens. 2. Run `_analyze` with `field` on the *query* text. Note those tokens. 3. Compare. If the two token sets do not intersect, the problem is analysis, not query type. 4. If they do intersect and there are still no hits, look elsewhere: the query might be a `term` query, which does not analyze its input at all and therefore must be given an indexed term verbatim; or you might be querying a multi-field such as `title.keyword` instead of `title`; or the document may not be searchable yet because the index has not refreshed. ## Practical notes `_analyze` accepts `text` as either a string or an array of strings — an array is analyzed as if the values were separate values of the same field, which is useful for checking position gaps between array entries. The API is cheap and read-only, so it is safe to run against production. And because it accepts arbitrary analyzer definitions inline, it is the right place to iterate on a chain before writing it into index settings, where changes are far more expensive to apply.

  • The _analyze output looks correct but your query still returns nothing — what do you check next?
    Check whether the query analyzes at all: a `term` query is not analyzed, so it must be handed an exact indexed term. Then confirm you are querying the field you think you are rather than a `.keyword` sub-field, and confirm the documents were indexed *after* the current mapping — `_analyze` shows today's analyzer, while old documents keep the terms they were written with until reindexed.
  • How can you inspect the terms actually stored for a document rather than what an analyzer would produce?
    Use the `_termvectors` API for a single document (or `_mtermvectors` for several), which returns the real terms, their frequencies and positions for the requested fields. For a field-wide view of stored values on a keyword field, a `terms` aggregation shows what is actually there. Both read the index; `_analyze` only simulates.
  • Why would you pass an array of strings to _analyze instead of one string?
    Because an array is analyzed as multiple values of the same field, which reproduces how Elasticsearch handles array-valued fields — including the position gap inserted between values. That is exactly what you need when a phrase query unexpectedly matches or fails across two array entries.

saying these in an interview costs you the question

  • Thinks _analyze dumps the terms stored in the index
  • Uses the cluster-level endpoint and expects field resolution to work
  • Confuses _analyze explain with the _explain scoring API
  • Believes _analyze needs a document to exist first
  • Never checks the query side, only the indexed side

context