skip to content

Custom Analysis Chains

A custom analyzer is a character-filter list, exactly one tokenizer, and a token-filter chain — that composition is how you build autocomplete, synonyms, or accent-insensitive matching. Interviewers ask you to assemble one out loud, in the right order.

part ofElasticsearchoverview, primer and where to startread it →
on this pageshow

questions

6

How is a custom analyzer composed in Elasticsearch, and in what order do its parts run?

level: juniorimportance: must knowfreq 78%

answer

  1. Three stages, one of them singular
  2. Raw characters, then splitting, then tokens
  3. Only one component may split text
  4. Lists are pipelines, not sets
  5. char_filter, tokenizer, filter under analysis

basics

~20 s

A custom analyzer is zero or more character filters, exactly one tokenizer, then zero or more token filters. Character filters rewrite the raw string, the tokenizer splits it into tokens, and token filters transform that token stream in listed order.

solid answer

~40 s

An Elasticsearch analyzer always has the same three-stage shape: **character filters** (zero or more), **one tokenizer**, and **token filters** (zero or more). Character filters see the raw character stream before anything is split — `html_strip` removes markup, `mapping` substitutes strings. The tokenizer is the only stage that decides token boundaries, and there is exactly one; it emits tokens carrying positions and offsets. Token filters then add, remove, or rewrite tokens one stage at a time: `lowercase`, `asciifolding`, `stop`, `stemmer`, `synonym_graph`, `shingle`. You define the analyzer under `settings.analysis.analyzer` at index creation, referencing built-in components by name or referencing named component definitions you created under `settings.analysis.char_filter`, `analysis.tokenizer`, and `analysis.filter`. Order matters within every list: the arrays are pipelines, not sets, and the same filters in a different order produce different terms.

code

json · 21 lines
json
PUT /articles
{
  "settings": {
    "analysis": {
      "char_filter": {
        "amp_to_and": { "type": "mapping", "mappings": ["& => and"] }
      },
      "filter": {
        "en_stem": { "type": "stemmer", "language": "english" }
      },
      "analyzer": {
        "article_body": {
          "type": "custom",
          "char_filter": ["html_strip", "amp_to_and"],
          "tokenizer": "standard",
          "filter": ["lowercase", "asciifolding", "en_stem"]
        }
      }
    }
  }
}

go deeper

for a junior

Memorize the shape and be able to say it in one breath: zero or more character filters, exactly one tokenizer, zero or more token filters, applied in that order.

for a middle

Explain why each stage exists — character filters change what the tokenizer can see, the tokenizer assigns positions and offsets, token filters rewrite the stream — and walk a concrete string through a chain out loud.

for a senior

Be ready to assemble a chain on the spot for a stated requirement and to justify the placement of each filter, including which components need their own named definition because they carry parameters.

for a principal

Own the convention: how many analyzers an index should carry, whether chains are shared through component templates or copy-pasted per index, and the maintenance cost of every bespoke chain a team invents.

## The fixed shape of an analyzer Every Elasticsearch analyzer — built-in or custom — is the same pipeline: a character-filter list, exactly one tokenizer, and a token-filter list. Text goes in as a string and comes out as a stream of tokens, and it is those tokens, not the original text, that end up in the inverted index. Understanding the shape is most of what an interviewer wants, because everything else (autocomplete, synonyms, accent-insensitive search) is just a choice of components inside it. ## Stage 1 — character filters Character filters receive the raw character stream, before any splitting has happened. They can add, remove, or replace characters. The three built-ins are `html_strip` (removes HTML elements and decodes entities), `mapping` (replaces literal strings you list, such as `"& => and"`), and `pattern_replace` (regex substitution). An analyzer may declare several; they run in the order listed, each feeding the next. They exist as a separate stage precisely because the tokenizer's decisions depend on the characters it is handed. If you want `C#` to survive the standard tokenizer, you must rewrite it to something the tokenizer will not split *before* tokenization; no token filter can bring back a character the tokenizer already discarded. ## Stage 2 — the tokenizer Exactly one tokenizer is required and exactly one is allowed. It converts characters into tokens and assigns each token a position (used by phrase queries and by span-style matching) and a character offset range (used by highlighting to point back into the original text). The `standard` tokenizer splits on Unicode text-segmentation rules and drops most punctuation; `whitespace` splits only on spaces; `pattern` splits on a regex; `ngram` and `edge_ngram` emit substrings; `path_hierarchy` emits successive path prefixes; `keyword` emits the entire input as a single token. Because there is only one tokenizer, "I want to split both on whitespace and on slashes" is solved with a character filter or a regex tokenizer, not by chaining two tokenizers. ## Stage 3 — token filters Token filters operate on the token stream, one token (or one position) at a time, in the order listed. Typical members: `lowercase`, `asciifolding` (folds `café` to `cafe`), `stop` (drops stopwords), `stemmer` (reduces `running` to `run`), `synonym_graph`, `shingle` (word n-grams), `ngram`/`edge_ngram` (character grams applied per token), `keyword_marker` (protects words from stemming), `length`, `unique`. The list is a pipeline. `["lowercase", "stop"]` and `["stop", "lowercase"]` do different things, because the default English stopword list is lowercase and `stop` does not ignore case by default. ## Defining versus referencing Built-in components with default settings can be named directly in the analyzer definition: `"char_filter": ["html_strip"]`, `"filter": ["lowercase"]`. A component you need to configure gets a named definition first, under the matching section of `settings.analysis`, and the analyzer then references that name: ```json "analysis": { "filter": { "en_stem": { "type": "stemmer", "language": "english" } }, "analyzer": { "body_en": { "tokenizer": "standard", "filter": ["lowercase", "en_stem"] } } } ``` The names you invent live in the index's namespace, so two indices can each have a `body_en` meaning different things. The `type` on the analyzer itself may be set to `custom` or omitted entirely — omitting it with a `tokenizer` present is the same declaration. ## Degenerate and edge cases An analyzer with only a tokenizer is legal and common (`whitespace` alone). Zero character filters is the normal case. What is not legal is zero tokenizers or two of them. A tokenizer name that is not a built-in and not defined under `analysis.tokenizer` fails index creation, which is the usual cause of a "failed to find tokenizer" error on a first attempt. ## Why the order questions get asked The pipeline is deterministic and cheap to reason about, so interviewers use it to check whether you actually understand indexing rather than having copied a recipe. Being able to say aloud "character filters on the raw string, one tokenizer, then filters in order" and then walk a concrete input — `"<b>Cafés</b> & bars"` through `html_strip`, a `mapping` filter turning `&` into `and`, the `standard` tokenizer, then `lowercase` and `asciifolding` — is the answer they are listening for.

  • Is an analyzer with no character filters and no token filters valid?
    Yes. The tokenizer is the only mandatory stage, so an analyzer that declares just `whitespace` or `keyword` and nothing else is a complete, legal definition. Zero character filters is in fact the common case; most chains only reach for them when the source text carries markup or punctuation the tokenizer would mangle.
  • What happens if you list two character filters in a custom analyzer?
    They run in the order listed, each consuming the previous one's output. `["html_strip", "amp_to_and"]` strips markup first, so the `mapping` filter sees the decoded, tag-free text. Reversing them means the mapping rules are applied to text that still contains tags and entities, which usually changes the result.
  • When do you have to define a component under analysis.filter instead of naming it inline?
    Whenever you need non-default settings. `lowercase` with defaults can be named directly in the analyzer's filter list, but a stemmer for a specific language, a stopword list, or an `edge_ngram` with chosen gram sizes needs a named definition carrying `type` plus its parameters, which the analyzer then references by that name.

Think of an assembly line: the character filters clean the raw material, a single saw cuts it into pieces, and the finishing stations sand, stamp, and sort each piece in a fixed order.

saying these in an interview costs you the question

  • Says an analyzer can chain several tokenizers together
  • Thinks token filters can undo what the tokenizer discarded
  • Treats the filter array as an unordered set
  • Confuses character filters with token filters
  • Believes a custom analyzer needs at least one character filter

context

open as a page

What tokens do the ngram and edge_ngram tokenizers emit, and how does that drive index size?

level: middleimportance: must knowfreq 64%

basics

~20 s

edge_ngram emits only prefixes anchored at the start, so Quick with grams 2 to 4 gives Qu, Qui, Quic. ngram emits substrings at every position, producing far more terms and a much larger index for infix matching.

open as a page

Why must html_strip and the mapping character filter run before the tokenizer rather than after it?

level: middleimportance: should knowfreq 44%

basics

~20 s

Character filters rewrite the raw text, so they change what the tokenizer can see and how it splits. A token filter runs after splitting and cannot recover characters the tokenizer already dropped, such as the hash in C#.

open as a page

An Elasticsearch analyzer lists filters as ["stop", "lowercase"]; why do capitalized stopwords survive?

level: middleimportance: should knowfreq 50%

basics

~10 s

The stop filter's default English stopword list is lowercase and it does not ignore case, so a token like The does not match and is kept, then lowercased afterwards. Putting lowercase first removes it.

open as a page

Why does synonym_graph handle multi-word synonyms correctly where the older synonym filter does not?

level: seniorimportance: should knowfreq 40%

basics

~10 s

synonym_graph emits a graph token stream, so a multi-word replacement occupies the right span of positions. The plain synonym filter stacks the replacement at one position, corrupting positions and breaking phrase and proximity matching.

open as a page

When would you add a shingle token filter to an Elasticsearch analysis chain, and what does it cost?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Shingles turn adjacent tokens into word n-grams, so word adjacency becomes an ordinary term you can match and boost. The cost is a much larger term dictionary, so they belong on a dedicated subfield rather than the main text field.

open as a page