skip to content

Why must html_strip and the mapping character filter run before the tokenizer rather than after it?

level: middleimportance: should knowfreq 44%

answer

  1. Runs on characters, not tokens
  2. Splitting depends on what it sees
  3. Some characters never reach the filters
  4. Markup and symbols are the usual targets
  5. Offsets are corrected for highlighting

basics

~20 s

Character filters rewrite the raw text, so they change what the tokenizer can see and how it splits. A token filter runs after splitting and cannot recover characters the tokenizer already dropped, such as the hash in C#.

solid answer

~40 s

Character filters act on the character stream, and the tokenizer's splitting decisions depend on exactly those characters. `html_strip` removes markup and decodes entities, so `<p>Tom &amp; Jerry</p>` reaches the tokenizer as clean text instead of producing tokens like `p` and `amp`. The `mapping` filter substitutes literal strings you list, which is the standard trick for symbols the `standard` tokenizer discards: mapping `"C# => csharp"` preserves the term, because after tokenization the `#` is simply gone and no token filter can bring it back. `pattern_replace` does the same with a regex. Elasticsearch tracks offset corrections through character filters so highlighting still points at the original text, though large substitutions can make highlight spans approximate. Multiple character filters are allowed and run in listed order, each consuming the previous one's output.

code

json · 20 lines
json
PUT /docs
{
  "settings": {
    "analysis": {
      "char_filter": {
        "tech_terms": {
          "type": "mapping",
          "mappings": ["C# => csharp", "C++ => cplusplus", "& => and"]
        }
      },
      "analyzer": {
        "tech_body": {
          "char_filter": ["html_strip", "tech_terms"],
          "tokenizer": "standard",
          "filter": ["lowercase"]
        }
      }
    }
  }
}

go deeper

for a junior

Know that character filters clean the text before it is split, and that html_strip removes markup so tag names do not become searchable terms.

for a middle

Explain the causal link: the tokenizer's splitting depends on the characters it receives, so symbol preservation such as C# has to happen in a character filter, not later.

for a senior

Discuss the operational side — offset correction and highlight accuracy, the cost of pattern_replace on ingest throughput, and distributing a mappings_path file to every node.

for a principal

Decide where text normalization belongs across the system: some of it is cheaper and more testable in the ingest pipeline or the producing service than baked into every index's analysis settings.

## Where character filters sit An Elasticsearch analyzer runs character filters, then one tokenizer, then token filters. The first stage is unusual in that it does not work on tokens at all — it works on the raw character stream, the same string the document field contained. That placement is the whole point of the stage. ## Why the position is load-bearing The tokenizer decides boundaries from characters. The `standard` tokenizer implements Unicode text segmentation, which treats most punctuation and symbols as separators to discard. So `C#` becomes the token `c`, `AT&T` becomes `at` and `t`, and a smiley face vanishes entirely. Once that has happened the information is gone: token filters can lowercase, fold, drop, or duplicate tokens, but they cannot reach back into text the tokenizer never handed them. A character filter fixes the problem upstream. Mapping `"C# => csharp"` and `"C++ => cplusplus"` gives the tokenizer a string it will keep whole. Mapping `":) => _happy_"` turns an emoticon into something indexable. The rule of thumb: if the tokenizer would destroy it, rewrite it before the tokenizer. ## html_strip `html_strip` removes HTML elements and replaces HTML entities with their decoded characters. Indexing raw HTML without it produces a term dictionary full of `div`, `href`, `nbsp`, and `amp`, which both bloats the index and pollutes relevance — a document is not "about" the tag names in its markup. The filter accepts `escaped_tags`, a list of element names to leave intact when you deliberately want the markup indexed. A subtlety worth knowing: stripping is not the same as sanitizing. `html_strip` is an analysis-time convenience for search quality, not a security control for what you render back to a browser; the stored `_source` still contains the original markup. ## mapping The `mapping` character filter takes a list of `key => value` rules and replaces every occurrence of the key with the value. Keys are literal strings, not patterns, and longer keys win over shorter overlapping ones. Rules can also be loaded from a file with `mappings_path`, resolved relative to the node's config directory — useful when the list is long, at the cost of having to distribute the file to every node. Typical uses: normalizing symbols the tokenizer eats, transliterating digit systems (`"٣ => 3"`), unifying spellings before tokenization, and collapsing separators so `wi-fi` and `wifi` become the same term. ## pattern_replace The third built-in, `pattern_replace`, applies a Java regular expression with a replacement string. It is the flexible option and also the slow one — a badly written pattern runs on every character of every document — so prefer `mapping` when a literal list will do. ## Offsets and highlighting Tokens carry character offsets pointing into the original text, and highlighting uses those offsets to wrap fragments of the stored value. Because character filters change the text's length, Elasticsearch maintains offset correction so the tokenizer's offsets map back to the pre-filter string. This works well for stripping tags and for one-for-one substitutions; drastic replacements where a short key becomes a long value can make highlight boundaries imprecise, which is a known trade-off rather than a bug you can configure away. ## Ordering among character filters The `char_filter` array is a pipeline like the others. `["html_strip", "my_mapping"]` means the mapping rules see decoded, tag-free text — so a rule keyed on `&amp;` will never fire, because the entity was already decoded to `&`. Reverse the order and the opposite is true. Interviewers like this because it is a small, concrete demonstration of whether you think of these arrays as pipelines. ## When not to use one If the transformation is per-token and does not affect splitting — lowercasing, accent folding, stemming, stopword removal — it belongs in a token filter, which is cheaper to reason about and does not disturb offsets. Character filters are for the cases where the tokenizer's own behaviour is the problem, or where the input is not plain prose to begin with.

  • How would you keep certain HTML elements while stripping the rest?
    `html_strip` takes an `escaped_tags` parameter listing element names to leave in place; everything else is removed and entities are still decoded. It is worth asking why you want them indexed at all — tag names rarely help relevance, and keeping them usually signals that the markup carries semantics that would be better extracted into their own field.
  • Why prefer the mapping character filter over pattern_replace when both would work?
    `mapping` matches literal strings and is fast and predictable; `pattern_replace` compiles a regular expression that runs over every character of every document and can degrade indexing throughput badly on pathological patterns. Reach for the regex only when the transformation genuinely cannot be expressed as a finite list of substitutions.

saying these in an interview costs you the question

  • Thinks html_strip is a token filter
  • Claims a token filter could restore the discarded hash
  • Believes html_strip sanitizes the stored _source
  • Treats the char_filter array as order-independent
  • Uses pattern_replace where a literal mapping list suffices

context