Why can't Elasticsearch's standard analyzer distinguish between "C++", "C#", and "C"?
answer
- Unicode word segmentation is doing this
- Symbols are treated as separators
- Three different inputs, one identical token
- Match over-matches; term finds nothing
- The fix happens before the tokenizer runs
basics
~20 sThe standard analyzer tokenizes with Unicode text segmentation, which treats punctuation and symbols as separators and discards them, so all three inputs reduce to the single token "c". The distinguishing characters never reach the index.
solid answer
~40 sRun the text through `_analyze` with the `standard` analyzer and you get one token, `c`, for each of the three. Unicode word segmentation drops `+` and `#`, so the index cannot tell them apart, and a `match` query for `C++` — analyzed the same way — matches every document mentioning plain C. Related casualties are `AT&T` (becomes `at`, `t`) and `.NET` (becomes `net`). The fixes all restore the missing characters before tokenization or avoid tokenization: use the `whitespace` analyzer so `C++` survives intact, accepting that it no longer lowercases; add a `keyword` field for exact filtering; or map the problem tokens to letter-only forms with a character filter. The production shape is usually a `text` field for general recall plus a precise companion field, queried together so the exact form outranks the fuzzy one.
code
json · 7 linesPOST /_analyze
{ "analyzer": "standard", "text": "C++ C# C" }
// tokens: c, c, c
POST /_analyze
{ "analyzer": "whitespace", "text": "C++ C# C" }
// tokens: C++, C#, Cgo deeper
Recall that the standard analyzer throws punctuation away, so C++ and C end up as the same indexed token. Knowing to check with the _analyze API is most of the expected answer here.
Explain the mechanism — Unicode word segmentation treats symbols as separators — and describe both symptoms: a match query over-matches on the bare letter, while a term query for C++ finds nothing at all.
Diagnose before prescribing: prove it with _analyze on both the indexed and query text, then choose a remedy and state its cost, whether that is losing case-insensitivity with whitespace analysis or maintaining a character-filter mapping list.
Own the rollout and the pattern: symbol-bearing identifiers need a modelled exact field rather than ad-hoc fixes, the change means reindex behind an alias, and relevance impact should be measured before the switch rather than discovered by users.
## Reproducing it in ten seconds ```json POST /_analyze { "analyzer": "standard", "text": "C++ C# C" } ``` Three tokens come back and all three are `c`. That single observation explains the whole class of bug reports: "searching for C++ returns C# jobs", "our .NET pages rank against generic content", "AT&T matches everything with the word at". ## Why it happens The standard analyzer's tokenizer implements Unicode text segmentation (UAX #29 word boundaries). That algorithm is designed for natural language across scripts, and in natural language `+`, `#`, `&` and a leading `.` are punctuation — they mark boundaries and carry no lexical content. So the tokenizer treats them as separators and does not emit them. The subsequent lowercase filter has nothing to do with the loss; the characters were already gone. This is not a bug. It is a deliberate trade in favour of language-neutral prose handling, and it is right for the overwhelming majority of text. It is wrong for the minority of tokens where a symbol *is* the identity: programming languages, product SKUs, ticker symbols, chemical formulae, version strings. ## Two different failure modes It is worth being precise about what actually breaks, because candidates often describe the wrong symptom. **A `match` query over-matches.** Query analysis mirrors index analysis, so a `match` for `C++` also becomes `c`, which finds every document containing plain C. You get hits — the wrong ones. Relevance looks broken rather than empty. **A `term` query returns nothing.** Term-level queries are not analyzed. A `term` query for `C++` looks for the literal indexed term `C++`, which does not exist in a standard-analyzed field: only `c` does. Zero hits, and the naive conclusion ("term queries are broken") is wrong; the index simply never held that term. Diagnosing means checking both sides with `_analyze` — the indexed text and the query text — and comparing the token sets. If they do not intersect, or intersect too generously, the problem is analysis and no amount of query tuning will fix it. ## The remedies, and their costs **Whitespace analysis.** The `whitespace` analyzer splits only on spaces, so `C++` stays `C++`. The cost is everything else the standard analyzer was doing: no lowercasing, so matching becomes case-sensitive, and no punctuation trimming, so `C++,` and `C++.` at the end of a sentence become distinct terms from `C++`. This works well for controlled, machine-generated content and badly for prose. **An exact field.** Model the value as a `keyword` field — optionally with a lowercase normalizer — and filter on it exactly. This is the right answer when the field is a tag or a facet (`skills: ["C++", "C#"]`) rather than free text. Users pick from a list; you filter on a term. No analysis, no ambiguity. **Rewrite before tokenizing.** A character filter can map `C++` to a letter-only form such as `cplusplus` before the tokenizer ever sees it, preserving the distinction inside otherwise-normal text analysis. This is exact and cheap, but it is a maintained list: every new symbol-bearing token needs an entry, and index-time and query-time must use the same mapping or nothing matches. **Both fields at once.** The pattern that survives production is indexing the content twice — general text analysis for recall and a precise variant for exactness — and querying both, boosting the precise field so an exact `C++` match outranks a loose one. It costs index size and a slightly more complex query, and it avoids the false choice between "finds nothing" and "finds everything". ## The operational reality Whichever fix you choose, it is a mapping change, and mapping changes are not retroactive: documents already indexed keep the terms they were written with. So the change is really "build a new index with the new analysis, reindex, switch the alias" — plan the rollout, not just the setting. Before doing any of that, confirm the diagnosis with `_analyze` on the real field, because the same symptom can also come from querying the wrong sub-field or from a stemmer elsewhere in the chain. ## What a strong answer sounds like Name the mechanism (Unicode word segmentation discards symbols), prove it with `_analyze` rather than asserting it, distinguish over-matching from zero-matching, and then pick a remedy with its trade-off stated out loud. Reaching straight for a wildcard or an escaping trick in the query string is the weak answer: escaping affects query parsing, not the fact that the plus signs were never indexed.
- How would you prove this diagnosis in thirty seconds before touching the mapping?Run `_analyze` twice against the index with the `field` parameter — once on a sample of the indexed text, once on the query string — and compare token lists. Seeing `c` for all three inputs settles it immediately. If the tokens look right and hits are still wrong, the problem is elsewhere: an unanalyzed term query, the wrong sub-field, or documents predating the current mapping.
- What breaks if you simply switch the field to the whitespace analyzer?You lose lowercasing and punctuation trimming. Matching becomes case-sensitive, so `c++` no longer finds `C++`, and trailing punctuation makes `C++.` a different term from `C++`. That is acceptable for controlled, machine-generated values and usually unacceptable for prose, where a custom chain or a second exact field is the better answer.
- Why doesn't escaping the plus signs in the query string fix this?Escaping only tells the query parser to treat `+` as a literal rather than as an operator. It has no effect on analysis: the field's analyzer still strips the symbols from the query text, and in any case the index never contains a term with plus signs. The problem is on the indexing side and must be fixed there.
saying these in an interview costs you the question
- Blames the query type instead of the analysis chain
- Thinks escaping the plus signs in the query fixes it
- Assumes a term query would work without checking indexed terms
- Believes the standard analyzer keeps symbols like plus and hash
- Reaches for wildcard queries instead of fixing the mapping