In Elasticsearch, why does a keyword field need a normalizer instead of an analyzer?
answer
- This field type refuses the analyzer parameter
- No tokenizer is allowed here
- One value in, exactly one term out
- It runs on the query value too
- Doc values change, so facets change
basics
~20 skeyword fields are indexed verbatim and reject the analyzer parameter. A normalizer applies character-level filters such as lowercase or asciifolding and must emit exactly one token, so exact matching, sorting and aggregations become case- or accent-insensitive without tokenizing the value.
solid answer
~50 sA `keyword` field exists to hold a value atomically — one document, one term — so it is not analyzed and the `analyzer` mapping parameter is not accepted on it. A **normalizer** is the constrained equivalent: no tokenizer, and only filters that work character by character, such as `lowercase`, `asciifolding`, `uppercase`, `trim` and `elision`. Anything that could change the token count, like a stemmer or shingle filter, is rejected. Elasticsearch ships a built-in `lowercase` normalizer, and you can define custom ones under `settings.analysis.normalizer`. The important property is that the normalizer runs at **both** index and query time, so a `term` query for `"New York"` still matches a document stored as `New York` on a lowercased field. Because doc_values hold the normalized form, sorting and `terms` aggregations also use it — so the original casing is no longer available from that field.
code
json · 8 linesPUT /places
{
"mappings": {
"properties": {
"city": { "type": "keyword", "normalizer": "lowercase" }
}
}
}go deeper
Remember the rule: text fields take an analyzer, keyword fields take a normalizer. Knowing that a lowercase normalizer gives case-insensitive exact matching covers the usual question.
Explain the constraint that makes normalizers different — no tokenizer, only character-level filters, exactly one output token — and that the normalizer is applied to query values as well as indexed values.
Show the operational consequences: doc_values hold the normalized form so sorting and facet buckets change, display casing must be preserved elsewhere, and rolling this out on an existing index means a reindex behind an alias.
Own the modelling convention: which fields are exact-match by contract, how case and accent folding are standardized across the estate via templates, and how you avoid per-team drift where the same logical value normalizes differently in different indices.
## The problem You have a `status`, `city` or `tag` field. You want exact matching — not full-text search — but you also want `"NEW YORK"`, `"New York"` and `"new york"` to be the same value for filtering, sorting and faceting. Mapping the field as `text` gives you the case-insensitivity but destroys atomicity: `New York` becomes two terms and a filter for it also matches `York, New`. Mapping it as `keyword` gives you atomicity but makes matching case-sensitive. The normalizer resolves this. It is analysis for fields that must not be tokenized. ## What a normalizer is allowed to contain A normalizer has **no tokenizer**. It may contain character filters and a restricted set of token filters — the ones that operate character by character and cannot change how many tokens exist. `lowercase`, `uppercase`, `asciifolding`, `trim`, `elision`, `decimal_digit` and the various per-language normalization filters qualify; `pattern_replace` as a character filter is also common. Filters that split, join or drop tokens — stemmers, `shingle`, `ngram`, `stop` — are rejected when the index is created, because they would break the one-value-one-term guarantee. Elasticsearch ships a built-in `lowercase` normalizer, so the simplest case needs no analysis settings at all: ```json "city": { "type": "keyword", "normalizer": "lowercase" } ``` Custom normalizers are declared in index settings under `analysis.normalizer` with `type: custom`, a `char_filter` list and a `filter` list, then referenced by name from the mapping. ## Both sides get normalized The property that trips people up is symmetry. A `term` query does **not** analyze its input on an ordinary field — that is the whole point of term-level queries. But on a keyword field with a normalizer, the normalizer *is* applied to the query value as well. So searching `term: { city: "New York" }` against a lowercase-normalized field works: the query value becomes `new york`, and it matches the stored term `new york`. You do not have to lowercase in application code, and doing so "just in case" is harmless but unnecessary. ## What it does to sorting and aggregations Keyword fields build doc_values, the columnar structure used for sorting, aggregating and scripting. Those doc_values hold the **normalized** value. Practical consequences: - Sorting is now case-insensitive, which is usually what users expect — without a normalizer, a byte-order sort puts every capitalized value before every lowercase one. - A `terms` aggregation returns one bucket for `new york` rather than three buckets differing only in casing. That is generally the goal of adding the normalizer in the first place. - The bucket keys and sort keys are the normalized strings, so you lose the display casing from that field. If you need to show `New York` in a facet UI, keep the raw value in a second field (a multi-field is the usual shape) and use the normalized one only for grouping and filtering. ## Common mistakes **Putting `analyzer` on a keyword field.** It is not a valid parameter there; the mapping is rejected. People reach for it because analyzer is the familiar name. **Confusing the keyword *analyzer* with the keyword *type*.** The keyword analyzer is a no-op analyzer for `text` fields. The keyword type is a field type that skips analysis entirely and supports normalizers and doc_values. They solve overlapping problems by different means. **Expecting existing documents to change.** Adding a normalizer changes how values are processed going forward. Documents already in the index keep the terms they were written with, so the field behaves inconsistently until the data is reindexed. Analysis settings are also not freely mutable on an open index, which is why this is normally handled by building a new index and switching an alias. **Assuming a normalizer can do stemming or splitting.** It cannot, by design. If you need tokenization you need a `text` field — and often the right answer is both: a `text` sub-field for search and a normalized `keyword` sub-field for filtering and faceting. ## When you do not need one Elasticsearch's term-level queries accept a `case_insensitive` option, which handles ad-hoc case-insensitive matching against a plain keyword field without changing the mapping. It is convenient for one-off queries, but it does nothing for sorting or aggregation bucketing, and it cannot exploit the index the way a pre-normalized term lookup can. When case-insensitivity is a permanent property of the field, the normalizer is the right mechanism; the query option is for exceptions.
- Which token filters are and are not allowed inside a normalizer?Allowed: filters that work character by character and cannot change the token count — `lowercase`, `uppercase`, `asciifolding`, `trim`, `elision`, `decimal_digit`, the per-language normalization filters, plus character filters such as `mapping` and `pattern_replace`. Rejected at index creation: anything that splits, merges or removes tokens, such as stemmers, `stop`, `shingle` and `ngram`.
- What happens to a terms aggregation on a field after you add a lowercase normalizer?Buckets collapse: values differing only by case merge into one bucket, and the bucket key is the normalized (lowercased) string. That is usually the intent, but it means the field can no longer supply display casing. Keep the raw value in a separate sub-field if the UI needs to show it, and remember existing documents keep their old terms until reindexed.
- When would you use a term query's case_insensitive option instead of a normalizer?For occasional, ad-hoc case-insensitive lookups where changing the mapping is not worth it — an admin search, a one-off report. It does not affect sorting or aggregation bucketing and does not benefit the whole workload. If case-insensitivity is a permanent property of the field, normalize at index time instead.
saying these in an interview costs you the question
- Adds analyzer to a keyword field and expects it to take effect
- Thinks the normalizer applies only at index time, not to queries
- Confuses the keyword analyzer with the keyword field type
- Believes a normalizer can tokenize or stem values
- Expects existing documents to be renormalized without reindexing