skip to content

Why does an Elasticsearch term query on a text field usually return no hits?

level: juniorimportance: must knowfreq 88%

answer

  1. Two pipelines that must agree
  2. One side transforms, the other does not
  3. Ask what the index actually stores
  4. Try the _analyze API on the value
  5. The dot-suffix sub-field holds the whole string

basics

~20 s

A term query looks up the exact bytes you supply with no analysis, while a text field stores analyzed tokens that are split and lowercased. Searching for New York finds no such token. Use match, or a keyword sub-field.

solid answer

~40 s

Elasticsearch runs two different pipelines. At index time a `text` field is passed through an analyzer, so `"New York"` is stored as the tokens `new` and `york`. At query time, term-level queries — `term`, `terms`, `prefix`, `wildcard`, `fuzzy` — are **not** analyzed: they look the supplied string up in the term dictionary byte for byte. `{"term":{"title":"New York"}}` therefore asks for a token literally spelled `New York`, which was never indexed, and the query matches nothing. Even a single word fails on case alone (`Apple` vs the indexed `apple`). The fixes are: use `match` (which analyzes the input with the field's search analyzer) for full-text intent, or target the exact-value sub-field — `title.keyword` from the default multi-field — when you really want exact matching, filtering, or sorting.

go deeper

for a junior

Be ready to say that a term query is not analyzed and a text field is, and to name the two fixes: use match, or target the keyword sub-field.

for a middle

Explain the index-time versus query-time pipelines concretely — tokenizer plus lowercase filter — and demonstrate the _analyze API as the diagnostic that settles it in one request.

for a senior

Show the judgment of mapping fields correctly up front: identifiers and enums as keyword, prose as text with a keyword sub-field, and know that ignore_above silently drops long values.

for a principal

Own the convention across teams: a mapping policy and a query-construction layer that keeps exact-match and full-text intents from being mixed up by every new service that indexes data.

## The two pipelines Elasticsearch stores and searches text through two paths that must agree, and this question is about what happens when they do not. **Index time.** When a document with a `text` field arrives, the field's analyzer runs: a character filter chain, a tokenizer, then token filters. The default `standard` analyzer splits on word boundaries and lowercases. The value `"New York City"` becomes three tokens — `new`, `york`, `city` — and those tokens, not the original string, are what land in the inverted index. The original string survives only in `_source`, which is returned to you but is not searchable. **Query time.** Full-text queries (`match`, `match_phrase`, `multi_match`) analyze the query string with the field's search analyzer before looking anything up, so they produce tokens that are directly comparable with what is indexed. Term-level queries do not. `term`, `terms`, `prefix`, `wildcard`, `regexp`, `fuzzy` and `range` take your input as an exact byte sequence and probe the term dictionary with it. ## Why the query silently returns zero `{"term": {"title": "New York City"}}` asks: is there a document containing a single indexed token whose bytes are exactly `New York City`? There is not — the analyzer never produced such a token. Elasticsearch does not warn you; a term that does not exist in the dictionary simply matches no documents, and you get `hits.total: 0` with a `200 OK`. This is what makes the mistake so common: nothing is broken, there is just nothing there. The failure has three flavours, in increasing order of subtlety: - **Case.** `{"term":{"title":"Apple"}}` misses the indexed `apple`. One capital letter is enough. - **Tokenization.** Any multi-word value fails, because the indexed unit is a single word. - **Token filters.** Stemming, stop words, ASCII folding or synonyms mean the indexed token may not resemble the input at all: `running` may be stored as `run`, so a term query for `running` misses even though the word is visibly in the document. And the reverse trap: a term query for a single lowercase word *does* work on a `text` field, which is exactly why teams ship code that appears correct until someone searches a two-word value. ## The fixes **Use `match` when you mean full text.** `match` analyzes the input, produces `new`, `york`, `city`, and by default ORs the resulting terms; `operator: "and"` or `minimum_should_match` tightens it. **Use the keyword sub-field when you mean exact.** Under dynamic mapping a string field becomes a `text` field with a `keyword` sub-field, so `title.keyword` holds the whole untouched string as one term. `{"term":{"title.keyword":"New York City"}}` matches. Note that the dynamically created sub-field carries `ignore_above: 256`, so values longer than 256 characters are not indexed there at all and will never match. **Map the field as `keyword` in the first place** when it is an identifier, status code, tag or enum. Those fields are never analyzed, so `term` behaves exactly as you expect, and they are also what you need for aggregations and sorting. A `match` query against a `keyword` field is safe too — the field's analyzer is a no-op, so `match` and `term` behave the same there apart from any `normalizer` you configured. ## Diagnosing it The fastest confirmation is the `_analyze` API: run your value through the field's analyzer and look at the tokens it produces. If your term query string does not appear verbatim in that token list, it cannot match. `GET /my-index/_analyze` with `{"field":"title","text":"New York City"}` settles the argument in one request. Checking `GET /my-index/_mapping` to see whether the field is `text`, `keyword`, or a multi-field is the other half of the diagnosis. ## Case-insensitive exact matching If you want exact matching on a `keyword` field but do not want to care about case, there are two clean options: set `case_insensitive: true` on the `term` query, or map the `keyword` field with a `normalizer` that lowercases at index time so the stored term is already folded. The normalizer route is preferable when the pattern is systematic, because it keeps the index and the query in agreement rather than paying for case folding on every search.

  • A term query for the single word apple works on a text field, but Apple returns nothing. Why?
    The default `standard` analyzer lowercases, so the indexed token is `apple`. A term query is not analyzed, so `Apple` is looked up byte for byte and misses. The single lowercase word working is a coincidence of the analyzer being a near no-op for that input — it is not evidence that term queries are safe on text fields.
  • How do you make a term query on a keyword field ignore case?
    Either set `case_insensitive: true` on the `term` query, or map the `keyword` field with a lowercasing `normalizer` so the indexed term is already folded. The normalizer is the better default when every search on that field should be case-insensitive; the query flag is for one-off needs.
  • Why does a term query on title.keyword miss a document with a very long title?
    The `keyword` sub-field that dynamic mapping creates sets `ignore_above: 256`. Values longer than 256 characters are skipped entirely for that sub-field — nothing is indexed, so no term query can match, and an `exists` query on it also fails. The value is still present in `_source`, which makes the gap easy to miss.

saying these in an interview costs you the question

  • Claims term queries are analyzed like match queries
  • Says text and keyword fields behave identically for term
  • Blames the empty result on missing data rather than analysis
  • Suggests lowercasing the query string as the general fix
  • Thinks _source contents are what gets searched

context