skip to content

In Elasticsearch, what does a match query do with the query string before it looks up terms?

level: juniorimportance: must knowfreq 82%

answer

  1. the string is not taken literally
  2. the document went through the same treatment
  3. analyzer first, then term lookup
  4. one clause per resulting token
  5. or is the default combinator

basics

~10 s

A match query runs the string through the field's search analyzer, turning it into terms, then builds a boolean query over those terms. By default a document matching any single term is a hit.

solid answer

~50 s

`match` is the standard full-text query. It passes the query string through the analyzer configured for the target field (the search analyzer, if the mapping defines one), and gets back a list of normalized tokens — lowercased, stemmed, stopword-filtered, whatever that chain does. It then builds a boolean query with one term lookup per token. The default `operator` is `or`, so a document containing any one of the tokens matches, and documents containing more of them score higher; `operator: "and"` requires all of them, and `minimum_should_match` sits between the two. Because the query text goes through the same kind of analysis the document went through at index time, `match` on a `text` field is naturally case- and morphology-insensitive. If analysis strips every token — for example the string is entirely stopwords — the default `zero_terms_query` of `none` means the query matches nothing.

code

json · 10 lines
json
{
  "query": {
    "match": {
      "title": {
        "query": "Quick Brown Foxes",
        "operator": "or"
      }
    }
  }
}

go deeper

for a junior

Be ready to state plainly that the query text is analyzed into tokens first and that any one token matching is enough by default. Knowing you can inspect this with the _analyze API already puts you ahead.

for a middle

Explain the mechanics: which analyzer is chosen, what boolean query is built from the tokens, and how operator and minimum_should_match change the required overlap. Expect to be asked why the same string behaves differently on text and keyword fields.

for a senior

Show you debug with terms, not guesses — reach for _analyze and _validate/query when recall looks wrong, and connect a sudden drop in matches to an analyzer change made without a reindex.

for a principal

Own the analysis contract across the system: which fields use which chains, how analyzer changes are rolled out via reindex-and-alias, and how much query-side complexity you are willing to trade for index-time normalization.

## What "full text" means here Elasticsearch splits its query DSL into two families. Term-level queries look for a byte-for-byte term in the inverted index. Full-text queries — `match`, `match_phrase`, `multi_match`, `match_bool_prefix`, `query_string` and friends — first *analyze* the string you supply, and only then look terms up. That single step is the whole difference, and it is the reason a search for `Running Shoes` finds a document that literally says `run shoe`. ## The analysis step Every `text` field has an analyzer: a character-filter chain, one tokenizer, and a token-filter chain. At index time the field's value went through it, and the resulting tokens were written into the inverted index. At search time a `match` query sends your query string through the analyzer of that same field — the `search_analyzer` if the mapping sets one, otherwise the field's `analyzer`, otherwise the index default, otherwise `standard`. With the `standard` analyzer, `"Quick Brown Foxes"` becomes `[quick, brown, foxes]`; add the `english` analyzer and it becomes `[quick, brown, fox]`. Nothing in the query is matched literally: punctuation may vanish, case is normalised, stopwords may be dropped, and stems may replace whole words. You can see exactly what happens with the `_analyze` API, which is the first thing to reach for when a query surprises you. ## The boolean query it builds After analysis, `match` is a thin wrapper: it constructs a boolean query with one clause per token. With the default `operator: "or"` every clause is a `should`, so one hit is enough to return the document, and each additional matching term adds to the score — that is why full-text search degrades gracefully instead of returning nothing. `operator: "and"` promotes all clauses to required. `minimum_should_match` lets you demand a fraction (`"75%"`) or a count (`"2"`) instead of all-or-one. A few other parameters commonly come up: - `analyzer` overrides which analyzer processes the query string. - `fuzziness`, `prefix_length` and `max_expansions` turn each token into a fuzzy lookup. - `zero_terms_query` (`none` by default, or `all`) decides what happens when analysis leaves nothing. - `lenient` suppresses format errors, for example a non-numeric string against a numeric field. - `auto_generate_synonyms_phrase_query` controls how multi-word synonyms are turned into phrase clauses. ## Why the field type changes everything On a `text` field the value was tokenized, so a `match` for one word finds it inside a sentence. On a `keyword` field the value was indexed whole and untokenized; a `match` query still goes through the (usually empty) normalizer, but the resulting single term has to equal the entire field value. So `match` on a `keyword` field behaves like an exact-value query, not a full-text one. This is why the standard mapping pattern is a `text` field with a `.keyword` multi-field: search the `text` side, sort and aggregate and filter on the `keyword` side. The mirror-image failure is running an unanalyzed lookup against a `text` field: the raw string `"Quick Brown"` was never indexed as one term, so it finds nothing even though the document plainly contains those words. Reasoning about that is the same reasoning as here — always ask what terms are in the index and what terms the query produces. ## When the two sides disagree Because matching is term-to-term, index-time and search-time analysis have to agree. Change a field's analyzer without reindexing and the existing documents keep their old tokens while queries generate new ones; the overlap silently shrinks. Deliberate asymmetry is sometimes useful — indexing with an edge-ngram filter and searching with a plain analyzer is the classic autocomplete recipe — but it must be a decision, not an accident. ## Empty analysis results If every token is removed, `match` has nothing to look up. The default `zero_terms_query: "none"` returns no documents, which surprises people whose stopword list swallowed the query `"to be or not to be"`. Setting `zero_terms_query: "all"` turns that case into a `match_all` instead. ## Interview traps Weak answers describe `match` as a substring or `LIKE` scan, or claim it compares against `_source`. `_source` is the stored original document returned in hits; it plays no part in matching. Matching happens entirely against the analyzed terms in the inverted index, and `match` is simply the query that promises to analyze your input the same way.

  • How would you confirm which terms Elasticsearch actually produced from a query string?
    Call the `_analyze` API with the field name and the text — it returns the exact token stream the field's analyzer produces, which you can compare against what you expect to be in the index. For the whole query, `_validate/query?rewrite=true` and the `_explain` API show the rewritten Lucene query, including the terms each clause is looking for.
  • What goes wrong if the index-time and search-time analyzers produce different tokens?
    Matching is term-to-term, so any mismatch silently reduces recall: the documents hold one set of tokens and the query generates another, and only the accidental overlap matches. Changing a field's analyzer therefore requires reindexing. Deliberate asymmetry — index-time edge ngrams with a plain search analyzer, for example — is fine, but it has to be an intentional choice.
  • Does a match query work against a keyword field?
    It runs, but a `keyword` field is not tokenized: the value was indexed as one whole term, so after any normalizer the entire query string has to equal the entire field value. In practice it behaves like an exact-value lookup. The usual mapping keeps a `text` field for full-text search and a `.keyword` multi-field for exact filters, sorting and aggregations.

saying these in an interview costs you the question

  • Describes match as a substring or LIKE-style scan of the field
  • Thinks match requires the whole string to appear as a phrase
  • Believes matching compares the query against the stored _source
  • Cannot say which analyzer processes the query text
  • Assumes match on a keyword field behaves like full-text search

context