skip to content

Full-text Queries

match, match_phrase, and multi_match run the query string through the field's analyzer before looking up terms, which is what makes them full text. Interviewers ask about multi_match types and minimum_should_match because those knobs are what actually decide recall.

part ofElasticsearchoverview, primer and where to startread it →
on this pageshow

questions

6

In Elasticsearch, what does a match query do with the query string before it looks up terms?

level: juniorimportance: must knowfreq 82%

answer

  1. the string is not taken literally
  2. the document went through the same treatment
  3. analyzer first, then term lookup
  4. one clause per resulting token
  5. or is the default combinator

basics

~10 s

A match query runs the string through the field's search analyzer, turning it into terms, then builds a boolean query over those terms. By default a document matching any single term is a hit.

solid answer

~50 s

`match` is the standard full-text query. It passes the query string through the analyzer configured for the target field (the search analyzer, if the mapping defines one), and gets back a list of normalized tokens — lowercased, stemmed, stopword-filtered, whatever that chain does. It then builds a boolean query with one term lookup per token. The default `operator` is `or`, so a document containing any one of the tokens matches, and documents containing more of them score higher; `operator: "and"` requires all of them, and `minimum_should_match` sits between the two. Because the query text goes through the same kind of analysis the document went through at index time, `match` on a `text` field is naturally case- and morphology-insensitive. If analysis strips every token — for example the string is entirely stopwords — the default `zero_terms_query` of `none` means the query matches nothing.

code

json · 10 lines
json
{
  "query": {
    "match": {
      "title": {
        "query": "Quick Brown Foxes",
        "operator": "or"
      }
    }
  }
}

go deeper

for a junior

Be ready to state plainly that the query text is analyzed into tokens first and that any one token matching is enough by default. Knowing you can inspect this with the _analyze API already puts you ahead.

for a middle

Explain the mechanics: which analyzer is chosen, what boolean query is built from the tokens, and how operator and minimum_should_match change the required overlap. Expect to be asked why the same string behaves differently on text and keyword fields.

for a senior

Show you debug with terms, not guesses — reach for _analyze and _validate/query when recall looks wrong, and connect a sudden drop in matches to an analyzer change made without a reindex.

for a principal

Own the analysis contract across the system: which fields use which chains, how analyzer changes are rolled out via reindex-and-alias, and how much query-side complexity you are willing to trade for index-time normalization.

## What "full text" means here Elasticsearch splits its query DSL into two families. Term-level queries look for a byte-for-byte term in the inverted index. Full-text queries — `match`, `match_phrase`, `multi_match`, `match_bool_prefix`, `query_string` and friends — first *analyze* the string you supply, and only then look terms up. That single step is the whole difference, and it is the reason a search for `Running Shoes` finds a document that literally says `run shoe`. ## The analysis step Every `text` field has an analyzer: a character-filter chain, one tokenizer, and a token-filter chain. At index time the field's value went through it, and the resulting tokens were written into the inverted index. At search time a `match` query sends your query string through the analyzer of that same field — the `search_analyzer` if the mapping sets one, otherwise the field's `analyzer`, otherwise the index default, otherwise `standard`. With the `standard` analyzer, `"Quick Brown Foxes"` becomes `[quick, brown, foxes]`; add the `english` analyzer and it becomes `[quick, brown, fox]`. Nothing in the query is matched literally: punctuation may vanish, case is normalised, stopwords may be dropped, and stems may replace whole words. You can see exactly what happens with the `_analyze` API, which is the first thing to reach for when a query surprises you. ## The boolean query it builds After analysis, `match` is a thin wrapper: it constructs a boolean query with one clause per token. With the default `operator: "or"` every clause is a `should`, so one hit is enough to return the document, and each additional matching term adds to the score — that is why full-text search degrades gracefully instead of returning nothing. `operator: "and"` promotes all clauses to required. `minimum_should_match` lets you demand a fraction (`"75%"`) or a count (`"2"`) instead of all-or-one. A few other parameters commonly come up: - `analyzer` overrides which analyzer processes the query string. - `fuzziness`, `prefix_length` and `max_expansions` turn each token into a fuzzy lookup. - `zero_terms_query` (`none` by default, or `all`) decides what happens when analysis leaves nothing. - `lenient` suppresses format errors, for example a non-numeric string against a numeric field. - `auto_generate_synonyms_phrase_query` controls how multi-word synonyms are turned into phrase clauses. ## Why the field type changes everything On a `text` field the value was tokenized, so a `match` for one word finds it inside a sentence. On a `keyword` field the value was indexed whole and untokenized; a `match` query still goes through the (usually empty) normalizer, but the resulting single term has to equal the entire field value. So `match` on a `keyword` field behaves like an exact-value query, not a full-text one. This is why the standard mapping pattern is a `text` field with a `.keyword` multi-field: search the `text` side, sort and aggregate and filter on the `keyword` side. The mirror-image failure is running an unanalyzed lookup against a `text` field: the raw string `"Quick Brown"` was never indexed as one term, so it finds nothing even though the document plainly contains those words. Reasoning about that is the same reasoning as here — always ask what terms are in the index and what terms the query produces. ## When the two sides disagree Because matching is term-to-term, index-time and search-time analysis have to agree. Change a field's analyzer without reindexing and the existing documents keep their old tokens while queries generate new ones; the overlap silently shrinks. Deliberate asymmetry is sometimes useful — indexing with an edge-ngram filter and searching with a plain analyzer is the classic autocomplete recipe — but it must be a decision, not an accident. ## Empty analysis results If every token is removed, `match` has nothing to look up. The default `zero_terms_query: "none"` returns no documents, which surprises people whose stopword list swallowed the query `"to be or not to be"`. Setting `zero_terms_query: "all"` turns that case into a `match_all` instead. ## Interview traps Weak answers describe `match` as a substring or `LIKE` scan, or claim it compares against `_source`. `_source` is the stored original document returned in hits; it plays no part in matching. Matching happens entirely against the analyzed terms in the inverted index, and `match` is simply the query that promises to analyze your input the same way.

  • How would you confirm which terms Elasticsearch actually produced from a query string?
    Call the `_analyze` API with the field name and the text — it returns the exact token stream the field's analyzer produces, which you can compare against what you expect to be in the index. For the whole query, `_validate/query?rewrite=true` and the `_explain` API show the rewritten Lucene query, including the terms each clause is looking for.
  • What goes wrong if the index-time and search-time analyzers produce different tokens?
    Matching is term-to-term, so any mismatch silently reduces recall: the documents hold one set of tokens and the query generates another, and only the accidental overlap matches. Changing a field's analyzer therefore requires reindexing. Deliberate asymmetry — index-time edge ngrams with a plain search analyzer, for example — is fine, but it has to be an intentional choice.
  • Does a match query work against a keyword field?
    It runs, but a `keyword` field is not tokenized: the value was indexed as one whole term, so after any normalizer the entire query string has to equal the entire field value. In practice it behaves like an exact-value lookup. The usual mapping keeps a `text` field for full-text search and a `.keyword` multi-field for exact filters, sorting and aggregations.

saying these in an interview costs you the question

  • Describes match as a substring or LIKE-style scan of the field
  • Thinks match requires the whole string to appear as a phrase
  • Believes matching compares the query against the stored _source
  • Cannot say which analyzer processes the query text
  • Assumes match on a keyword field behaves like full-text search

context

open as a page

How do the multi_match types best_fields, most_fields and cross_fields differ in Elasticsearch?

level: seniorimportance: must knowfreq 64%

basics

~20 s

best_fields takes the best-scoring single field, so all query terms should be in one field. most_fields sums the scores of several analyses of the same text. cross_fields is term-centric: it treats the listed fields as one big field so terms may be spread across them.

open as a page

How does the fuzziness parameter behave in an Elasticsearch match query, and what does AUTO mean?

level: middleimportance: should knowfreq 55%

basics

~20 s

fuzziness allows each analyzed term to match index terms within an edit distance, capped at 2 edits. AUTO scales with term length: no edits for terms of 1-2 characters, one edit for 3-5, two edits above that.

open as a page

What must a document contain to match an Elasticsearch match_phrase query, and what does slop change?

level: middleimportance: should knowfreq 70%

basics

~20 s

match_phrase requires all analyzed terms in the same field, in the given order, at consecutive positions. slop is the number of position moves allowed, so a non-zero slop tolerates intervening words and, at a cost of two moves, reversed pairs.

open as a page

In an Elasticsearch match query, how do operator and minimum_should_match change which documents match?

level: middleimportance: should knowfreq 66%

basics

~10 s

Both control how many of the analyzed terms a document must contain. operator switches between any term (or) and every term (and); minimum_should_match sets a count or percentage in between, trading recall for precision.

open as a page

When would you expose query_string versus simple_query_string to end users in Elasticsearch?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Prefer simple_query_string for anything user-facing: it never fails on bad syntax and its operator set can be restricted. Reserve query_string, which parses full Lucene syntax and returns an error on malformed input, for trusted power users.

open as a page