In an Elasticsearch match query, how do operator and minimum_should_match change which documents match?
answer
- both decide how many terms are required
- one is all-or-nothing, one is a dial
- or and and are the two extremes
- percentages round down
- counts apply to analyzed tokens, not typed words
basics
~10 sBoth control how many of the analyzed terms a document must contain. operator switches between any term (or) and every term (and); minimum_should_match sets a count or percentage in between, trading recall for precision.
solid answer
~40 sA `match` query analyzes the input into tokens and turns each one into a clause of a boolean query. `operator` decides how those clauses combine: `"or"` (the default) makes them all optional, so one hit returns the document; `"and"` makes them all required. That is a blunt choice — `or` on a five-word query returns almost everything, `and` returns nothing as soon as one word is missing or was stemmed differently. `minimum_should_match` is the dial between them: `"3"` requires three tokens, `"75%"` requires three quarters of them rounded down, `"-1"` requires all but one. It also accepts conditional forms such as `"2<75%"`, meaning "if there are at most two tokens require all of them, otherwise require 75%". With `operator: "and"` every clause is already required, so `minimum_should_match` has nothing left to relax.
code
json · 10 lines{
"query": {
"match": {
"description": {
"query": "waterproof hiking boots for winter",
"minimum_should_match": "2<75%"
}
}
}
}go deeper
Know that a match query defaults to requiring only one of the analyzed terms, and that operator: "and" requires them all. That is enough at this stage.
Explain that both parameters gate the boolean clauses, walk through the integer, percentage and negative forms of minimum_should_match, and note that percentages round down and count analyzed tokens.
Show judgment about the recall–precision trade-off: pick a value from query-log evidence, watch the zero-result rate as well as click depth, and describe a tiered fallback rather than one global setting.
Own the policy across query classes — navigational versus exploratory queries deserve different strictness — and define how a change to it is measured and rolled back before it ships to a whole search estate.
## Where these parameters act A `match` query is a two-stage thing: analyze the string into tokens, then build a boolean query with one term clause per token. `operator` and `minimum_should_match` both act on that second stage. Neither changes the analysis, and neither has anything to say about scoring — a document either clears the bar or it is not a hit at all. ## operator `operator` takes `"or"` (default) or `"and"`. With `or`, every token becomes an optional (`should`) clause and a single match is enough. With `and`, every token becomes a required clause. The two ends of the dial have opposite failure modes. `or` maximises recall and, for a long query, effectively stops filtering: a five-word search returns every document containing any one of the five common words, and you are relying entirely on ranking to keep the junk off page one. `and` maximises precision and is brittle: one typo, one stopword removed on one side, one word the stemmer treated differently, and a perfectly relevant document disappears. Product search teams discover this the day a user searches for a five-word product title and gets zero results. ## minimum_should_match `minimum_should_match` applies to the optional clauses and states how many of them must match. The accepted forms: - **Positive integer** (`"3"`) — at least this many clauses must match. If there are fewer clauses than the number, all of them are required. - **Negative integer** (`"-1"`) — all but this many may be missing. - **Percentage** (`"75%"`) — this share of the clauses, rounded down, must match. - **Negative percentage** (`"-25%"`) — this share may be missing. - **Combination** (`"2<75%"`) — if the number of clauses is at most 2, all are required; otherwise 75% applies. Several conditions can be chained, for example `"2<-25% 9<-3"`. The rounding matters in interviews: four tokens with `"75%"` requires three; five tokens with `"75%"` requires three as well, because 3.75 rounds down. The conditional form exists because a fixed percentage behaves badly at both ends of the query-length range. Requiring 75% of a two-word query means requiring both words — usually right. Requiring 75% of a twelve-word query means requiring nine, which is far too strict for conversational input. `"3<75%"` says "short queries must match completely, long ones only mostly". ## Interaction between the two Setting `operator: "and"` already promotes every clause to required, so a `minimum_should_match` alongside it has no optional clauses left to govern. Express the intent one way: use `operator` for the two extremes, `minimum_should_match` for anything in between. Note also that `operator: "and"` and `minimum_should_match: "100%"` describe the same requirement. ## Interaction with analysis The count that percentages apply to is the number of tokens *after* analysis, not the number of words the user typed. If the analyzer removes stopwords, `"the best of the coffee makers"` may reduce to three tokens, and `"75%"` then requires two of those three rather than five of seven. Synonym expansion can move the number the other way. This is why a relevance change that looks like a scoring problem is often a `minimum_should_match` interacting with an analyzer change. ## Choosing a setting in production There is no universally correct value; the honest answer in an interview is that you measure. A common starting point for a consumer search box is a conditional `minimum_should_match` that requires everything for one- and two-word queries and relaxes for longer ones, then judge the result against real query logs — checking both the zero-result rate (too strict) and the click depth or abandonment rate (too loose). Another pattern is a tiered strategy: run the strict interpretation first and fall back to a looser one only when the strict one returns too few hits, so you never trade precision away on queries that did not need it. ## Common misconceptions The most frequent one is that `minimum_should_match` boosts documents that match more terms. It does not: it is a hard threshold. Documents that match more terms do tend to score higher, but that comes from the scoring model summing the per-clause contributions, not from this parameter. The second most frequent is treating percentages as counts of the user's words rather than of the analyzed tokens.
- What does minimum_should_match set to "3<75%" mean?If the query analyzes to three tokens or fewer, all of them are required; above three, 75% of them (rounded down) must match. The conditional form exists because one fixed percentage suits neither short nor long queries: two-word queries usually need both words, while a ten-word conversational query would return nothing if you demanded eight of them.
- Why does a percentage sometimes require fewer terms than the user typed?The percentage applies to the clauses produced after analysis, not to the raw words. Stopword removal, token filters that drop or merge tokens, and multi-word synonym handling can all change the count. A seven-word query reduced to three tokens with "75%" requires two of three — which is why relevance shifts often trace back to an analyzer change rather than the query.
- How would you avoid returning zero results for a long query without loosening short ones?Use a conditional minimum_should_match such as "2<75%" so short queries stay strict, or run a tiered strategy: execute the strict interpretation, and only if it returns too few hits re-run a relaxed one. Both keep precision on the queries that can afford it and buy recall only where the strict form fails.
saying these in an interview costs you the question
- Says minimum_should_match boosts documents matching more terms
- Thinks percentages apply to the words the user typed
- Sets operator and minimum_should_match together expecting both to bite
- Believes operator: and is always the safer default
- Cannot state that or is the default combinator