What must a document contain to match an Elasticsearch match_phrase query, and what does slop change?
answer
- order and adjacency, not just presence
- the index stores where each term occurred
- zero tolerance by default
- a number of position moves, not words
- swapping two terms costs two of them
basics
~20 smatch_phrase requires all analyzed terms in the same field, in the given order, at consecutive positions. slop is the number of position moves allowed, so a non-zero slop tolerates intervening words and, at a cost of two moves, reversed pairs.
solid answer
~50 s`match_phrase` analyzes the query string exactly like `match` does, but instead of independent term lookups it uses the **positions** recorded in the inverted index. All the resulting terms must appear in the same field, in order, with no gaps — position 5, 6, 7. `slop` relaxes that: it is the maximum total number of position moves you may apply to the query terms to line them up with the document. `slop: 1` lets `"quick fox"` match the text `quick brown fox`, because `fox` moves one position. Swapping two adjacent terms costs two moves, so `slop: 2` makes `"fox quick"` match `quick fox`. The default `slop` is 0. Positional matching is only possible if the field indexed positions, which `text` fields do by default; it also costs more than a plain `match`, since the engine must read and intersect position lists.
code
json · 10 lines{
"query": {
"match_phrase": {
"content": {
"query": "quick fox",
"slop": 1
}
}
}
}go deeper
Recall that match_phrase demands the words together and in order, while a plain match does not, and that the default allows no gap at all.
Explain that matching uses the positions stored in the index, define slop as a budget of position moves, and work out why an adjacent swap costs two of them.
Bring the production shape: phrase clauses boosted inside a bool rather than used alone, awareness that position lists make these queries expensive, and the mapping conditions that make them possible at all.
Weigh phrase precision against index cost and latency across the estate — index_phrases and shingle strategies, where proximity truly earns its keep, and when the requirement really calls for the intervals query instead.
## What the query actually requires `match_phrase` starts the same way `match` does: the query string goes through the field's analyzer and comes out as an ordered token stream. The difference is what is done with it. Where `match` issues independent term lookups, `match_phrase` builds a Lucene phrase query, which consults the **term positions** stored in the postings list. A document matches only if every term appears in that field with positions forming the requested sequence. So `"quick brown fox"` matches a document whose `title` contains those three tokens at positions *n*, *n+1*, *n+2*. It does not match `the fox is quick and brown`, and it does not match a document where `quick` appears in the title and `brown fox` in the body — positions are per field. ## slop `slop` (default `0`) is the budget of position moves you allow. Think of the query terms as counters you may slide along the document's position axis; `slop` is the maximum total distance you may slide them to make the sequence fit. Two consequences follow, and interviewers ask about both: - **Gaps cost one move per position.** Query `"quick fox"` against text `quick brown fox` needs `fox` shifted one position, so `slop: 1` suffices. - **Reordering costs two moves per swap.** Query `"fox quick"` against text `quick fox` requires swapping adjacent terms, which is two moves, so you need `slop: 2`. That is why people describe non-zero slop as "proximity search": it stops being a strict phrase and becomes "these words, near each other, preferably in order". Note that slop does not switch order off — it just makes out-of-order matches expensive enough that they only pass at higher values, and Lucene scores sloppier matches lower than exact ones, so a document with the tight phrase still outranks one that only barely fit within the budget. ## Analysis still applies Everything that happens to a `match` query happens here. Stemming applies, so `"running shoes"` can match indexed `run shoe`. Case is normalized. And stopword handling can be surprising: a stop filter removes tokens but leaves a position gap where they were, so phrases spanning a stopword often still line up — but if the query side and index side disagree about which words are stopwords, phrase matching quietly breaks. When phrase behaviour looks wrong, run `_analyze` on both sides before touching `slop`. ## Cost and mapping requirements Positional matching needs positions in the index. `text` fields index positions by default, but a field mapped with `index_options` set to `docs` or `freqs` has thrown them away and phrase queries on it will fail or find nothing. Phrase queries are also more expensive than term queries: the engine must load position lists for every candidate document and intersect them, and a high `slop` widens the window it has to consider. Frequent, latency-sensitive phrase queries on very common terms are one of the classic slow-query patterns. The `text` field mapping offers `index_phrases`, which additionally indexes two-word shingles so that exact two-term phrases can be answered without walking positions — faster queries at the price of a bigger index. It only helps exact phrases, not sloppy ones. ## Where it fits in a real query In practice `match_phrase` is rarely the whole query, because on its own it is brittle: one missing word returns nothing. The common production shape is a `bool` whose `should` clauses combine a permissive `match` (which decides *which* documents come back) with a boosted `match_phrase` on the same text (which decides which of them rise to the top). Users who type a phrase get the phrase-matching documents first, and users who type keywords still get results. Also worth knowing: `fuzziness` is not supported on `match_phrase`. If you need typo tolerance and adjacency at once, you are combining clauses or moving to the `intervals` query. ## When slop is too blunt `slop` gives you one number for the whole phrase. The `intervals` query gives per-rule control: `match` rules with their own `max_gaps` and `ordered` flags, combined with `all_of` and `any_of`, and filtered with `containing`, `contained_by`, `not_containing` and similar. That is the tool for requirements like "these two terms within three words of each other, in order, but not if a third term sits between them" — the kind of precision legal and medical search needs and `slop` cannot express.
- Why does slop 1 not let "fox quick" match a document containing "quick fox"?Slop counts position moves, and swapping two adjacent terms takes two of them: one term has to move past the other. A gap of one intervening word costs only one move, which is why slop: 1 handles "quick fox" against "quick brown fox" but not a reversal. Reordering an adjacent pair needs slop: 2.
- The user is still typing the last word of the phrase — which query would you use?`match_bool_prefix` analyzes the input and treats every term but the last as an ordinary term clause, with the final one as a prefix query, so word order is not required — good for a general search-as-you-type box. `match_phrase_prefix` keeps phrase semantics and only prefixes the last term; it expands that prefix to a limited number of terms, which can cost recall on common prefixes.
- What would you use when slop cannot express the proximity rule you need?The `intervals` query. It composes `match`, `prefix`, `wildcard` and `fuzzy` rules with `all_of` and `any_of`, each carrying its own `max_gaps` and `ordered` settings, and it supports filters such as `containing`, `contained_by` and `not_containing`. That lets you say "A within three words of B, in order, unless C appears between them", which a single slop number cannot.
- How do you keep phrase matching from returning zero results when one word is missing?Don't let the phrase decide recall. Put a permissive `match` in a bool `should` to determine which documents qualify, and add a boosted `match_phrase` on the same text as another `should` so exact phrases rank first. Users who type a phrase still get it on top, and keyword-style queries still return something.
saying these in an interview costs you the question
- Thinks slop is a count of words allowed between terms
- Says reversing two adjacent terms works with slop 1
- Believes match_phrase can span two different fields
- Assumes fuzziness works on match_phrase
- Forgets that the phrase is analyzed and expects a literal match