skip to content

What must a document contain to match an Elasticsearch match_phrase query, and what does slop change?

level: middleimportance: should knowfreq 70%

answer

  1. order and adjacency, not just presence
  2. the index stores where each term occurred
  3. zero tolerance by default
  4. a number of position moves, not words
  5. swapping two terms costs two of them

basics

~20 s

match_phrase requires all analyzed terms in the same field, in the given order, at consecutive positions. slop is the number of position moves allowed, so a non-zero slop tolerates intervening words and, at a cost of two moves, reversed pairs.

solid answer

~50 s

`match_phrase` analyzes the query string exactly like `match` does, but instead of independent term lookups it uses the **positions** recorded in the inverted index. All the resulting terms must appear in the same field, in order, with no gaps — position 5, 6, 7. `slop` relaxes that: it is the maximum total number of position moves you may apply to the query terms to line them up with the document. `slop: 1` lets `"quick fox"` match the text `quick brown fox`, because `fox` moves one position. Swapping two adjacent terms costs two moves, so `slop: 2` makes `"fox quick"` match `quick fox`. The default `slop` is 0. Positional matching is only possible if the field indexed positions, which `text` fields do by default; it also costs more than a plain `match`, since the engine must read and intersect position lists.

code

json · 10 lines
json
{
  "query": {
    "match_phrase": {
      "content": {
        "query": "quick fox",
        "slop": 1
      }
    }
  }
}

go deeper

for a junior

Recall that match_phrase demands the words together and in order, while a plain match does not, and that the default allows no gap at all.

for a middle

Explain that matching uses the positions stored in the index, define slop as a budget of position moves, and work out why an adjacent swap costs two of them.

for a senior

Bring the production shape: phrase clauses boosted inside a bool rather than used alone, awareness that position lists make these queries expensive, and the mapping conditions that make them possible at all.

for a principal

Weigh phrase precision against index cost and latency across the estate — index_phrases and shingle strategies, where proximity truly earns its keep, and when the requirement really calls for the intervals query instead.

## What the query actually requires `match_phrase` starts the same way `match` does: the query string goes through the field's analyzer and comes out as an ordered token stream. The difference is what is done with it. Where `match` issues independent term lookups, `match_phrase` builds a Lucene phrase query, which consults the **term positions** stored in the postings list. A document matches only if every term appears in that field with positions forming the requested sequence. So `"quick brown fox"` matches a document whose `title` contains those three tokens at positions *n*, *n+1*, *n+2*. It does not match `the fox is quick and brown`, and it does not match a document where `quick` appears in the title and `brown fox` in the body — positions are per field. ## slop `slop` (default `0`) is the budget of position moves you allow. Think of the query terms as counters you may slide along the document's position axis; `slop` is the maximum total distance you may slide them to make the sequence fit. Two consequences follow, and interviewers ask about both: - **Gaps cost one move per position.** Query `"quick fox"` against text `quick brown fox` needs `fox` shifted one position, so `slop: 1` suffices. - **Reordering costs two moves per swap.** Query `"fox quick"` against text `quick fox` requires swapping adjacent terms, which is two moves, so you need `slop: 2`. That is why people describe non-zero slop as "proximity search": it stops being a strict phrase and becomes "these words, near each other, preferably in order". Note that slop does not switch order off — it just makes out-of-order matches expensive enough that they only pass at higher values, and Lucene scores sloppier matches lower than exact ones, so a document with the tight phrase still outranks one that only barely fit within the budget. ## Analysis still applies Everything that happens to a `match` query happens here. Stemming applies, so `"running shoes"` can match indexed `run shoe`. Case is normalized. And stopword handling can be surprising: a stop filter removes tokens but leaves a position gap where they were, so phrases spanning a stopword often still line up — but if the query side and index side disagree about which words are stopwords, phrase matching quietly breaks. When phrase behaviour looks wrong, run `_analyze` on both sides before touching `slop`. ## Cost and mapping requirements Positional matching needs positions in the index. `text` fields index positions by default, but a field mapped with `index_options` set to `docs` or `freqs` has thrown them away and phrase queries on it will fail or find nothing. Phrase queries are also more expensive than term queries: the engine must load position lists for every candidate document and intersect them, and a high `slop` widens the window it has to consider. Frequent, latency-sensitive phrase queries on very common terms are one of the classic slow-query patterns. The `text` field mapping offers `index_phrases`, which additionally indexes two-word shingles so that exact two-term phrases can be answered without walking positions — faster queries at the price of a bigger index. It only helps exact phrases, not sloppy ones. ## Where it fits in a real query In practice `match_phrase` is rarely the whole query, because on its own it is brittle: one missing word returns nothing. The common production shape is a `bool` whose `should` clauses combine a permissive `match` (which decides *which* documents come back) with a boosted `match_phrase` on the same text (which decides which of them rise to the top). Users who type a phrase get the phrase-matching documents first, and users who type keywords still get results. Also worth knowing: `fuzziness` is not supported on `match_phrase`. If you need typo tolerance and adjacency at once, you are combining clauses or moving to the `intervals` query. ## When slop is too blunt `slop` gives you one number for the whole phrase. The `intervals` query gives per-rule control: `match` rules with their own `max_gaps` and `ordered` flags, combined with `all_of` and `any_of`, and filtered with `containing`, `contained_by`, `not_containing` and similar. That is the tool for requirements like "these two terms within three words of each other, in order, but not if a third term sits between them" — the kind of precision legal and medical search needs and `slop` cannot express.

  • Why does slop 1 not let "fox quick" match a document containing "quick fox"?
    Slop counts position moves, and swapping two adjacent terms takes two of them: one term has to move past the other. A gap of one intervening word costs only one move, which is why slop: 1 handles "quick fox" against "quick brown fox" but not a reversal. Reordering an adjacent pair needs slop: 2.
  • The user is still typing the last word of the phrase — which query would you use?
    `match_bool_prefix` analyzes the input and treats every term but the last as an ordinary term clause, with the final one as a prefix query, so word order is not required — good for a general search-as-you-type box. `match_phrase_prefix` keeps phrase semantics and only prefixes the last term; it expands that prefix to a limited number of terms, which can cost recall on common prefixes.
  • What would you use when slop cannot express the proximity rule you need?
    The `intervals` query. It composes `match`, `prefix`, `wildcard` and `fuzzy` rules with `all_of` and `any_of`, each carrying its own `max_gaps` and `ordered` settings, and it supports filters such as `containing`, `contained_by` and `not_containing`. That lets you say "A within three words of B, in order, unless C appears between them", which a single slop number cannot.
  • How do you keep phrase matching from returning zero results when one word is missing?
    Don't let the phrase decide recall. Put a permissive `match` in a bool `should` to determine which documents qualify, and add a boosted `match_phrase` on the same text as another `should` so exact phrases rank first. Users who type a phrase still get it on top, and keyword-style queries still return something.

saying these in an interview costs you the question

  • Thinks slop is a count of words allowed between terms
  • Says reversing two adjacent terms works with slop 1
  • Believes match_phrase can span two different fields
  • Assumes fuzziness works on match_phrase
  • Forgets that the phrase is analyzed and expects a literal match

context