skip to content

How do the multi_match types best_fields, most_fields and cross_fields differ in Elasticsearch?

level: seniorimportance: must knowfreq 64%

answer

  1. field-centric versus term-centric is the axis
  2. one asks which field is best
  3. one adds up several views of one text
  4. one pretends the fields are concatenated
  5. operator and means different things across them

basics

~20 s

best_fields takes the best-scoring single field, so all query terms should be in one field. most_fields sums the scores of several analyses of the same text. cross_fields is term-centric: it treats the listed fields as one big field so terms may be spread across them.

solid answer

~50 s

`multi_match` runs one query string against several fields, and `type` decides how the per-field evidence is combined. **best_fields** (the default) is field-centric: it scores each field separately and keeps the highest, plus `tie_breaker` times the others. It is right when one field should plausibly contain the whole query — `title` versus `body`. **most_fields** sums the per-field scores; it suits the same text indexed several ways, such as `title`, `title.english` and `title.shingles`, where each analysis is more evidence about the same content. **cross_fields** is term-centric: for every term it looks across the group of fields as if they were one concatenated field, blending term statistics so a term is not scored oddly because it happens to be rare in one small field. That is what you want for `first_name` / `last_name` or split address fields, where no single field holds all the terms. There are also `phrase`, `phrase_prefix` and `bool_prefix` types, which run the corresponding per-field query.

code

json · 10 lines
json
{
  "query": {
    "multi_match": {
      "query": "peter smith",
      "type": "cross_fields",
      "fields": ["first_name", "last_name"],
      "operator": "and"
    }
  }
}

go deeper

for a junior

Know that multi_match runs one string against several fields and that field boosts are written as title^3; the type distinctions come later.

for a middle

Explain how each type combines per-field evidence — max, sum, or a blended term-centric view — and name best_fields as the default.

for a senior

Diagnose with it: recognise the albino elephant symptom, know that operator and minimum_should_match shift meaning under cross_fields, and pick a type from the shape of the data rather than by habit.

for a principal

Own the field strategy — which fields are searchable, how they are analyzed, how boosts are governed, and how a type change is validated against query logs before it ships.

## The problem multi_match solves Real documents scatter searchable text across fields: `title`, `description`, `brand`, `tags`, or `first_name` and `last_name`. `multi_match` takes one query string and one list of fields (boostable as `title^3`, and wildcards like `*_name` are allowed) and runs it against all of them. The interesting part is `type`, which decides *how evidence from several fields becomes one score* — and, less obviously, how `operator` and `minimum_should_match` are interpreted. The fundamental split is **field-centric** versus **term-centric**. ## best_fields — field-centric, winner takes most `best_fields` is the default. It runs a `match` per field and wraps them in a disjunction that takes the **maximum** score, optionally adding `tie_breaker` times the sum of the other fields' scores (`tie_breaker` defaults to 0, meaning the losers contribute nothing). The intuition is: the best single field is the best signal. A document whose `title` reads "Java Concurrency in Practice" should beat one that mentions "Java" in the title and "concurrency" three paragraphs into the body. Use it when the query is likely to be satisfied by one field on its own. Its famous failure is the *albino elephant* problem. With `operator: "and"`, the requirement lands **inside each field**: a document must have every term in *the same* field. So searching `"albino elephant"` will not match a document with `albino` in one field and `elephant` in another, even though a human would call that a match. ## most_fields — field-centric, evidence adds up `most_fields` runs the same per-field `match` queries but **sums** the scores. That makes sense when the fields are not independent content but different *analyses of the same content*: a multi-field with a plain analyzer, a language analyzer, and a shingle or ngram variant. A document matching the exact form and the stemmed form and the shingled form is more confidently relevant, so scores accumulate. Using `most_fields` across genuinely different content fields tends to reward documents that repeat the same words everywhere — a document with the term in `title`, `description` and `tags` outscores one that only has a strong title — which is often not what a product owner means by relevant. ## cross_fields — term-centric `cross_fields` inverts the loop. Instead of "for each field, match all terms", it does "for each term, look across the fields". Conceptually the listed fields are treated as one large concatenated field, and term statistics are blended across them so that a term is not scored strangely just because it is rare in one short field and common in a long one. The practical consequences: - `operator: "and"` and `minimum_should_match` now apply **per term across the whole group**, not inside each field. That solves the albino elephant problem: every term must appear *somewhere* in the group. - Fields are grouped by analyzer. Because the group is treated as one field, fields analyzed differently cannot be blended together, so they form separate groups combined like `best_fields`. If you intend cross-field matching, give the fields the same analyzer. - `fuzziness` is not supported by `cross_fields`. The canonical use case is structured data split across columns: person names, address lines, a product whose brand and model live apart. ## The other types - **phrase** — runs `match_phrase` on each field and combines them `best_fields`-style. - **phrase_prefix** — the same with `match_phrase_prefix`, for search-as-you-type where order matters. - **bool_prefix** — runs `match_bool_prefix` on each field: every term but the last is an ordinary term clause and the last becomes a prefix, so order is not required. ## Choosing in an interview A good answer names the decision rule rather than reciting the list. Ask: *can one field plausibly contain the whole query?* If yes, `best_fields`. *Are the fields different analyses of the same text?* Then `most_fields`. *Is the query naturally spread across structured fields?* Then `cross_fields` with a shared analyzer. And say out loud that these are hypotheses to be measured against real query logs, not settings to be guessed at once and forgotten. A very common production shape is not one `multi_match` at all but a `bool` combining a `cross_fields` or `best_fields` clause for recall with boosted phrase clauses on the important fields for precision — one query decides who is in the result set, the other decides who is at the top. ## Traps Expecting `operator: "and"` to mean the same thing under `best_fields` and `cross_fields` is the single most common misunderstanding, and it produces a bug that looks like missing documents rather than a scoring quirk. Second is mixing analyzers under `cross_fields` and wondering why cross-field matching stopped working.

  • Why does multi_match with best_fields and operator "and" fail to match a document with one term in the title and the other in the body?
    best_fields is field-centric: it builds a separate match per field and takes the highest score, so operator: "and" is enforced inside each field individually. No single field holds both terms, so no clause matches. cross_fields moves the requirement to the term level across the whole field group and matches that document, which is why it exists.
  • What does tie_breaker do in a best_fields multi_match?
    best_fields keeps the highest-scoring field's score; tie_breaker adds that fraction of the other matching fields' scores on top. At the default of 0 the other fields contribute nothing, so a document matching two fields ties with one matching only the best. A small value acts as a gentle bonus for matching in several places without letting field count dominate.
  • Which multi_match type suits a title field indexed three ways — plain, stemmed and shingled?
    most_fields. Those are three analyses of the same text, so matching several of them is corroborating evidence about one piece of content and summing the scores is the right combination. best_fields would ignore the corroboration by keeping only the top one, and cross_fields is inappropriate because the fields are not different content to spread terms across.
  • What breaks if the fields in a cross_fields query use different analyzers?
    cross_fields blends term statistics as if the fields were concatenated, which is only meaningful for fields producing comparable terms. Fields are therefore grouped by analyzer, and separate groups are combined the way best_fields would be. The cross-field requirement holds within a group but not across groups, so terms split between differently analyzed fields stop matching as intended.

saying these in an interview costs you the question

  • Says operator: and works the same under best_fields and cross_fields
  • Treats most_fields as the safe default for unrelated content fields
  • Thinks cross_fields physically concatenates the fields at index time
  • Claims best_fields sums the scores of all matching fields
  • Expects fuzziness to work with cross_fields

context