skip to content

In an Elasticsearch dis_max query, what does the tie_breaker parameter change?

level: middleimportance: should knowfreq 40%

answer

  1. Scoring, not matching, is what changes
  2. One clause decides the score by default
  3. Others can count for a fraction
  4. That fraction defaults to zero
  5. Formula: max plus fraction of the rest

basics

~20 s

dis_max scores a document by its single best-matching sub-query. tie_breaker, default 0.0, adds that fraction of every other matching sub-query's score to the total, so documents matching several clauses edge ahead of equally-good single-clause matches.

solid answer

~50 s

`dis_max` (disjunction max) takes a list of queries under `queries` and returns any document matching at least one of them — but it scores each document by the **single highest-scoring** sub-query rather than by the sum. `tie_breaker`, default `0.0`, controls how much the *other* matching sub-queries count: the final score is `max + tie_breaker * (sum of the remaining matching scores)`. At 0.0, only the best clause matters and documents matching two clauses tie with documents matching one just as well. A small value such as 0.1 or 0.3 breaks that tie in favour of the document matching more clauses without letting the extra clauses dominate. Contrast this with a `bool` `should` array, which simply sums every matching clause — that rewards documents matching many clauses weakly over documents matching one clause strongly, which is wrong when the clauses are alternative views of the same text across different fields.

code

json · 11 lines
json
{
  "query": {
    "dis_max": {
      "queries": [
        { "match": { "title": "quick fox" } },
        { "match": { "body":  "quick fox" } }
      ],
      "tie_breaker": 0.3
    }
  }
}

go deeper

for a junior

Recall that dis_max returns documents matching at least one of its queries and scores them by the best-matching one, and that tie_breaker is the parameter governing the rest.

for a middle

State the formula — maximum plus tie_breaker times the sum of the other matching clauses — with the default of 0.0, and work a two-clause arithmetic example on demand.

for a senior

Justify choosing dis_max over bool should from the structure of the signals, and diagnose a ranking complaint by reading explain output to see whether a second field contributed anything.

for a principal

Own when multi-field search is modelled as competing views versus accumulating evidence, and how tie_breaker is tuned against relevance judgements rather than guessed once and frozen.

## The problem dis_max solves Suppose a user types one phrase and you want to search it across `title`, `body` and `tags`. The obvious construction is a `bool` with three `should` clauses. But `bool` **sums** the scores of every matching should clause, and that produces a specific failure: a document that matches the phrase mediocrely in all three fields can outrank a document that matches it perfectly in `title` alone. Since the three fields are alternative places to find the *same* user intent, not three independent pieces of evidence, summing double-counts. `dis_max` encodes the other assumption: the fields are competing views of one query, so the best view should decide the score. ```json { "dis_max": { "queries": [ { "match": { "title": "quick fox" } }, { "match": { "body": "quick fox" } } ], "tie_breaker": 0.3 }} ``` ## The scoring formula For a document that matches a set of the sub-queries with scores s1..sn: **score = max(s1..sn) + tie_breaker × (sum of all the others)** The default `tie_breaker` is **0.0**, which reduces the formula to the plain maximum. Note that non-matching sub-queries contribute nothing at all — the sum runs only over clauses the document actually matched. Work an example. A document scores 3.0 in `title` and 2.0 in `body`. - `tie_breaker` 0.0 → score 3.0. Identical to a document scoring 3.0 in `title` and nothing in `body`. - `tie_breaker` 0.3 → 3.0 + 0.3 × 2.0 = 3.6. The multi-field match now wins the tie. - A `bool` `should` over the same two clauses → 5.0, letting the secondary field carry a third of the weight. ## Choosing a value `tie_breaker` ranges from 0.0 (pure maximum, `bool` semantics discarded) to 1.0 (equivalent to summing, i.e. `bool` `should` semantics). Small values — commonly in the 0.1 to 0.3 region — are the useful middle: they say "matching a second field is a genuine but secondary signal". There is no universally correct value; it is a relevance parameter you tune against judgements, not a setting with a right answer. ## Matching semantics `dis_max` is a **disjunction**: a document matching any one sub-query is a hit. It has no notion of required clauses and no `minimum_should_match`. If you need mandatory predicates alongside it, wrap the `dis_max` in a `bool` and put the predicates in `filter`. ## Relationship to bool should The two are the same disjunction with different score aggregation: | | matching | scoring | |---|---|---| | `bool` with `should` | at least `minimum_should_match` clauses | sum of matching clause scores | | `dis_max` | at least one query | max + `tie_breaker` × sum of the rest | Use `bool` `should` when the clauses are **independent evidence** that should accumulate — in stock, on promotion, in the preferred brand. Use `dis_max` when they are **alternative expressions of one thing** — the same phrase looked for in several fields. ## Where it shows up implicitly `dis_max` is not only a hand-written query: the `best_fields` behaviour of `multi_match` is built on it, and `tie_breaker` is exposed there for the same reason. So even teams who never type `dis_max` are using its semantics, which is why understanding the formula pays off when someone asks why adding a field to a multi-field search did not change the ranking as expected — with the default `tie_breaker` of 0.0, it may genuinely change nothing for documents whose best field is unchanged. ## Diagnosing with explain Run the query with `"explain": true`. A `dis_max` clause shows the winning sub-query's contribution plus, when `tie_breaker` is non-zero, a separate small contribution attributed to the remaining matches. If you expected a multi-field bonus and see only one contribution, `tie_breaker` is at its default. ## Common mistakes - Assuming `dis_max` requires all sub-queries to match. It requires one. - Assuming the default `tie_breaker` gives a small multi-field bonus. The default is 0.0 — no bonus at all. - Setting `tie_breaker` to 1.0 and being surprised it behaves like a `bool` — that is precisely what 1.0 means. - Reaching for `dis_max` to combine unrelated signals such as recency and stock status, where summing is the correct model.

  • When should you use a bool with should clauses instead of dis_max?
    When the clauses are independent pieces of evidence that should accumulate — in stock, on promotion, recently updated — summing is the right model and bool should does exactly that. dis_max fits the opposite case: several clauses are alternative views of one user intent, usually the same text across several fields, where summing double-counts a single signal.
  • What happens if you set tie_breaker to 1.0?
    The formula becomes max plus the full sum of the other matching clauses, which is just the sum of all matching clauses — the same aggregation a bool should array performs. At that point dis_max adds nothing over bool, so 1.0 is a signal that dis_max was the wrong construct for the problem.
  • Does dis_max require every sub-query to match?
    No, it is a disjunction: a document matching any single query in the list is a hit. Only the scoring differs from a disjunction built with bool should. If you need mandatory predicates, wrap the dis_max inside a bool query and put those predicates in the filter array.

Grading a student on their best exam rather than the sum of all exams, with tie_breaker as the small amount of credit their other papers still earn when two students post the same top mark.

saying these in an interview costs you the question

  • Says dis_max requires all sub-queries to match
  • Thinks tie_breaker defaults to a small non-zero value
  • Believes dis_max sums the sub-query scores
  • Cannot distinguish dis_max from bool should on scoring
  • Uses dis_max to combine independent signals that should accumulate

context