skip to content

What does wrapping a query in Elasticsearch's constant_score query do to matching and scoring?

level: middleimportance: should knowfreq 45%

answer

  1. It changes the score, not the hits
  2. The inner query runs unscored
  3. Every match gets the same number
  4. That number is the boost, default 1.0
  5. Unlike bool filter, it does contribute

basics

~20 s

constant_score runs its inner query in filter context and gives every matching document the same score, taken from its boost parameter, which defaults to 1.0. Matching is unchanged; the inner query's own relevance calculation is discarded.

solid answer

~50 s

`constant_score` takes a single inner query under a `filter` key and executes it in **filter context**: the inner query's own relevance calculation is thrown away and every document it matches receives a flat `_score` equal to the `constant_score` clause's `boost`, which defaults to 1.0. Matching semantics are untouched — the same documents match as if the query ran alone. Its value is that it turns a scored clause into a fixed-value one *while keeping it in a scoring position*. Inside a `bool` `should` array, it lets you say "documents in this category get exactly 2.0 more, no matter what the term statistics say", which a bare `filter` clause cannot do because filter contributes nothing at all. It also stops rare-term inflation: a `match` on a low-frequency value would otherwise score far higher than a common one just because IDF says so.

code

json · 11 lines
json
{
  "query": {
    "bool": {
      "must": [ { "match": { "title": "running shoes" } } ],
      "should": [
        { "constant_score": { "filter": { "term": { "in_stock": true } }, "boost": 2 } },
        { "constant_score": { "filter": { "term": { "brand": "acme" } }, "boost": 1 } }
      ]
    }
  }
}

go deeper

for a junior

Recall the shape: constant_score wraps one query under a filter key and gives every match the same score, defaulting to 1.0. The hit set does not change.

for a middle

Explain why it exists next to bool's filter slot — filter contributes zero, constant_score contributes its boost — and why that matters inside a should array.

for a senior

Show the ranking-hygiene use: replacing a match on a categorical field so that a rare value stops outscoring a good text match, and reasoning about additive boost budgets across several should clauses.

for a principal

Own the scoring policy: which signals get flat, tunable bonuses with an explicit budget versus which stay genuinely relevance-driven, so ranking stays explainable as signals accumulate.

## What it is `constant_score` is a compound query with exactly one child, given under the key `filter`: ```json { "constant_score": { "filter": { "term": { "status": "published" } }, "boost": 1.2 } } ``` It does two things at once: 1. It runs the inner query in **filter context** — no similarity computation, and the resulting per-segment doc-id set is a candidate for the node query cache. 2. It assigns every matching document a `_score` of exactly `boost`, which defaults to **1.0**. The hit set is identical to running the inner query on its own. Only the score changes, and it changes to a constant. ## Why not just use bool's filter slot? Because the two live in different places in the score arithmetic. A clause in `bool`'s `filter` array contributes **nothing** — zero — to `_score`. A `constant_score` clause contributes its `boost`. So: - "this must be true, and it says nothing about ranking" → `bool` `filter`. - "this is optional, and matching it is worth exactly this much" → `constant_score` inside `bool` `should`. The second shape is the reason `constant_score` exists in practice. It is how you attach a flat, predictable bonus to a signal. ## Killing term-statistic noise The subtler use is defending ranking against BM25's own behaviour on non-text fields. Suppose you put `{ "match": { "category": "electronics" } }` in a `should` array as a mild preference. The contribution it makes depends on how many documents share that category: a rare category produces a large inverse-document-frequency contribution, a common one a small contribution. So the bonus for "is in the user's preferred category" silently varies by category — an obscure category can outweigh a genuinely good title match. Wrapping it in `constant_score` removes that variance. Every document in the category gets the same bonus, and you tune it with one number you control rather than with the shape of your data. ## Interaction with bool A few consequences worth being able to state: - A `bool` whose only clause is a `constant_score` in `must` gives every hit the same score — an intentional way to say "return these, ranking is decided elsewhere by an explicit sort". - Two `constant_score` clauses in a `should` array are **additive**: a document matching both gets the sum of the two boosts. That makes score budgets easy to reason about — the ceiling of the bonus stack is the sum of the boosts. - Because the inner query runs in filter context, it can be served from the node query cache under the same rules as a `bool` `filter` clause. Wrapping an expensive but repeated predicate in `constant_score` therefore buys both a predictable bonus and cache eligibility. - Putting a full-text `match` on the user's query text inside `constant_score` is almost always wrong: you have deliberately destroyed the relevance signal you needed. `constant_score` is for signals whose *presence* matters, not signals whose *quality* matters. ## Versus other flat-boost tools Setting `boost` directly on a leaf query is not the same thing: `boost` multiplies whatever score the query already produces, so the variance stays — a rare term still scores higher, just scaled. `constant_score` replaces the score outright. When someone says "I set a boost of 3 and it did not behave predictably", this is usually the reason, and `constant_score` is the fix. Richer, multi-signal score shaping — decays over recency, field-value factors — belongs to `function_score`, which is a different and heavier tool. ## Reading it back in explain output When you run the query with `"explain": true`, a `constant_score` clause shows a single flat contribution with no term-frequency or field-length detail beneath it. That is a fast way to confirm a clause is really running unscored: if you still see BM25 sub-details, the wrap is not where you thought it was. ## Summary of the decision Ask what the clause is worth. Nothing to ranking, but required → `bool` `filter`. A fixed amount, optionally → `constant_score` in `should`. An amount that genuinely depends on how well the text matched → leave it as a scored `match` in `must` or `should`.

  • How is constant_score different from setting boost on the inner query directly?
    boost multiplies the score the query already produced, so variation from term statistics survives, just rescaled. constant_score discards the inner score entirely and substitutes a flat value. If you want every matching document to get the same bonus regardless of how rare the matched term is, only constant_score gives that.
  • When would you prefer constant_score inside a should clause over simply putting the clause in bool's filter array?
    When the signal is optional and worth a specific amount. A bool filter clause is mandatory and contributes zero to _score, so it cannot express a bonus at all. constant_score in should makes the clause optional and gives matching documents exactly the boost value, which stacks additively with other should clauses.
  • Is a query inside constant_score eligible for the node query cache?
    Yes. The inner query runs in filter context, so it produces a doc-id set rather than scores, and that set is a caching candidate under the same usage and segment-size heuristics that govern a bool filter clause. Wrapping a repeated, expensive predicate therefore buys predictable scoring and cache eligibility together.

saying these in an interview costs you the question

  • Thinks constant_score changes which documents match
  • Says it always produces a score of zero
  • Confuses it with setting boost on the inner query
  • Wraps the user's main text query in it, destroying relevance
  • Believes the inner query still contributes its BM25 score

context