When would you set an Elasticsearch field's similarity to boolean instead of BM25?
answer
- An alternative built-in similarity, not a query type
- Score equals the boost, nothing else
- Removes the rarity advantage of unusual values
- Best on tag, status and category fields
- Set once in the mapping, applies to every query
basics
~20 sUse the boolean similarity on filter-like fields where only matching should count. It scores a match at the query boost, ignoring term frequency, inverse document frequency and field length, so common and rare tag values contribute equally.
solid answer
~40 sThe `boolean` similarity is a built-in alternative to BM25 that you attach with the `similarity` mapping parameter. A matching document scores the query boost and nothing else — no term frequency, no idf, no length normalisation. That is what you want on tag, category, status or identifier fields you include in a scored query to nudge ranking: with BM25 those clauses score by idf, so an unusual tag silently outweighs a common one, which is rarely the intent. The alternative is to move the clause into filter context or wrap it in `constant_score`, which achieves a similar effect per query. Choose the mapping-level similarity when you want the behaviour to hold for every query against that field; choose query-level filtering when only some queries should treat it that way.
code
json · 6 lines{
"properties": {
"title": { "type": "text" },
"tags": { "type": "keyword", "similarity": "boolean" }
}
}go deeper
Know that BM25 is not the only similarity Elasticsearch offers, and that a field can be told to score matches flatly instead.
Explain that the boolean similarity scores a match at the query boost with no term frequency, idf or length component, and that it is chosen per field in the mapping.
Justify the choice against filter context and constant_score, and articulate the failure it prevents: idf letting a rare tag value dominate ranking for reasons nobody chose.
Decide where scoring policy lives — in mappings that bind every future query, or in query construction that each team controls — and encode that decision in index templates and review standards.
## What the boolean similarity does Elasticsearch ships more than one similarity implementation. Alongside `BM25` it exposes `boolean`, which you select with the `similarity` mapping parameter on a field: ```json { "tags": { "type": "keyword", "similarity": "boolean" } } ``` Its rule is deliberately trivial: a matching term scores the query boost, and a non-matching one scores nothing. There is no term frequency, no inverse document frequency and no field-length normalisation. The field becomes a yes/no signal whose weight you set explicitly through boosts rather than one whose weight the corpus statistics decide for you. ## The problem it solves Consider a product search that adds tag matching to the main text query so that a document tagged `waterproof` gets a lift when the user types "waterproof jacket". Under BM25, that clause is scored by idf. If only 12 of 400,000 products carry a niche tag, matching that tag produces an enormous score; if 300,000 carry a common one, matching it produces almost nothing. You did not choose that weighting — the tag distribution did, and it changes every time the catalogue changes. The same applies to status fields, identifiers, enumerated categories and any keyword field whose values are administrative rather than descriptive. Under the `boolean` similarity every matching value contributes the same amount, and you decide how much through the clause's boost. Relevance becomes something you tune rather than something that drifts with data. ## The alternatives, and how to choose There are three ways to stop a field's statistics from steering ranking: 1. **Filter context.** Put the clause in a `bool` query's `filter` section. It then decides matching without contributing score at all, and gains the benefit of caching. Correct when the field should never influence ranking. 2. **`constant_score`.** Wrap the clause so it contributes a fixed score you choose. Correct when it should influence ranking, but only for this query. 3. **`similarity: boolean` on the field.** The behaviour becomes a property of the field for every query anyone writes against it, including queries written by other teams later. The deciding question is whether the behaviour belongs to the field or to the query. A field whose values are a controlled vocabulary and which is never meant to be ranked by rarity is genuinely a boolean-scoring field — encoding that in the mapping stops the next engineer from reintroducing idf weighting by accident. A field that some queries want scored and others do not should keep BM25 and be handled per query. Be aware of the scope: the mapping parameter applies to *every* query on that field, and you cannot ask for BM25 on it for one particular search. If you need both behaviours, index the value into two fields with different similarities via a multi-field. ## Where it does not belong Do not put the boolean similarity on analysed text fields you actually rank on — `title`, `body`, `description`. There, term frequency and rarity are exactly the signals you want, and flattening them destroys relevance rather than stabilising it. The similarity is for fields whose values behave like labels, not like prose. Also remember what it does not remove. A `bool` query still sums its scoring clauses, so a document matching three boolean-scored tag clauses still outscores one matching a single tag — the per-clause weight is fixed, not the total. That is usually the desired behaviour, but say it out loud when explaining the design so nobody expects a single flat score. ## Configuring it Because the per-field `similarity` parameter is set in the mapping, plan it when the index is created or when you introduce the field; changing scoring behaviour on an existing field in a long-lived index is normally handled by reindexing into a new mapping behind an alias. On a mixed index it is entirely normal to see BM25 with default parameters on `title`, a lower-`b` BM25 on `body`, and `boolean` on `tags` — each field scored by the model that suits what the field actually is.
- How is this different from putting the clause in filter context?Filter context removes the clause's score contribution entirely and makes it cacheable, so the field cannot influence ranking at all. The boolean similarity keeps a contribution but fixes it at the query boost, so the field still nudges ranking by an amount you control. Filter context is a per-query decision; the similarity is a permanent property of the field.
- What surprises people about a scored term query on a plain keyword field?That the score varies at all. Keyword fields have norms disabled and hold a single value, so BM25 reduces to idf — matching a rare value scores far higher than matching a common one. Candidates who expect all exact matches to score alike are describing the boolean similarity's behaviour, not the default.
- Can you use BM25 on the same field for one query and boolean for another?No. The similarity is fixed in the mapping and applies to every query against that field. If both behaviours are genuinely needed, index the value twice as a multi-field, giving each sub-field its own similarity, and target whichever one the query needs.
BM25 pays you more for a rare word the way a scoring system rewards an unusual answer; the boolean similarity pays a flat fee per match, so you decide the fee instead of the corpus deciding it.
saying these in an interview costs you the question
- Thinks boolean similarity means the field stops being searchable
- Applies it to analysed text fields that are actually ranked on
- Says every matching document scores 1.0 regardless of boost
- Confuses it with a bool query or with filter context
- Believes it can be chosen per query rather than per field