skip to content

How do you override BM25's k1 and b for an Elasticsearch index or for a single field?

level: middleimportance: should knowfreq 52%

answer

  1. Two named parameters, one similarity block
  2. One controls repetition, one controls length
  3. Attached per field through the mapping
  4. Static setting — closed index to change
  5. No reindex needed; norms hold length separately

basics

~20 s

Define a named similarity in the index settings with type BM25 and your k1 and b values, then reference that name from a field's similarity mapping parameter. k1 controls term-frequency saturation, b controls field-length normalisation.

solid answer

~40 s

You declare a similarity under `index.similarity.<name>` with `"type": "BM25"` plus `k1` and `b`, then attach it to a field with the `similarity` mapping parameter — or override the index-wide default by naming the similarity `default`. `k1` sets how quickly repeated occurrences of a term stop adding score: higher k1 lets extra occurrences keep counting, `k1: 0` makes only presence matter. `b` sets how strongly a long field is penalised: `b: 0` disables length normalisation entirely, `b: 1` applies it fully. These are static index settings, so you set them at index creation or update them on a closed index. Because norms store field length independently of these parameters, changing k1 or b re-scores existing documents at query time without a reindex — verify with the explain output.

code

json · 16 lines
json
{
  "settings": {
    "index": {
      "similarity": {
        "long_text_bm25": { "type": "BM25", "k1": 1.2, "b": 0.3 }
      }
    }
  },
  "mappings": {
    "properties": {
      "title": { "type": "text" },
      "body":  { "type": "text", "similarity": "long_text_bm25" },
      "tags":  { "type": "keyword", "similarity": "boolean" }
    }
  }
}

go deeper

for a junior

Know that Elasticsearch's default scoring is BM25 and that its two parameters are named k1 and b, even if you have never changed them.

for a middle

Explain what each parameter does to the scoring curve, and show that you declare a named similarity in index settings and reference it from the field's mapping.

for a senior

Discuss the operational cost: these are static settings needing a closed index or a template change, and the tuning is only defensible if you can measure the relevance effect on a fixed query set.

for a principal

Frame k1 and b as a low-leverage lever compared with analysis, field boosts and query structure, and decide whether a corpus-wide scoring change is worth the coordination it forces across index templates.

## What you are actually configuring Every field in an Elasticsearch index is scored by a Lucene `Similarity`. The default is BM25 with the standard parameters `k1 = 1.2` and `b = 0.75`. Both are exposed as configuration, so you can change the shape of the scoring curve without writing code. ## k1 — term-frequency saturation The term-frequency part of BM25 is `freq / (freq + k1 * (1 - b + b * dl / avgdl))`. The `freq` in the numerator grows linearly, but it also sits in the denominator, so the whole fraction approaches 1 as `freq` grows — repeated occurrences give diminishing returns. `k1` controls how fast that ceiling is approached: - `k1 = 0` — the fraction becomes 1 for any non-zero frequency, so repetition is ignored completely and only presence matters. - Low `k1` (say 0.5) — saturation is quick; the second occurrence adds little. - High `k1` (say 2.0) — occurrences keep adding meaningful score for longer, which suits corpora where genuine repetition signals aboutness. The practical reason this knob exists is keyword stuffing: without saturation, a document repeating a term fifty times would swamp a genuinely relevant one. ## b — field length normalisation `b` scales how much the field's length relative to the average length dampens the score. In the denominator, `(1 - b + b * dl / avgdl)` interpolates between ignoring length (`b = 0`) and applying it fully (`b = 1`). - `b = 0` — a 5,000-word body matching once scores the same as a three-word title matching once. Useful when field length is genuinely uninformative, for example product names of wildly differing verbosity that you do not want penalised. - `b = 1` — full penalty; long fields must match more to compete. A common tuning move on long-form `body` fields is to lower `b` a little, because otherwise long but authoritative documents lose to short thin ones. ## Declaring the similarity Similarities are named blocks in the index settings: ```json { "settings": { "index": { "similarity": { "long_text_bm25": { "type": "BM25", "k1": 1.2, "b": 0.3 } } } } } ``` Then the field opts in through the `similarity` mapping parameter: ```json { "body": { "type": "text", "similarity": "long_text_bm25" } } ``` If you name a similarity `default`, every field that does not specify one uses it — that is how you change the whole index at once. Fields keep their own similarity when they declare one, so a mixed index is normal: BM25 with default parameters on `title`, a low-`b` BM25 on `body`, and the `boolean` similarity on tag-like keyword fields. Besides `BM25` and `boolean`, Elasticsearch exposes several other Lucene models (DFR, DFI, IB, LM Dirichlet, LM Jelinek-Mercer) and a `scripted` similarity for full control. In practice interviews stay on BM25 and boolean; the others are almost never used in production without a research reason. ## Static settings and what a change costs Index-level similarity settings are static: you supply them when creating the index, or close the index, update the settings, and reopen it. That is an availability event for that index, so on hot indices the usual route is to bake the setting into an index template so new backing indices pick it up, or to reindex into a fresh index. The good news is what a k1/b change does **not** require. Norms store the field length that BM25 needs, and that stored value does not depend on k1 or b — so changing the parameters changes scoring for already-indexed documents immediately at query time, with no reindex. (Disabling norms is a different matter and is not reversible without reindexing.) ## When it is worth doing Honestly: rarely as a first move. Tuning k1 and b is a small, corpus-wide adjustment; field boosts, better analysis, and query structure usually move relevance far more. Reach for these parameters when you have a specific diagnosed symptom — long documents systematically losing (lower `b`), or repetition being over-rewarded in a spammy corpus (lower `k1`) — and when you have a way to measure the result. Change one parameter at a time, and use the explain output on a handful of known documents to confirm the tf component moved in the direction you expected.

  • Does changing k1 or b require reindexing the existing documents?
    No. Norms store the field length that BM25 needs, and that value is independent of k1 and b, so the new parameters apply to already-indexed documents as soon as queries run against the updated settings. The change itself is a static index setting, so you set it at creation or update it on a closed index.
  • What does setting k1 to 0 do?
    It removes term-frequency saturation entirely by making the tf fraction 1 for any non-zero frequency: a document matching the term once scores the same as one matching it twenty times. Scoring then depends on idf and, if b is non-zero, still on field length. It is a blunt way to make repetition irrelevant.
  • Which field would you most likely give a lower b, and why?
    A long-form body or description field. With the default b, long documents are penalised relative to short ones for the same match, so an authoritative long article can lose to a thin stub. Lowering b softens that penalty. The opposite case — short, uniform fields like titles — gains nothing from the change.

k1 is a volume limiter on repetition and b is a handicap for long fields: turn the limiter down and shouting the same word stops helping, turn the handicap off and a long document competes on equal terms with a short one.

saying these in an interview costs you the question

  • Thinks k1 and b are query-time parameters passed per search
  • Claims changing k1 or b requires a full reindex
  • Says b controls term frequency and k1 controls length
  • Assumes the similarity is index-wide and cannot be per field
  • Tunes k1 and b before fixing analysis or field boosts

context