skip to content

How do you use Elasticsearch's _explain API to see why a document got the BM25 score it did?

level: middleimportance: must knowfreq 70%

answer

  1. Two entry points: one document or every hit
  2. Response is a nested value/description/details tree
  3. Formulas appear literally in the description strings
  4. k1 and b are visible in the tf node
  5. The statistics shown are shard-local

basics

~20 s

Call GET /<index>/_explain/<id> with the query body, or add "explain": true to a search. Elasticsearch returns a nested tree breaking the score into per-clause weights with the idf and tf formulas and the actual N, n, freq, dl and avgdl values used.

solid answer

~50 s

There are two entry points. `GET /<index>/_explain/<id>` with a query body explains one document and also tells you whether it matched at all (`"matched": false` is the fastest way to prove a query never hit the document). Adding `"explain": true` to a `_search` body attaches an `_explanation` to every hit instead. The output is a tree: each leaf is a term weight described as `score(freq=...), computed as boost * idf * tf from:`, with child nodes spelling out `idf, computed as log(1 + (N - n + 0.5) / (n + 0.5))` and `tf, computed as freq / (freq + k1 * (1 - b + b * dl / avgdl))`, each carrying the real numbers. Parent nodes show how clause scores were combined — summed for `bool` clauses, maxed for `dis_max`. The crucial caveat: the statistics are the ones on the shard holding that document, so they will not match a cluster-wide count.

code

bash · 4 lines
bash
GET /articles/_explain/1
{
  "query": { "match": { "title": "elasticsearch" } }
}

go deeper

for a junior

Know that Elasticsearch can tell you why a document scored what it did, via GET /<index>/_explain/<id> or explain: true on a search.

for a middle

Walk through the output tree: the term weight, the idf formula with N and n, the tf formula with freq, k1, b, dl and avgdl, and the sum or max node that combines clauses.

for a senior

Use it as a diagnostic instrument — verify a custom similarity applied, catch a stemmer that made a term common, and remember that the statistics shown are shard-local so scores differ across shards.

for a principal

Treat explain output as evidence in a relevance-change process: capture query, mapping and shard count together, and insist that relevance claims are backed by explanations rather than intuition.

## The two ways in **Single document.** `GET /<index>/_explain/<id>` takes a query body and answers two questions at once: did this document match, and if so what did the score consist of. The response has `"matched": true|false` plus an `explanation` object. If the document lives on a shard chosen by custom routing, you must pass the same `routing` value or the request looks on the wrong shard. **Whole result set.** Adding `"explain": true` to a `_search` body attaches an `_explanation` to each hit. This is convenient for comparing why document A beat document B, but it is expensive — never leave it on in a production search path. ## Reading the tree The explanation is recursive: a `value`, a human-readable `description`, and `details` containing the children whose values produced it. For a single term on a single field you will see something like: - `weight(title:elasticsearch in 0) [PerFieldSimilarity], result of:` — the top-level weight for this term on this field. `PerFieldSimilarity` is a reminder that the similarity is resolved per field, so a field with a custom similarity is scored by that one. - `score(freq=1.0), computed as boost * idf * tf from:` — the BM25 product for this term. - `idf, computed as log(1 + (N - n + 0.5) / (n + 0.5)) from:` — with child values `n` (documents containing the term) and `N` (documents that have this field). Both come from the shard. - `tf, computed as freq / (freq + k1 * (1 - b + b * dl / avgdl)) from:` — with child values for `freq` (occurrences in this document's field), `k1` and `b` (the configured similarity parameters — this is where you verify a custom similarity actually took effect), `dl` (this document's field length) and `avgdl` (the average field length on the shard). Above the leaves sit combination nodes. A `bool` query's scoring clauses appear as `sum of:`; a `dis_max` or a `multi_match` in `best_fields` mode appears as `max of:` with a tie-breaker term; a `constant_score` shows a fixed value with no BM25 breakdown; a `function_score` shows the base query score and the function applied to it. ## What you diagnose with it - **"Why did the wrong document win?"** Compare the two trees. Usually one document has a much shorter field (`dl` far below `avgdl`) or matched an extra clause the other missed. - **"Why is this rare-looking term not boosting?"** Look at `n` in the idf node. A stemmer or a synonym filter may be collapsing the term into a far more common one, making it much less rare than you assumed. - **"Did my custom similarity apply?"** Read `k1` and `b` in the tf node. If they are the defaults, your custom similarity never reached the field. - **"Did my boost land where I thought?"** The boost appears as a multiplier in the weight node, so a field boost that never took effect is visible immediately. - **"Did the query even match?"** `"matched": false` from `_explain` ends the argument. From there the usual culprit is analysis, which you chase with the `_analyze` API rather than with explain. ## The distributed caveat The `N`, `n` and `avgdl` in the output are **shard-local**. On a multi-shard index the numbers in one document's explanation come from that document's shard, and a second copy of the same document on another shard would show different statistics and a different score. That is not a bug in explain; it is how term statistics are gathered. It also means a `_count` of documents containing a term across the index will not equal the `n` you see. A related subtlety: term and collection statistics include documents that have been deleted but whose segments have not yet been merged away. So scores drift slightly after merges, and an explanation taken before and after a force merge on a static index can differ. ## What explain is not - It is **not** `GET /<index>/_validate/query?explain=true`, which shows how the query was parsed and rewritten into Lucene syntax but says nothing about scores. - It is **not** the profile API. `"profile": true` reports where *time* went across query and collector phases; explain reports where *score* came from. Use profile for latency problems, explain for relevance problems. ## Cost and habit Explaining is real work: it re-runs scoring with instrumentation. Use `_explain` on individual documents while debugging, keep `"explain": true` out of production traffic, and when you do capture an explanation for a bug report, capture the query, the mapping of the fields involved, and the shard count alongside it — the tree is much harder to interpret without them.

  • The idf node reports n = 40000 for a term you expected to be rare. What do you check next?
    The analysis chain. Run the `_analyze` API on the field to see what token the query text actually produced — a stemmer, a synonym filter or an ascii-folding filter may be collapsing your rare term into a common one, so the indexed term is far more frequent than the surface word. The explain output is telling you about the token, not the word you typed.
  • How does the explain output differ between a bool query and a dis_max query?
    A `bool` query's scoring clauses appear under a `sum of:` node, so matching more clauses adds score. A `dis_max` — including `multi_match` in `best_fields` mode — appears under a `max of:` node, taking the single best clause plus a fraction of the others via the tie-breaker. Seeing which node you got tells you immediately whether extra field matches are accumulating.
  • Why can an explanation's N not match a _count of the index?
    Explain runs on one shard and reports that shard's statistics, so N is the number of documents on that shard that have the field, not the cluster-wide total. Deleted documents whose segments have not been merged away are still counted too, so even a single-shard index can report more than a live _count.
  • When would you reach for the profile API instead of explain?
    When the problem is latency, not ranking. Profiling with `"profile": true` reports the time spent in each query component and collector on each shard; explain reports how a score was assembled and costs extra work to produce. Slow query goes to profile, wrong ordering goes to explain.

saying these in an interview costs you the question

  • Confuses _explain with _validate/query?explain=true
  • Thinks explain reports timing rather than score composition
  • Reads N and n as cluster-wide document counts
  • Leaves "explain": true enabled on production searches
  • Cannot say which part of the tree shows field length

context