What does the _score field in an Elasticsearch search hit represent, and how is it produced?
answer
- A number attached to every hit
- Decides ordering, not absolute quality
- BM25 is the default similarity
- Depends on corpus statistics that drift
- Null when you sort by a field
basics
~20 s_score is a relevance number Elasticsearch computes for a document against one specific query, using the BM25 similarity by default. Hits are sorted by it descending. It is not a percentage and is not comparable across queries.
solid answer
~50 sEvery hit carries a `_score`: a positive float produced by the field's similarity implementation, which is **BM25** by default in modern Elasticsearch. BM25 combines three ingredients — how often the term occurs in the field (with diminishing returns), how rare the term is across the shard's documents, and how long the field is compared with the average. Scores from multiple matching terms or `bool` clauses are summed. The key thing to say is what the number is *not*: it is not normalised to 0–1, not a percentage, and not stable across queries or indices, because it depends on corpus statistics that change as documents are added and deleted. Use it to order results, never as an absolute quality threshold. If you sort by a field instead of relevance, `_score` comes back as `null` unless you set `track_scores: true`.
go deeper
Be ready to say that _score is the relevance number Elasticsearch attaches to each hit, that results are sorted by it descending, and that BM25 produces it by default.
Explain the three ingredients BM25 combines — saturating term frequency, inverse document frequency, and field length — and why the resulting number is unbounded rather than a percentage.
Show the operational consequence: score magnitudes drift with corpus statistics and shard layout, so min_score cutoffs and score assertions in tests are fragile. Diff rankings instead.
Own the framing that raw relevance scores are not a product-level quality signal. If downstream systems need calibrated numbers, that calibration is a deliberate layer you design, not something BM25 gives you.
## What the number is Every entry in `hits.hits` of an Elasticsearch search response carries a `_score` field: a positive floating-point number computed for that document against that query. By default the result list is ordered by `_score` descending, so the score is what decides result order. The number answers exactly one question: *relative to the other documents this query matched, how good a match is this one?* ## Where the number comes from For full-text queries Elasticsearch delegates scoring to Lucene's `Similarity` implementation attached to the field. Since Elasticsearch 5.0 the default has been **BM25**, replacing the older classic TF-IDF model. BM25 combines three signals per query term: - **Term frequency (tf)** — how many times the term appears in that field of that document. The contribution *saturates*: the tenth occurrence adds far less than the second. - **Inverse document frequency (idf)** — how rare the term is. A term appearing in 3 documents is worth much more than one appearing in 300,000. It is computed from the shard's own statistics as `log(1 + (N - n + 0.5) / (n + 0.5))`, where `n` is the number of documents containing the term and `N` the number of documents that have the field. - **Field length normalisation** — a short `title` matching "kotlin" is a stronger signal than a 5,000-word `body` matching it once. The field length is stored at index time in a per-document structure called **norms**. A query with several terms, or a `bool` query with several scoring clauses, sums the per-clause scores. So a document matching two query terms typically outranks one matching only one. ## Why it is not a percentage This is the part interviewers actually probe. BM25 scores are unbounded and unnormalised. Their magnitude depends on corpus statistics — the number of documents, how many contain each term, the average field length — all of which drift as you index and delete documents. Consequences: - The same document and the same query on a different index give a different number. - Two different queries produce numbers that cannot be compared. `_score` 12.4 for query A says nothing about `_score` 9.0 for query B. - A `min_score` cutoff or a hardcoded rule like "above 5 is relevant" is brittle and will silently change behaviour as the corpus grows. The honest formulation: `_score` is an ordering device within one query, not a calibrated confidence. ## When _score is null, constant, or absent - **Sorting by a field.** If you supply a `sort` on a field, Elasticsearch skips score computation and returns `"_score": null` for every hit. Add `"track_scores": true` if you still want the relevance value alongside a field sort. - **Filter context.** Clauses placed in a `bool` query's `filter` or `must_not` sections, and clauses wrapped in `constant_score`, do not contribute a BM25 value — they only decide matching. A query made entirely of filters yields a constant score for every hit. - **Term-level queries in query context** still score: a `term` query on a `keyword` field produces a BM25 score driven almost entirely by that term's idf, which is why rare exact values score higher than common ones. ## Inspecting it You never have to guess where a score came from. Adding `"explain": true` to a search body attaches a `_explanation` tree to each hit, and `GET /<index>/_explain/<id>` explains one document in detail. Both break the number into its idf and tf components with the actual values that went in. ## Practical guidance Compare *rankings*, not raw scores. When relevance changes, diff the ordering of a fixed set of test queries rather than asserting on score values — score assertions in tests are notoriously flaky because they move with the corpus. If a downstream system genuinely needs a bounded number, derive it yourself (for example from rank position) rather than trusting BM25's magnitude. And when someone asks why the top hit "only" scored 3.2, the answer is that the absolute magnitude carries no information at all.
- Why is asserting an exact _score value in an automated test a bad idea?Because the value depends on corpus statistics — document count, how many documents contain the term, average field length — and on which shard the document landed. Adding, deleting or merging documents shifts it, and a change of shard count shifts it more. Assert on result ordering, or on which document is first, instead.
- If you sort search results by a date field, what happens to _score?Elasticsearch stops computing relevance and returns `"_score": null` for each hit, since the score is not needed for ordering. Setting `"track_scores": true` in the search body forces it to compute and return the score anyway, at the cost of the scoring work you just opted out of.
- Does a term query on a keyword field produce a meaningful score?Yes, in query context it is scored by BM25, but the term frequency is effectively one and `keyword` fields have norms disabled by default, so the score is driven by idf — rarer values score higher. If you do not want that skew, put the clause in filter context or give the field the `boolean` similarity.
saying these in an interview costs you the question
- Says _score is a 0-to-100 percentage match
- Claims scores are comparable between two different queries
- Thinks Elasticsearch still uses classic TF-IDF by default
- Believes a higher score means more query terms matched
- Hardcodes a min_score threshold as a relevance rule