When would you use Elasticsearch's script_score query rather than function_score?
answer
- It is the escape hatch, not the default
- The inner query's score is available to it
- Runs once per matching document
- One field access path is far cheaper
- Compiled and cached by source text
basics
~20 sscript_score wraps a query with a Painless script that computes the score directly, so it fits formulas the built-in function_score functions cannot express. The script runs for every matching document with no skipping, making it the expensive option.
solid answer
~50 sThe `script_score` query takes an inner `query` and a `script` whose return value becomes the document's score; the inner query's score is available in Painless as `_score`. Reach for it when the ranking formula is genuinely custom — combining several fields arithmetically, applying a business rule that no built-in function encodes, or blending signals in a way `boost_mode` cannot express. Prefer `function_score` when a built-in function already fits, because it is clearer and better optimised. The cost is real: the script executes once per matching document with no early termination, so a query matching millions of documents pays for millions of script invocations. Read fields through doc values (`doc['likes'].value`), never through `params._source`, which forces a per-document source fetch and parse. Scripts must return a non-negative value or the query errors. Painless gives you helpers such as `saturation`, `sigmoid` and `randomScore` in this context, so common shapes need no hand-rolled maths.
code
json · 11 lines{
"query": {
"script_score": {
"query": { "match": { "body": "kafka" } },
"script": {
"source": "_score * saturation(doc['likes'].value, params.pivot)",
"params": { "pivot": 50 }
}
}
}
}go deeper
Know that script_score lets a Painless script compute the score, that the inner query's score is exposed as _score, and that it is not the everyday tool.
Explain when a built-in function_score function is the better choice, and describe the cost model: one script execution per matching document with no early termination.
Demonstrate the operational habits — doc values over source, constants in params, narrowing the match set or rescoring a top-k window — and the non-negative-score constraint.
Own where custom scoring logic should live at all: an ever-growing Painless formula is code without tests or review, and at some point the signal belongs in an indexed feature or a separate re-ranking stage.
## The shape of the query `script_score` is a top-level query, not a function inside `function_score`. It takes an inner `query` that decides which documents match, and a `script` that computes each matching document's score: ```json { "script_score": { "query": { "match": { "body": "kafka" } }, "script": { "source": "_score * saturation(doc['likes'].value, params.pivot)", "params": { "pivot": 50 } } } } ``` Inside the script, `_score` is the score of the inner query, `doc['field']` gives access to that field's doc values, and `params` carries constants. There is also a `min_score` parameter that drops documents whose computed score falls below a threshold. ## When it is the right tool Use it when the formula cannot be expressed by composing the built-in functions. Typical genuine cases are arithmetic across several fields (`price` relative to `category_median_price`), a conditional business rule that branches on field values, or a combination that `boost_mode`'s six options cannot express. If what you need is "multiply by a compressed popularity field" or "decay by recency", `function_score` already has `field_value_factor` and the decay functions, and using a script instead trades clarity and performance for nothing. The reverse framing is useful in an interview: `function_score` is a small library of pre-optimised, self-documenting functions; `script_score` is the escape hatch when the library does not cover your case. Escape hatches are correct to use and wrong to reach for first. ## What it costs The script executes for every document the inner query matches. There is no way to skip documents that cannot reach the top hits, because an arbitrary script could promote any of them — the engine cannot reason about the formula. On a broad query this per-document cost dominates the search. Three things make that cost worse than it needs to be. Reading `params._source` forces the engine to fetch and parse the stored source for every matching document, which is orders of magnitude more expensive than reading a doc value; use `doc['field'].value` instead, which reads a columnar structure built for exactly this. Second, scripts are compiled and cached by their source text, so building the source string with values interpolated into it produces a new compilation per request and will eventually hit the compilation rate limit — put varying values in `params`. Third, a heavy script over a match-everything query is a self-inflicted wound; narrow the inner query, or apply the expensive formula only to a top-k window with a `rescore` block so the script runs on hundreds of documents rather than millions. ## Correctness constraints Scores must be non-negative. A script that can return a negative number causes the query to fail, so guard subtractions with `Math.max(0, ...)`. Fields that may be absent need a check — `doc['likes'].size() == 0` — before reading `.value`, or documents without the field will error. ## Built-in helpers The script_score context provides Painless functions for the shapes people would otherwise write by hand, and they are both more correct and faster than a hand-rolled version. `saturation(value, pivot)` returns `value / (value + pivot)`, a bounded curve that approaches 1 — the right shape for popularity, because it makes the difference between 100,000 and 1,000,000 views small. `sigmoid(value, k, a)` gives an S-curve with a tunable inflection. `randomScore(seed, fieldName)` produces a reproducible pseudo-random score for shuffling results. There are also decay helpers for numeric, date and geo fields mirroring the `function_score` decay functions. ## Relationship to function_score's script_score function Confusingly, `function_score` also has a `script_score` entry in its `functions` array. It runs the same kind of script, but as one function among several, whose result is then folded in by `score_mode` and `boost_mode`. Use that form when a script is one signal alongside a decay function and a weight; use the standalone `script_score` query when the script is the entire scoring formula. The standalone query is the newer and simpler surface, and it is what most current code should use for the whole-formula case. ## The judgment an interviewer is checking They want to hear that you know the escape hatch exists, that you can name its cost model in concrete terms (per-document execution, no skipping, doc values not source, params not string interpolation), and that you reach for a bounded, indexed alternative when the signal is a static per-document number. Someone who answers "script_score, because it can do anything" has given the wrong answer even though the sentence is true.
- Why is doc['likes'].value so much cheaper than params._source.likes inside a script?Doc values are a columnar, per-segment structure built at index time for exactly this access pattern, so reading one is close to an array lookup. `params._source` fetches the original JSON document from stored fields and parses it, per document, which is heavy enough to dominate the query on any broad match set.
- What goes wrong if you interpolate a changing value into the script source instead of using params?Scripts are compiled and cached keyed by their source text, so every distinct source string is a fresh compilation. A per-request value baked into the source produces a new compile each time, thrashes the cache and eventually trips the script compilation rate limit, causing requests to fail. Keeping the source constant and passing values in `params` avoids all of that.
- How do you keep an expensive scoring script off millions of documents?Score cheaply first and apply the script only to a top-k window using a `rescore` block, so the expensive formula runs on the hundreds of documents that could plausibly make the final page. Narrowing the inner query with filters is the other half. Both trade a small amount of ranking fidelity for a large amount of latency.
saying these in an interview costs you the question
- Reaches for a script when field_value_factor already fits
- Reads fields through params._source inside the script
- Builds the script source string per request instead of using params
- Assumes Elasticsearch can skip documents under an arbitrary script
- Returns a negative value from the script and expects it to rank last