skip to content

Why can Elasticsearch's rank_feature and distance_feature queries be cheaper than an equivalent function_score?

level: seniorimportance: should knowfreq 34%

answer

  1. One is a wrapper, the others are queries
  2. Top-k pruning needs an upper bound
  3. Their contributions are bounded by design
  4. They add into a bool rather than multiply
  5. Two signal shapes only: popularity and distance

basics

~10 s

rank_feature and distance_feature are ordinary queries whose per-document contribution is indexed and bounded, so Elasticsearch can skip documents that cannot reach the top hits. function_score must score every document its inner query matches.

solid answer

~50 s

Both are real queries rather than scoring wrappers, so they sit in a `bool` alongside the text query and contribute additively. A `rank_feature` field stores a single strictly positive value per document in a form the query can read with a known upper bound, and the `rank_feature` query applies a bounded curve to it — `saturation` (the default, taking a `pivot`), `log` (taking a `scaling_factor`) or `sigmoid` (taking `pivot` and `exponent`). Because the maximum possible contribution is known, the engine can stop considering documents that cannot enter the top hits. `distance_feature` does the same for proximity on `date`, `date_nanos` and `geo_point` fields, scoring `boost × pivot / (pivot + distance)` from an `origin`. `function_score` gives no such bound: its function could promote any document, so every match must be scored. The trade-off is expressiveness — these two cover popularity and proximity, not arbitrary formulas.

code

json · 11 lines
json
{
  "query": {
    "bool": {
      "must": [ { "match": { "body": "kafka rebalance" } } ],
      "should": [
        { "rank_feature": { "field": "pagerank", "saturation": { "pivot": 10 } } },
        { "distance_feature": { "field": "published_at", "origin": "now", "pivot": "30d", "boost": 2 } }
      ]
    }
  }
}

go deeper

for a junior

Recall that rank_feature and distance_feature are queries that add a popularity or proximity signal to a score, unlike function_score which wraps and rewrites a query's scores.

for a middle

Explain the bounded-contribution idea: because the maximum score each can add is known, the engine can prune documents that cannot reach the top hits.

for a senior

Show the composed bool shape in production, pick a sensible saturation pivot from the field's distribution, and know the mapping constraints that make rank_feature a query-only field.

for a principal

Own the decision of which signals get promoted into indexed, bounded features versus left in a general scoring function, and what that costs in indexing pipeline complexity and re-scoring when a signal changes.

## Wrapper versus query `function_score` wraps a query and rewrites its scores. That wrapping is what costs: for every document the inner query matches, the functions must run, because the engine has no way to know that a given document's function result could not lift it into the top hits. Broad queries therefore pay a per-document cost across the whole match set. `rank_feature` and `distance_feature` are queries in their own right. They can appear directly in a `bool`'s `should` array beside the text query, and their contributions add into the same score sum as any other clause. Crucially, each has a computable ceiling on how much it can contribute, which lets the top-k collection logic discard documents whose best possible total is already below the current worst hit. That skipping is the source of the performance difference. ## rank_feature A `rank_feature` field holds one strictly positive numeric value per document — a popularity score, a PageRank-like authority, a quality signal. `rank_features` is its sibling for a sparse map of many named features on one document. The field takes `positive_score_impact`, which defaults to `true`. Setting it to `false` inverts the relationship, so a lower value scores higher — the way to express "faster page load ranks better". The `rank_feature` query applies one of several bounded functions to the stored value: - `saturation` is the default and computes a curve that rises quickly at small values and approaches a ceiling. It takes a `pivot`, the value at which the function reaches half its maximum; if you omit it, an approximation is derived from the values in the index. This is the shape most popularity signals want: the difference between 10 and 100 should matter far more than the difference between 100,000 and 1,000,000. - `log` requires a `scaling_factor` and grows logarithmically, useful when the signal keeps meaning something at large magnitudes. - `sigmoid` takes a `pivot` and an `exponent` and gives an S-curve whose steepness you control. The cost of this design is narrowness. A `rank_feature` field exists to be queried by the `rank_feature` query; if you also need to sort, aggregate or filter on the same number, index a plain numeric copy of it alongside. The value must also be strictly positive, and updating it means updating the document. ## distance_feature `distance_feature` covers the recency and proximity cases. It works on `date`, `date_nanos` and `geo_point` fields and takes an `origin` (`"now"`, a timestamp, or coordinates), a `pivot`, and an optional `boost`. Its score is `boost × pivot / (pivot + distance)`, so `pivot` is the distance at which the contribution falls to half of `boost`. Like `rank_feature`, its output is bounded above by `boost`, which is exactly what makes skipping possible. Against a `gauss` decay function it is less expressive — one curve shape, no `offset` plateau — but it is a query, so it composes into `bool` naturally and adds rather than multiplies. That additive placement is often the better relevance design anyway: a fresh document gains a bounded amount, and no amount of freshness can bury a far better text match. ## Composing them The idiomatic shape is a single `bool` where the text query is the `must` and the signals are `should` clauses: ```json { "bool": { "must": [ { "match": { "body": "kafka rebalance" } } ], "should": [ { "rank_feature": { "field": "pagerank", "saturation": { "pivot": 10 } } }, { "distance_feature": { "field": "published_at", "origin": "now", "pivot": "30d", "boost": 2 } } ] }} ``` Every contribution is bounded, each is independently tunable via its `boost`, and no clause can dominate the others by orders of magnitude the way an unbounded multiplicative factor can. ## When function_score is still right These two queries cover a specific pair of signal shapes. When the signal is not a static per-document value or a distance — when it is a conditional business rule, a formula across several fields, a random shuffle, or a case where multiplying rather than adding is genuinely correct — `function_score` (or `script_score`) remains the tool. The engineering judgment is to use the bounded, indexed form for the common popularity and recency cases, which are also the highest-traffic ones, and keep the general mechanism for the genuinely irregular remainder. ## What an interviewer is listening for The key sentence is that skipping requires a known upper bound on a document's possible contribution. Candidates who answer only "rank_feature is faster" have memorised a fact; candidates who explain why a bounded, indexed contribution enables top-k pruning while an arbitrary function does not have understood the mechanism, and the same reasoning transfers to other parts of the engine.

  • What does the pivot parameter mean in a rank_feature saturation function?
    It is the feature value at which the function returns half its maximum contribution, so it sets where the curve bends. Choose it near the median of the field's values so the signal discriminates across the bulk of the corpus. If you omit it, Elasticsearch derives an approximation from the values present in the index, which is a reasonable starting point but not tuned to your intent.
  • Can you sort or aggregate on a rank_feature field?
    Treat it as a query-only field: it exists to be scored by the rank_feature query, and it accepts only single, strictly positive values. If the same number also has to drive a sort, a range filter or an aggregation, index a plain numeric copy of it in a separate field and keep the two in sync at index time.
  • How does distance_feature differ from a gauss decay function for recency?
    distance_feature is a query, so it composes into a bool and adds a bounded amount, and it can skip non-competitive documents. A gauss decay lives inside function_score, multiplies by default, and supports an offset plateau and a choice of curve shapes. Decay is more expressive; distance_feature is cheaper and structurally safer against one signal dominating.

A bounded query is a bidder with a published maximum bid, so the auctioneer can ignore them once the price passes it. An arbitrary scoring function is a bidder who might offer anything, so nobody can be dismissed early.

saying these in an interview costs you the question

  • Says rank_feature is faster without explaining bounded contributions
  • Expects to aggregate or sort on a rank_feature field
  • Stores zero or negative values in a rank_feature field
  • Thinks these queries replace function_score for arbitrary formulas
  • Assumes function_score can skip low-scoring documents too

context