skip to content

In an Elasticsearch gauss decay function, what do origin, scale, offset and decay control?

level: middleimportance: should knowfreq 58%

answer

  1. Four parameters, one distance calculation
  2. One of them is a plateau of indifference
  3. Another names the half-life distance
  4. One curve shape can reach exactly zero
  5. Output never exceeds one

basics

~20 s

origin is the ideal value, offset is a grace band around it where the function stays at 1, scale is the distance beyond that band at which the function falls to the decay value, and decay defaults to 0.5. The gauss, exp and linear variants differ only in curve shape.

solid answer

~50 s

Decay functions score a document by how far a numeric, date or `geo_point` field sits from an ideal value. `origin` is that ideal — `"now"` for recency, a user's coordinates for proximity. `offset` (default 0) is a plateau: anything within that distance of `origin` scores a full 1.0. `scale` is the distance past the plateau at which the function returns exactly the `decay` value, which defaults to 0.5. So `origin: "now", offset: "7d", scale: "30d"` means documents up to a week old are untouched and a 37-day-old document scores 0.5. The three variants differ only in shape: `gauss` is a bell that falls slowly near the origin then steeply, `exp` drops fastest immediately and has a long tail, and `linear` is a straight line that reaches zero and stays there. Because the output is at most 1 and `boost_mode` defaults to `multiply`, a decay function demotes distant documents rather than promoting near ones. For multi-valued fields, `multi_value_mode` picks which value is used.

code

json · 18 lines
json
{
  "function_score": {
    "query": { "match": { "body": "outage" } },
    "functions": [
      {
        "gauss": {
          "published_at": {
            "origin": "now",
            "offset": "7d",
            "scale": "30d",
            "decay": 0.5
          }
        }
      }
    ],
    "boost_mode": "multiply"
  }
}

go deeper

for a junior

Recall that gauss, exp and linear score documents by distance from an ideal origin value, and that they work on numeric, date and geo_point fields.

for a middle

Be able to state that the function returns the decay value at origin plus offset plus scale, that decay defaults to 0.5, and that only linear reaches zero.

for a senior

Show how you pick a scale from the actual age or distance distribution of the corpus, and explain when you switch to boost_mode sum so freshness cannot swamp a strong text match.

for a principal

Own the question of whether recency belongs in the score at all versus in a separate result module or a sort, and how you would measure the ranking change before rolling it out.

## What a decay function expresses Decay functions answer "how close is this document to what the user wants?" for a dimension that has a natural distance: how recent, how near, how close to a target price. They belong to the `functions` array of a `function_score` query and come in three shapes — `gauss`, `exp` and `linear` — that share the same four parameters and differ only in the curve between them. The supported field types are numeric, `date` (and `date_nanos`) and `geo_point`, because those are the types on which "distance from a value" is meaningful. ## The four parameters **origin** is the point of maximum score — the ideal. For recency it is usually `"now"`; for proximity it is the searcher's coordinates; for a price preference it is the target price. A document whose field value equals `origin` scores 1.0. **offset** (default 0) creates a plateau of indifference around the origin. With `offset: "7d"`, everything published in the last seven days is treated as equally fresh and scores 1.0; decay only begins beyond that. This matters more than it looks: without an offset, a two-hour-old document outranks a two-day-old one on freshness alone, which is usually noise rather than signal. **scale** together with **decay** defines the steepness. The contract is precise: at a distance of `offset + scale` from the origin, the function returns exactly `decay`. `decay` defaults to 0.5, so the common reading of `scale` is "the half-life": the distance at which relevance halves. Setting `decay: 0.2` with the same scale makes the curve much steeper, since it must fall to 0.2 over the same distance. For dates and geo distances the values are written as time or distance strings — `"30d"`, `"12h"`, `"10km"` — and for numeric fields they are plain numbers. ## Choosing the shape All three curves pass through 1.0 at the origin and through `decay` at `offset + scale`. Between and beyond those points they behave differently. `gauss` is the usual default. Its bell shape is flat near the origin, so small differences close to the ideal barely matter, then it falls away steeply, then it flattens again into a long shallow tail. That matches most human intuitions about "recent" or "nearby". `exp` falls fastest right at the origin and then decays ever more slowly. Use it when even a small departure from the ideal should cost something and you still want far-away documents to keep a small non-zero score. `linear` is a straight line, and it is the only one of the three that reaches zero: past twice the given scale the function returns 0, which zeroes out the document's contribution entirely under `boost_mode: multiply`. That hard cutoff is either exactly what you want or a nasty surprise, so choose it deliberately. ## Decay demotes; it does not promote A decay function returns a value in the range 0 to 1. With the default `boost_mode: multiply`, the best a document can do is keep its full text score, and everything else is scaled down. Candidates often describe decay as "boosting fresh content" — mechanically it is demoting stale content, and the distinction shows up when a decay function multiplied against a weak text score produces a number so small the document falls off the page. If you want a bounded additive nudge instead, use `boost_mode: sum` with a `weight` on the decay function, so the freshest documents gain a fixed amount and old ones simply gain nothing. ## Multi-valued fields When the field holds several values — several event dates, several store locations — `multi_value_mode` decides which one feeds the distance calculation: `min`, `max`, `avg` or `sum`. The default is `min`, meaning the closest value wins, which is the right semantics for "is there a branch near me?" and the wrong one for "is the whole tour close to me?". ## Practical failure modes The two classic mistakes are a scale that is far too small — a `scale` of `"1d"` on a corpus spanning years means everything older than a couple of days scores near zero and relevance effectively becomes a sort by date — and stacking multiple decay functions under `score_mode: multiply`, where three values around 0.5 multiply to 0.125 and the combined demotion is far harsher than any of them intended. Start with a scale on the order of the age range that genuinely matters to users, add an offset so the newest bucket is not over-differentiated, and verify against judged queries rather than by looking at one search.

  • Why does adding a recency decay sometimes make relevance obviously worse?
    Usually the scale is far too small for the corpus. If `scale` is a day and documents span years, almost everything multiplies by a near-zero factor and ranking collapses into a date sort. Widen the scale to the age range users actually care about, add an offset so the newest documents are not over-separated, or switch to `boost_mode: sum` so freshness adds rather than multiplies.
  • How is a decay function different from just sorting by date?
    A sort by date ignores relevance completely; a decay function keeps BM25 in the ranking and lets a much better textual match beat a slightly fresher document. Decay is the tool when freshness is one signal among several. If freshness must strictly win, a sort is the honest implementation and far cheaper.
  • What does multi_value_mode default to and when is that wrong?
    It defaults to `min`, so the closest of the field's values determines the distance. That is right for "any nearby branch qualifies" and wrong when the aggregate matters — a multi-date event where the average or latest date is what users judge. Set `avg` or `max` explicitly in those cases.

It is a dimmer switch on a document's relevance, wired to distance: full brightness inside the plateau, then dimming at a rate you choose, with only one of the three curves able to switch off completely.

saying these in an interview costs you the question

  • Describes decay as boosting fresh documents rather than demoting old ones
  • Thinks scale is the point where the score reaches zero
  • Believes gauss and exp can return exactly zero
  • Stacks several decay functions under score_mode multiply
  • Leaves offset at zero so hours-old differences dominate freshness

context