How do you fold document freshness into a relevance score without letting recency dominate?
answer
- multiply the score, do not add to it
- bound the penalty with a floor
- a flat window before decay starts
- not every query wants the newest document
basics
~20 sMultiply the relevance score by a decay factor between zero and one that falls with age, instead of adding a freshness bonus. A bounded multiplier preserves relevance ordering among documents of similar age and caps how much recency can distort ranking.
solid answer
~50 sUse a **decay function** on age that yields a multiplier in a bounded range, and multiply the text relevance score by it. Multiplying is the key choice: relevance scores are unbounded and vary hugely between queries, so an additive freshness bonus that is negligible on one query swamps the text signal on another, whereas a multiplier rescales consistently. Give the function three properties: a **plateau** at the start so documents from the last few hours or days are not reordered by minutes of age difference, a **half-life** (or scale) matched to how fast content in your corpus actually goes stale, and a **floor** so an old but perfect match can still surface rather than decaying to zero. Then check that recency is what the query wants — "news today" and "how does TLS work" have opposite freshness needs. The same shape works for geographic distance from the user.
code
text · 4 lines# bounded exponential recency multiplier
lam = ln(2) / half_life_days # half the multiplier per half-life
decay = exp(-lam * max(0, age_days - plateau_days)) # plateau: no penalty at first
score = relevance * (floor + (1 - floor) * decay) # floor: never decay to zerogo deeper
Know that freshness is normally applied as a multiplier on the relevance score, and that sorting purely by date is a different feature from relevance ranking.
Explain why multiplying is safer than adding — relevance scores are unbounded and vary per query — and describe the shape knobs: plateau, half-life or scale, and a floor.
Show that you derive the half-life from engagement data, gate decay on query intent, and understand the operational fallout of a time-dependent score on caching and bug reproduction.
Frame freshness as a product decision with a cost: it trades away some best-answer quality for perceived currency, and the right curve differs per vertical, so argue for measured, per-surface settings rather than one global constant.
## Why freshness needs shaping at all Sorting by date is not relevance tuning; it throws the text signal away and returns whatever was published most recently regardless of whether it answers the query. The opposite extreme — pure text relevance — returns a five-year-old page for a query about an event last week. Freshness tuning is the middle ground: keep relevance as the primary signal and let age modulate it. ## Multiplicative, not additive Relevance scores are not normalised. A query on a rare phrase can produce scores an order of magnitude larger than a query on common words, and the spread between the first and tenth result differs per query too. That makes an additive freshness bonus unmanageable: `score + 5` is decisive on a query whose scores range from 2 to 6 and invisible on one whose scores range from 40 to 90. A multiplier in a bounded range — say 0.3 to 1.0 — behaves the same on both queries, because it rescales rather than displaces. Multiplication also has a useful invariant: among documents of identical age, the multiplier is constant, so their relative order is exactly the text ranking. Freshness only ever competes *across* age groups, which is what you meant. ## The three common shapes **Exponential decay** falls fastest at the origin and flattens out, with a constant half-life: every N days the multiplier halves. It suits content that loses value steadily and never quite becomes worthless — documentation, forum answers. **Gaussian decay** is flat near the origin, drops steeply in the middle, then flattens again. It suits a notion of "about this recent": everything within a window is roughly equivalent, then relevance falls off a cliff, then everything ancient is equally ancient. **Linear decay** falls at a constant rate and reaches zero at a defined point. Its distinguishing property is that it *does* hit zero, which is occasionally what you want — an events listing where anything past the event date is genuinely worthless — and usually what you don't. A useful parameterisation, whichever shape you pick, has four knobs: the **origin** (usually "now"), a **plateau** or offset where no penalty applies at all, a **scale** with a stated multiplier value at that distance (a half-life is just "multiplier 0.5 at the scale"), and a **floor** below which the multiplier never falls. ``` lam = ln(2) / half_life_days decay = exp(-lam * max(0, age_days - plateau_days)) score = relevance * (floor + (1 - floor) * decay) ``` The plateau prevents pointless churn — without it, two articles published four minutes apart get different multipliers and the result page reshuffles constantly. The floor prevents the pathological case where an obscure, perfectly-matching ten-year-old document is mathematically unable to reach page one. ## Picking the half-life from data, not taste The half-life is an empirical property of the corpus, and you can measure it: for a sample of queries, look at the age distribution of the documents users actually engaged with, and pick a half-life where the multiplier decays at roughly the rate that engagement does. A news corpus may want hours; a knowledge base may want a year or no decay at all. Different query classes in the same corpus often want different half-lives, which is an argument for detecting query intent rather than applying one global curve. ## Not every query wants recency Applying freshness uniformly is the most common error. "Champions league result" wants today; "how does TCP congestion control work" wants the best explanation ever written. If you can classify intent — by query pattern, by category, by a temporal-intent classifier, or simply by which vertical the search came from — apply decay only where it belongs, and consider making the half-life a per-vertical setting rather than a global constant. ## Geography is the same function Distance from the user is structurally identical: origin is the user's location, the plateau is "anything in the same neighbourhood is equally close", the scale is a distance at which the multiplier halves, and a floor keeps a genuinely excellent distant result reachable. The Gaussian shape is a natural fit for geography because "nearby" really does behave like a flat zone followed by a fall-off. ``` decay = exp(-ln(2) * (distance_km / scale_km)^2) # half the multiplier at scale_km score = relevance * (0.2 + 0.8 * decay) ``` ## Operational consequences A score that depends on the current time is not reproducible. Two things follow. First, result caching becomes age-sensitive — cache keys must incorporate a coarsened time bucket, or the cache serves yesterday's decay. Second, a bug report saying "this ranked wrong at 09:14" cannot be reproduced later unless you can pin the origin to a fixed timestamp; make the evaluation harness able to freeze it. Bucketising age (hour, day, week) rather than using continuous seconds mitigates both problems at negligible relevance cost, and it also stops every document's score changing on every tick. Finally, do not implement freshness by re-indexing documents with a recomputed decay value on a schedule. That turns a cheap query-time arithmetic operation into a corpus-wide write amplification problem, and it makes the decay stale between runs.
- How would you choose the half-life instead of guessing it?Measure it. For a sample of queries, plot the age distribution of the documents users actually engaged with and pick a half-life whose curve decays at roughly that rate. Expect different answers per query class in the same corpus — breaking news versus reference material — which is an argument for a per-vertical or intent-conditioned half-life rather than one global constant.
- What does a time-dependent score break in testing and caching?Reproducibility and cache correctness. Scores drift continuously, so a ranking bug reported this morning cannot be replayed unless the evaluation harness can pin the decay origin to a fixed timestamp. Result caches must include a coarsened time bucket in the key, otherwise they serve yesterday's decay. Bucketising age to the hour or day fixes both at negligible relevance cost.
- When should recency not be applied at all?When the query has no temporal intent. Reference and how-to queries want the best explanation ever written, not the newest one, and decaying them systematically buries the canonical answer. If you can classify intent — by query pattern, vertical, or a temporal-intent model — gate the decay on it rather than applying one curve to all traffic.
Think of the freshness multiplier as a dimmer switch on the relevance score rather than a spotlight of its own: it can dim an old result, but it never lights up a document that has nothing to say.
saying these in an interview costs you the question
- Sorts results by date and calls it relevance tuning
- Adds a large constant bonus to recent documents
- Lets decay reach zero, hiding perfect older matches
- Applies one half-life to every query intent
- Recomputes freshness by reindexing the corpus nightly