skip to content

Why do agent memory retrievers score recency and importance alongside relevance?

level: middleimportance: should knowfreq 54%

answer

  1. similarity is not usefulness
  2. three terms, not one
  3. time since last access decays
  4. the model rates it at write time
  5. weights are a product decision

basics

~20 s

Semantic similarity alone surfaces old, trivial memories that merely resemble the query. Adding a recency term favours what the agent learned lately, and an importance term favours consequential facts over small talk, so the top few slots go to memories that change the answer.

solid answer

~50 s

Relevance is a similarity measure, and similarity is not usefulness. A hundred near-identical logged observations all match a query equally well, and the oldest, most trivial one can win by a rounding error. The standard fix, popularized by the Generative Agents work, is a composite score: normalized semantic relevance, plus a recency term that decays exponentially with time since the memory was last accessed, plus an importance term the model assigns when the memory is written. Each term corrects a different failure. Recency breaks ties toward the current state of the world and demotes stale detail. Importance rescues the rare consequential memory — "the patient is allergic to penicillin" — that a chatty query would otherwise bury under scheduling chatter. In practice you tune the weights per product: an assistant tracking changing preferences leans on recency, a compliance-facing one leans on importance. Watch for the failure where recency alone starts overwriting durable facts with noise from the last session.

go deeper

for a junior

Be able to name the three common scoring terms — relevance, recency and importance — and say in one sentence what each contributes beyond the others.

for a middle

Explain why similarity alone ties and stales, how exponential decay on last access behaves, and where the importance score comes from. Expect to be asked how you would weight them for a given product.

for a senior

Show you have inspected real retrieval sets: recency crowding out durable facts, importance inflating toward the top of the scale, near-duplicates consuming the top slots, and what you changed in response.

for a principal

Own the position that no formula has won. Argue the trade between a tuned composite score, a temporal graph and a filesystem the agent searches, and set what the organization measures before it tunes anything.

## Why relevance alone ranks badly Semantic relevance answers one question: how close is this memory to the query, in meaning? That is necessary and nowhere near sufficient, because a memory store is not a document corpus. It is a pile of observations about one user or one task, many of which are near-duplicates of each other, accumulated over months, and differing mostly in when they happened and how much they matter. Three failures follow directly. **Ties.** Fifty logged observations about appointment scheduling are all roughly equally similar to "when should I book them?", so which five surface is decided by embedding noise. **Staleness.** A precise, well-phrased memory from a year ago often out-scores a terse recent one that supersedes it, because phrasing drives similarity and time does not enter the score at all. **Triviality.** A throwaway remark that happens to share vocabulary with the query beats a consequential fact phrased in different words. ## The composite score The canonical formulation comes from the Generative Agents simulation, and it remains the reference answer in interviews. Each candidate memory receives three sub-scores, each normalized to a common range, then combined as a weighted sum: - **Relevance** — semantic similarity between the query and the memory text. - **Recency** — an exponential decay over the time since the memory was last retrieved, so recently-used memories stay warm and untouched ones fade. - **Importance** — a score the model assigns at write time when asked how consequential the memory is, on a fixed scale; mundane observations land low, life-changing ones land high. The top-scoring memories are then what enters context. In the original work the three weights were all set equal; in a product you tune them, and the tuning is where the judgment lies. ## What each term is actually for **Recency encodes that the world changes.** An agent that assists a person is modelling a moving target — preferences, projects, team membership. Decay is a cheap prior that recent observations describe the present better than old ones. Note the subtlety in decaying by *last access* rather than creation: a fact retrieved every week stays warm even if it was recorded a year ago, which is a reasonable proxy for "still in play". **Importance encodes that not all facts are equal.** Without it, retrieval is a popularity contest among whatever the user talked about most. Importance is the term that keeps a rare safety-critical fact reachable from a query that does not lexically resemble it. It is also the most expensive term, because someone — usually a model call at write time — has to produce it, and model-assigned importance is noisy and drifts between model versions. **Relevance keeps the whole thing on topic.** Drop it and you get the agent's greatest hits regardless of what was asked. ## Tuning and its traps The weights are a product decision. A scheduling or preferences assistant weights recency heavily, because being right about *now* dominates. A clinical or compliance-facing assistant weights importance heavily and dampens decay, because a durable fact must not fall out of reach just because it has not come up lately. An agent doing a long single task weights relevance heavily, because everything in the store is from the same session and time barely separates it. Two traps recur. The first is **recency crowding out durability**: crank decay too high and last week's noise reliably outranks a standing fact, which reads to the user as the agent forgetting something it clearly knows. The second is **importance inflation**: ask a model to rate importance and it will drift upward over time and cluster near the top of the scale, flattening the term into a constant. Both are detectable only by looking at what actually gets retrieved, which is why you sample real retrieval sets and read them rather than trusting the formula. ## Beyond the three-term score Production systems add terms as their failure modes demand. A **redundancy penalty** stops five paraphrases of the same fact from consuming all the slots — you want the top five to be five different things, not one thing five times. A **confidence or provenance** term down-weights facts inferred by the model relative to facts the user stated outright. A **validity** term suppresses facts whose stated period has ended, which is a filter rather than a soft score. And graph-structured memories bring their own signals: how central a node is, and how many edges connect it to the entity under discussion. The honest position in 2026 is that no single scoring formula has won. Filesystem-style memories, temporal knowledge graphs and vector stores with composite scores all post competitive results on long-conversation benchmarks, and the practical answer is usually hybrid and routed. What an interviewer wants is that you can state the three terms, say what each one fixes, and describe how you would tune them for a named product rather than reciting the formula.

  • Recency decays by time since last retrieval rather than time since creation. Why does that matter?
    Decaying by last access treats repeated use as evidence that a fact is still live, so a year-old preference the agent consults weekly stays reachable, while a recent but never-used observation fades. Decaying by creation instead models pure ageing, which punishes durable facts that simply have not come up. The access-based version approximates relevance-in-practice; the creation-based one approximates freshness. Most assistants want the former, audit-style systems often want the latter.
  • Your importance scores have all drifted toward the top of the scale. What went wrong and what do you do?
    Model-assigned importance inflates: given a fixed scale and no calibration, an LLM rates most things as fairly consequential, and the term collapses into a constant that no longer discriminates. Fix it by anchoring the rubric with concrete examples at each level, scoring in batches so the model rates relatively rather than absolutely, or normalizing scores within the store so the term is a percentile rather than a raw value. Re-calibrate whenever the scoring model changes.
  • Why add a redundancy penalty on top of relevance, recency and importance?
    Because those three terms score each memory independently, so a fact recorded five times in slightly different words takes five of the top slots and crowds out four other things the agent needed. A redundancy penalty scores the set rather than the item: after picking a memory, down-weight candidates that are near-duplicates of what is already selected. It buys coverage without enlarging k.

Ranking memories by similarity alone is like picking what to tell a friend based purely on topic match — you would also weigh how recently it happened and how much it mattered.

saying these in an interview costs you the question

  • Higher cosine similarity always means a more useful memory
  • Recency is just sorting by timestamp descending
  • Importance can be inferred at read time from the query
  • Equal weights are correct because the paper used them
  • Adding more scoring terms always improves retrieval quality

context