How should agent memory retrieval handle facts that were only true for a period?
answer
- two clocks, not one timestamp
- when we learned it versus when it held
- filter before you rank
- the expired fact often matches best
- as-of instant defaults to now
basics
~20 sStore a validity period on each fact and filter at read time against the moment being reasoned about, so superseded facts never enter context. Ranking cannot fix this: an expired fact can be the most semantically relevant one in the store.
solid answer
~60 sTreat validity as a filter, not a score. Each memory carries a period during which the fact held — a valid-from and an optional valid-to — separate from when it was recorded. At read time the retriever takes an as-of instant, normally now, and drops anything whose period does not cover it before ranking begins. The reason it cannot be left to scoring is that a superseded fact is often the *best* semantic match. A clinic assistant whose store says a patient preferred morning appointments from March 2024 until they changed jobs in November 2025 will rank that memory top for "when do they like to come in?", and no recency weighting reliably beats it if the replacement is terse. Keeping both clocks also buys real capability: you can answer "what did we believe last spring?" by moving the as-of instant, and you can audit why the agent acted as it did. Superseded facts are usually hidden by default and surfaced only when the agent is explicitly reasoning about history.
code
python · 16 linesfrom datetime import date
MEMORIES = [
{"fact": "prefers morning appointments",
"valid_from": date(2024, 3, 1), "valid_to": date(2025, 11, 15)},
{"fact": "prefers late-afternoon appointments",
"valid_from": date(2025, 11, 15), "valid_to": None},
]
def holding(memories, as_of):
return [m for m in memories
if m["valid_from"] <= as_of
and (m["valid_to"] is None or as_of < m["valid_to"])]
print([m["fact"] for m in holding(MEMORIES, date(2026, 8, 19))])
print([m["fact"] for m in holding(MEMORIES, date(2025, 1, 10))])go deeper
Know that a stored fact can stop being true and that a memory should carry when it applied, not just when it was written. Say plainly that the agent should not be shown facts that no longer hold.
Explain the two clocks — when the system learned a fact versus the period it held — and why validity belongs in a pre-ranking filter rather than in the score. Describe the as-of instant defaulting to now.
Show judgment about where validity times come from, when to surface a superseded fact with explicit period annotations instead of hiding it, and the operational traps: everything left open-ended, a wrong as-of default in replays, silent expiry with no replacement.
Own the modelling decision. Argue when a bi-temporal store or temporal graph is worth its complexity against a flat store with write-time supersession, and what the organization owes users in auditability of what the agent believed when.
## Two clocks, not one Every memory has at least two timestamps and they are not the same thing. **Transaction time** is when the system recorded the fact. It is what a naive store keeps, and it answers "when did we learn this?" **Valid time** is the period during which the fact was true in the world — a valid-from, and a valid-to that is open until something supersedes it. It answers "when was this the case?" They diverge constantly. A patient tells the clinic in June that they changed jobs back in November; transaction time is June, valid time starts in November. An agent that only knows transaction time cannot distinguish a fact learned late from a fact that became true late, and both mistakes produce confidently wrong answers. Systems built on temporal knowledge graphs make this bi-temporal model explicit and it is a large part of why they lead on long-conversation benchmarks that test reasoning over changing facts. ## Why ranking cannot substitute for a filter The instinct is to lean on recency decay: newer memories score higher, so the current fact wins. It does not hold, for three reasons. First, a superseded fact is frequently the strongest semantic match. It was recorded when the topic was actively discussed, so it is phrased in exactly the vocabulary of the query, while its replacement may be a two-word correction. Second, decay is soft. A scoring nudge changes an ordering; it does not remove an item. With any k above one, both the expired and the current fact enter context, and the model is left to adjudicate a contradiction it has no evidence to resolve. Third, transaction-time recency is the wrong clock. A backfilled historical fact recorded yesterday looks maximally fresh under decay while describing something that stopped being true a year ago. So validity belongs in the filter stage, applied before ranking, exactly like a scope filter: choose an as-of instant, keep only memories whose valid period covers it, then score what remains. ## Where validity times come from This is the hard part in practice, and honesty about it is what an interviewer is listening for. Some facts carry explicit periods because the source has them — an employment record, a subscription, a prescription. Some are dated by the utterance that produced them, so valid-from is the conversation timestamp and valid-to stays open. Some are closed by a later contradicting fact: when a new preference is recorded, the previous one's valid-to is set to the new fact's valid-from. And some have no honest validity at all — an inference the model drew from tone — and are better stored with low confidence than with an invented period. Note the division of labour. Deciding *at write time* that a new fact supersedes an old one, and stamping the closing timestamp, is a write-policy concern. What retrieval owns is honouring those periods on the way out, defaulting to now, and never silently returning a fact whose period has closed. ## Suppress, or surface with a marker? Default to suppression: if the fact does not hold now, it does not belong in context, because the model will treat anything in its context as usable. But two cases want the expired fact back. When the agent is reasoning explicitly about history — "why did we book them at 9am last year?" — the as-of instant should move to the time under discussion rather than now. When the agent must explain a change, return both with explicit period annotations, formatted so the model cannot miss which is current: state the periods inline rather than relying on ordering, because ordering carries no meaning to a model reading a flat list. A third case is contested facts, where two memories claim overlapping periods. Retrieval should return both with their periods and provenance rather than picking a winner, and the agent should ask rather than guess. ## Failure modes to watch **Open-ended everything.** If nothing ever gets a valid-to, the filter is a no-op and you are back to hoping recency saves you. Detect it by measuring what fraction of retrieved facts about mutable attributes have a closing timestamp. **Wrong as-of default.** Batch jobs and replays that pass the current time while reprocessing last quarter's conversation will retrieve facts that did not exist then. The as-of instant should be an explicit parameter of the read. **Timezone and granularity sloppiness.** Day-granularity periods with local timezones produce off-by-one-day windows where a fact is neither old nor current. **Silent expiry.** A fact whose period ended with no replacement leaves the agent with nothing where it used to have something. That should be visible — the agent should be able to say "I knew their preference until last November and have not been told since" rather than behaving as though the subject never came up.
- Why not just let recency weighting demote the superseded fact instead of filtering it out?Because decay changes an ordering rather than removing an item, so with any k above one both the expired and the current fact reach the context and the model must adjudicate a contradiction with no evidence. Worse, the expired fact is often the better semantic match, having been recorded when the topic was discussed at length, and a backfilled historical fact looks fresh under transaction-time decay while describing something long over.
- How do you support a question like "what did we think their preference was last spring?"Make the as-of instant an explicit parameter of the read rather than hardcoding now. Answering a historical question means running the same validity filter with the as-of moved to that date, which returns the fact that held then. If the question is about what the system believed rather than what was true, you query on transaction time instead — that is why keeping both clocks matters.
- What do you do when two memories claim overlapping validity periods for the same attribute?Return both with their periods and provenance rather than silently picking one. Retrieval is the wrong layer to resolve a genuine contradiction: it has no evidence beyond timestamps, and choosing wrong is worse than surfacing the conflict. The agent should either ask the user to confirm, prefer the higher-provenance source if policy defines one, or act conservatively. Resolving the conflict permanently is a write-side decision.
saying these in an interview costs you the question
- Recency decay is enough to bury facts that stopped being true
- One created-at timestamp is all a memory needs
- The model will notice the contradiction and pick the newer fact
- Expired memories should be deleted rather than closed off
- Validity filtering can be applied after ranking without cost