skip to content

Memory Retrieval Policies

Reading memory back: the agent fetching a fact through a memory tool when it needs it, versus pre-ranking by relevance, recency and validity time. Interviewers want a concrete scoring story here.

part ofAI agentsoverview, primer and where to startread it →
on this pageshow

questions

5

When is agentic memory recall better than pre-injecting top-k memories in an agent?

level: middleimportance: must knowfreq 62%

answer

  1. two read paths, not one
  2. who decides: pipeline or model
  3. unconditional cost versus extra round trip
  4. model-authored query beats raw user message
  5. pre-inject the profile, make the tail agentic

basics

~20 s

Agentic recall — the agent calling a memory tool when it notices it needs a fact — wins when most turns need no memory and precision matters. Pre-injected top-k wins for a small always-relevant profile and tight latency budgets.

solid answer

~50 s

There are two read paths. **Pre-injection** runs a similarity search before the model is called and pastes the top k memories into the prompt every turn: one round trip, fully deterministic, but it fires whether or not the turn needs memory, uses the raw user message as the query, and dilutes context with near-misses. **Agentic recall** exposes the store as a tool — search, read, or a memory directory the agent can list and grep — and the model decides mid-turn that it needs a fact, writes its own query, and can refine or widen it. The trade is latency and determinism against precision and conditional cost. My default is to pre-inject a few hundred tokens of always-relevant profile and make everything else agentic, because on long-horizon work most turns need nothing and blind top-k is mostly noise. The risk agentic recall adds is a model that never looks, so the prompt must say memory exists, ideally with a cheap index of what it holds.

go deeper

for a junior

Know that stored memories reach the model in one of two ways: retrieved automatically before the call, or fetched by the agent through a tool. Be able to say which one a system you worked on used.

for a middle

Explain the mechanics of both paths and their costs — one round trip and unconditional tokens versus conditional cost, a model-authored query and an extra turn. Name fixed-k's inability to return nothing as a concrete weakness.

for a senior

Show you have tuned this on a live system: what fraction of turns actually used an injected memory, how you split a small pre-injected profile from an agentic tail, and how you caught an agent that never called the memory tool.

for a principal

Own the policy across surfaces with different latency budgets and the evaluation cost it implies. Argue when determinism is worth more than precision, and set the organizational default plus the exceptions rather than picking per feature.

## Two ways a memory reaches the prompt **Pre-injection (retrieve-then-prompt).** Before the model is invoked, the harness takes the user's latest message, searches the memory store, keeps the top k results — usually five to ten — and places them in the system prompt or a dedicated memory block. The model never asks for them; they are simply there. This is the classic pipeline shape borrowed from retrieval-augmented question answering. **Agentic recall.** The memory store is exposed to the model as a tool: a search call, a read call, or a plain directory the agent can list, open and grep. Mid-turn, the model notices it lacks a fact, issues a call, reads the result and continues. Anthropic's file-directory memory tool and Letta's filesystem-style memory are both this shape, and by mid-2026 it is the dominant pattern for long-horizon agents, with blind top-k pre-injection retreating to the small always-on profile. ## What pre-injection buys, and what it costs It buys a single round trip, which matters enormously when the product has a hard latency budget — a voice assistant cannot afford a recall turn before it starts speaking. It is deterministic and therefore easy to evaluate and replay: the same conversation state yields the same injected memories. It costs on three fronts. First, it is unconditional: it pays retrieval cost and context tokens on every turn, including the large majority of turns that need no memory at all. Second, the query is whatever the user just typed, which is often a terrible retrieval query — "and the other one?" carries no searchable content, so the top k is effectively random. Third, fixed k has no notion of "nothing relevant here": if the store holds no useful memory, top-k still returns k items, and those near-misses sit in the window as plausible-looking noise. That is the mechanism behind context rot and lost-in-the-middle degradation, and it is why more memory can make an agent worse. ## What agentic recall buys, and what it costs Cost becomes conditional — nothing is spent on turns that need nothing. The query is model-authored, so it can be a resolved, explicit query ("the appointment-time preference recorded for this patient") rather than the user's fragment. The model can iterate: search, find nothing, widen the query, or list a directory and open only the two files that look right. It can also filter and aggregate before anything enters context, which is exactly what makes a grep over a `memory/patients/` directory cheaper than a pre-injected top-8: the agent reads one file instead of eight snippets. And the model can conclude "I have no memory of this" and say so, which fixed-k retrieval structurally cannot express. The costs are real. Every recall is an extra model turn plus a tool round trip, so a turn that needs two lookups can triple its latency. The model may fail to look at all — an agent cannot retrieve a fact it does not suspect exists, and this silent miss is the characteristic failure of agentic recall. It is also nondeterministic, which makes regression testing harder: two runs of the same conversation may retrieve different things. ## Choosing, in practice A workable default is a split by size and hit rate. Pre-inject the small, high-hit-rate core — identity, language, standing preferences, active goals — because a few hundred tokens that are relevant most turns are effectively free and save a round trip. Make everything else agentic: the long tail of episodic detail, per-entity files, older sessions. Then tune with two numbers. If a large fraction of turns actually consume an injected memory, injection is earning its tokens. If the fraction is small — and it usually is — agentic recall is cheaper and cleaner. On the other side, if the agent frequently answers as though a stored fact does not exist, the problem is discovery, not the store: fix it by pre-injecting an index or table of contents of what memory holds, so the model knows what is worth asking for. Latency decides the edge cases. Real-time voice and autocomplete cannot absorb a recall turn; asynchronous agents doing multi-minute work absorb several without anyone noticing. ## Failure modes worth naming Blind top-k on a memory-irrelevant turn, poisoning an otherwise clean context. A model that never calls the memory tool because nothing in the prompt suggests memory exists. Recall loops, where the agent searches repeatedly with near-identical queries and burns its budget. And unbounded results, where an agentic read pulls a whole file that a pre-injected snippet would have trimmed — agentic recall removes the k cap, so the tool itself must impose size limits.

  • Your agent has memory available as a tool but keeps answering as if it has none. What do you change?
    That is a discovery failure, not a retrieval failure. Say in the system prompt that persistent memory exists and when to consult it, and pre-inject a cheap index — the directory listing or a one-line summary per memory file — so the model can see that something relevant is stored. Failing that, add a lightweight trigger: on turns matching known memory-dependent intents, force a recall call before answering.
  • How would you keep agentic recall from adding two seconds to every turn?
    Cap it structurally. Pre-inject the small always-relevant profile so common turns need no lookup at all; allow parallel recall calls so several lookups cost one round trip; cap recall calls per turn and return a bounded payload from the tool. For latency-critical surfaces, run a speculative pre-injection concurrently with the first model call and let the agent ignore it when it is not needed.
  • Does agentic recall make evaluation harder, and how do you handle that?
    Yes — the retrieval set is no longer a deterministic function of the input, so the same conversation can take different paths. Handle it by evaluating the outcome rather than the retrieved set, running each memory-dependent case several times and requiring consistent success, and recording the recall calls in the trace so a failure can be attributed to a missed lookup versus a bad answer given a correct lookup.

Pre-injection is handing someone a stack of possibly-relevant files before every meeting; agentic recall is letting them walk to the cabinet when a question actually comes up.

saying these in an interview costs you the question

  • Top-k always returns something useful because it returns the nearest neighbours
  • Injecting more memories per turn makes the agent smarter
  • Agentic recall costs the same as pre-injection, just later
  • The user's raw message is always a good retrieval query
  • Memory in the prompt is free once you have a long context window

context

open as a page

How do scope and metadata filters on agent memory reads prevent cross-user leakage?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Every memory read must be constrained by a scope key — user, tenant, session or project — derived from the authenticated session, applied inside the query before ranking. Semantic search has no notion of ownership, so nothing but an explicit filter keeps one user's memories out of another's context.

open as a page

Why do agent memory retrievers score recency and importance alongside relevance?

level: middleimportance: should knowfreq 54%

basics

~20 s

Semantic similarity alone surfaces old, trivial memories that merely resemble the query. Adding a recency term favours what the agent learned lately, and an importance term favours consequential facts over small talk, so the top few slots go to memories that change the answer.

open as a page

How should agent memory retrieval handle facts that were only true for a period?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Store a validity period on each fact and filter at read time against the moment being reasoned about, so superseded facts never enter context. Ranking cannot fix this: an expired fact can be the most semantically relevant one in the store.

open as a page

How would you measure whether an agent's memory retrieval is actually helping?

level: principalimportance: should knowfreq 38%

basics

~20 s

Measure usefulness, not recall. Track how often a retrieved memory is actually used in the reply, whether turns that used memory produced better outcomes than the same turns with memory disabled, and how often a retrieved memory made the answer wrong.

open as a page