skip to content

How do you stop an agent's long-term memory store from growing into noise?

level: principalimportance: should knowfreq 33%

answer

  1. storage is cheap, context is not
  2. default-deny on writes
  3. denser, not longer
  4. one global lifetime is always wrong
  5. measure the store or guess

basics

~20 s

Persist by exception rather than by default, consolidate asynchronously between sessions to merge duplicates and promote repeated patterns, and expire facts by class instead of one global lifetime. Growth is a policy problem: without one, every store eventually costs more than it returns.

solid answer

~50 s

Treat memory as a curated asset with an admission bar, not a log. Three levers do the work. **Admission**: persist only facts that are durable, reusable and would change a future answer — most conversational content fails all three. **Consolidation**: run an asynchronous pass between sessions that merges near-duplicates, resolves contradictions and promotes repeated episodes into a single semantic fact or a reusable procedure, so the store gets denser rather than longer. **Expiry**: set lifetimes by fact class — a dispatcher's language preference is effectively permanent, while "currently covering the north depot" is stale in a week — because one global TTL is either too aggressive for identity facts or useless for volatile ones. Then instrument it: store size per user, duplicate and contradiction rate, and the age distribution of what gets injected. Without those numbers you cannot tell a rich memory from an expensive one.

go deeper

for a junior

Know that a memory store cannot just grow forever: old and duplicated facts eventually crowd out useful ones and can be re-injected as if they were still true.

for a middle

Explain the three levers — an admission bar on writes, a background consolidation pass that merges and promotes, and expiry set per class of fact — and why one global lifetime fails in both directions.

for a senior

Show that you would instrument the store: size per user over time, duplicate and contradiction rate, age of what actually gets injected. Be ready to explain why consolidation runs offline and what a bad merge costs.

for a principal

Own the steady state and the tradeoffs: what the product promises to remember, what it deletes, how scope bounds blast radius, and how you would prove with an offline replay that memory is net positive rather than an expensive feature nobody measured.

## The failure this prevents A memory store that only ever grows has a predictable arc. In month one it feels magical. By month six a heavy user has thousands of entries, most of them stale operational trivia, several pairs of them contradictory. Every injection now carries more marginal and conflicting material, so answers get *worse* while cost goes up — the system is paying more to be less reliable. Nobody notices, because there is rarely a metric that would show it. So the principal-level question is not "how do we store more" but "what is the steady state". A store needs an equilibrium between what enters and what leaves, and that equilibrium is a policy decision with product, cost and privacy consequences. ## Lever one: admission The cheapest memory to manage is the one never written. A useful admission test asks three things of every candidate fact: - **Is it durable?** Will it still be true next month? "Prefers metric units" yes; "is looking at the Tuesday manifest" no. - **Is it reusable?** Will it apply outside this one task? A one-off calculation is not memory, it is a result. - **Would it change a future answer?** If injecting it would produce the same response as omitting it, it is decoration. A logistics dispatch assistant should persist "this operator plans in local depot time, not UTC" and should not persist "asked about route 14 on Tuesday". The instinct to keep everything "in case it's useful" is exactly what produces the month-six store. ## Lever two: asynchronous consolidation Even with a good admission bar, entries drift toward redundancy: five separate observations that all imply one underlying preference, two entries that contradict because the world changed. Consolidation is a background pass — run between sessions, on a schedule, or when a user's store crosses a size threshold — that rewrites the store rather than appending to it. It does three things. It **merges** near-duplicates into one canonical entry. It **resolves** contradictions by recency and provenance, keeping one current version. And it **promotes**: when the same episodic pattern recurs — the operator always wants the exceptions list before the summary — that becomes a single semantic fact or, if it is a repeatable multi-step routine, a procedure stored once and loaded when relevant. Promotion is the lever that makes a store denser instead of longer, and it is only affordable because it runs offline, where a slower and more expensive model call is fine. ## Lever three: expiry by class A single global TTL is always wrong. Thirty days deletes a user's language preference; two years keeps "currently covering the north depot" long after it stopped being true. Segment by volatility: - **Identity and preference facts** — effectively permanent, removed only by user action or explicit contradiction. - **Relationship and configuration facts** — long-lived but re-verifiable; worth a periodic check rather than a hard expiry. - **Operational state** — short lifetime by construction. Often this class should not be in long-term memory at all; it belongs to the session. - **Episodic records** — keep an event log if the product needs history, but consider whether it should be reachable at injection time or archived out of the way. The useful reframing is that expiry is not deletion for storage reasons — storage is cheap. It is deletion so that a stale fact cannot re-enter a context window and be presented as current. ## Instrument it or you are guessing Minimum viable telemetry: store size per user over time (is it converging or compounding?), duplicate and contradiction rate after consolidation, and the age distribution of entries that actually reach a prompt. If injected facts skew old and the store grows linearly with usage, consolidation is not keeping up. Pair that with an offline eval that replays real sessions with and without memory — the only way to answer whether the store is net positive at all, which is a question surprisingly few teams ask after launch. ## The tradeoffs to own This is contested ground and an honest answer says so. Aggressive expiry loses genuinely useful long-tail facts and produces the "it forgot again" complaint. Aggressive consolidation loses nuance and can launder a wrong inference into a confident canonical fact — a bad merge is harder to notice than a duplicate. Doing nothing is the worst option but has the best short-term demo. There are also non-technical constraints. Scope determines blast radius: a fact promoted from one user's session into a shared or global store is a leakage vector as well as a correctness risk. And retention interacts with deletion obligations — a store you can enumerate and purge per user is the reason to keep memory in an inspectable store in the first place, so consolidation must preserve provenance rather than melting facts into anonymous prose. The defensible position: default-deny on writes, consolidate offline, expire by class, measure the store, and accept that the right settings are found empirically per product rather than copied from anyone else.

  • How would you know whether the memory store is a net positive at all?
    Replay a sample of real sessions offline, once with memory injected and once without, and score the outputs on task-specific criteria. That gives a direct answer instead of an intuition. Segment the result — memory often helps returning heavy users and does nothing for one-off sessions — because an aggregate that nets to zero can hide a strong effect in the segment the feature was built for.
  • What is the risk of an aggressive consolidation pass?
    Bad merges. Collapsing several observations into one canonical fact can launder a wrong inference into something the system now states confidently, and the evidence that would have contradicted it has been deleted. Duplicates are visible and annoying; a confidently wrong merged fact is invisible. Keep provenance on merged entries, be conservative when the sources disagree, and prefer leaving two entries over inventing a third.
  • Where does scoping fit into a growth policy?
    Scope bounds blast radius. Per-session facts expire naturally; per-user facts are the main store and are individually deletable; anything promoted to a team or global scope is now a shared belief that a single bad write can propagate to everyone. Promotion across scopes deserves a much higher bar than a normal write — and in most products, a human in the loop.
  • Users complain the assistant forgot something after you tightened expiry. How do you respond?
    Treat it as evidence about class boundaries, not as proof the policy is wrong. Look at what was dropped: if it was an identity or preference fact, it was misclassified as volatile and the class definition needs fixing. If it was genuinely operational state the user expected to persist, that is a product question about what the assistant promises to remember, and the honest fix may be an explicit user-controlled pin rather than a longer global lifetime.

saying these in an interview costs you the question

  • Says storage is cheap so nothing needs deleting
  • Applies one global TTL to every kind of fact
  • Consolidates synchronously and calls memory slow
  • Never measures whether memory improves outputs
  • Promotes a single session's fact to a shared scope automatically

context