In agent memory, how do you decide which facts from a session are worth storing?
answer
- fewer facts, better facts
- would it change a later decision
- the model already knows the world
- one wrong fact never announces itself
- four from forty minutes is success
basics
~20 sStore only facts that are durable, specific to this user or account, and likely to change a future decision. Skip general world knowledge the model already has, and skip anything provisional. A wrong stored fact costs far more than a missed one.
solid answer
~50 sThree filters do most of the work. **Durability** — will this still be true and useful next session? A stated preference or a contract date passes; "I'm on the train right now" does not. **Specificity** — is it about this user, account or project? General world knowledge is already in the model's weights and storing it only crowds the store. **Actionability** — would knowing it change what the agent does or says later? If not, it is trivia. Then there is the asymmetry that drives the tuning: a missed fact costs one awkward re-ask, while a wrong stored fact silently steers every future session and nobody sees the source. So extraction is tuned for precision, not recall. Forty minutes of a discovery call producing four solid facts is a good outcome — the failure mode to fear is forty facts, half of them wrong.
go deeper
Be able to say that not everything in a conversation should be remembered, and give the test: is it durable, is it about this specific user, would it change a future answer?
Explain what a well-formed record looks like — atomic, self-contained after pronouns and relative dates are resolved, scoped to a user, and carrying its source and timestamps.
Lead with the asymmetry: a missed fact costs one re-ask, a wrong fact silently misleads every later session. Show how you tune extraction for precision and how provenance lets you contain a bad batch.
Own the policy layer: which categories may never be written in a regulated domain, whether users can inspect and delete their memories, how extraction quality is measured over sampled traffic, and how untrusted content is prevented from driving writes.
## Why this is the hard half of memory Storing a fact is trivial engineering. Deciding it deserves to outlive the session is a judgment call made hundreds of times a day by a model, unsupervised, with no one reading the output. That is why the interview question is about *criteria* rather than storage. ## The three filters **Durability.** Ask whether the fact will still be true and still be useful when the agent next sees this user. Preferences, constraints, identities, commitments, and structural facts about an account pass. Transient state — where someone is right now, what they are about to do in the next five minutes, the intermediate result of the current task — does not. Some facts are durable but time-bounded; those get stored with an expiry rather than dropped. **Specificity.** Memory is for what the model cannot already know. If a statement is general world knowledge, storing it wastes a retrieval slot and adds nothing: the weights already carry it. The test is whether the fact is about *this* user, account, project or environment. **Actionability.** Would this fact change a future response, plan or tool call? "Prefers all figures in EUR" changes output. "Mentioned they like hiking" usually does not, unless the product is about hiking. Teams that skip this filter end up with stores full of true, useless facts that dilute every recall. ## The asymmetry that sets the operating point Missing a fact and storing a wrong one are not symmetric errors. - A **missed** fact produces one visible, recoverable failure: the agent asks again, the user repeats themselves, mild friction. - A **wrong** fact is silent and compounding. It gets recalled into future sessions as ground truth, the model has no way to doubt it, and the user sees a confidently wrong agent with no visible cause. This is the memory-poisoning failure mode, and it is worse when the wrong fact came from untrusted content the agent read rather than from the user. So you tune the extractor toward precision. Instruct it to emit nothing when uncertain; "no memory" is always a valid extraction result. Ask for a confidence signal and drop low-confidence proposals. Prefer fewer, better-formed facts. ## What a well-formed stored fact looks like Extraction is not quoting; it is normalization. A good record is: - **Atomic** — one claim per record, so it can be updated or expired independently. "Renewal is in Q3 and Dana is the CTO" is two records. - **Self-contained** — it resolves pronouns and relative dates. "Renewal in Q3" becomes "renewal due 2026-09-30"; "she approves budget" names the person. A memory that only makes sense next to its original turn is not a memory. - **Scoped** — attached to a user, account or workspace, so it can never surface for the wrong party. - **Attributed** — carrying provenance: which session and message it came from, when it was stated, when it was written, and which extractor produced it. Provenance is what lets you audit a bad memory, trace it back, and delete the whole batch when an extractor turns out to be faulty. ## Worked shape: a discovery call A sales assistant processes a forty-minute call. The transcript contains thousands of words: pleasantries, a demo walkthrough, three tangents about the weather, and an unrelated internal complaint. Good extraction returns something like four records — the renewal window, who signs off on budget, the competing product under evaluation, and the integration constraint that blocks a deal. Everything else is either transient, general, or not actionable. If the same call produced forty records, the store would grow faster than anyone can validate it, dedup pressure would spike, and retrieval quality would fall for every other fact about that account. ## Guarding against poisoned writes When an agent reads documents, web pages or tool output, that content is untrusted input. Text encountered there can attempt to instruct the extractor to store something. Practical defences: only extract from turns you trust (user statements and verified tool results) rather than arbitrary retrieved text, keep provenance so a suspicious memory can be traced to its source, and treat memory writes originating from external content as a lower-trust class — or forbid them. ## Human and product controls Because extraction is imperfect, mature products expose the store: users can view, edit and delete what the agent remembers about them. That is partly a trust feature and partly an evaluation channel — user deletions are a direct signal about extraction precision. In regulated domains, extraction may also need a category filter so that sensitive attributes are never written at all, regardless of durability or usefulness. ## What interviewers listen for Named criteria rather than vibes, the precision-over-recall asymmetry stated explicitly, atomic self-contained records with provenance, and awareness that untrusted content can drive a write.
- How do you tell the extractor to prefer precision without it simply writing nothing at all?Give it explicit positive categories to look for — preferences, constraints, commitments, identities, account structure — plus an explicit instruction that emitting no memory is a valid outcome. Then measure: sample real sessions, hand-label the facts that should have been captured, and compare. Precision tuning without a labelled sample is guesswork, and the failure mode of an over-tightened extractor is a store that never grows.
- A memory turns out to be wrong and has been shaping answers for weeks. How do you contain it?Provenance makes this tractable: the record names its source session, message and extractor version, so you delete the record, look for siblings written by the same pass, and check whether the same extractor version produced similar errors elsewhere. Without provenance you can only delete the one record you happened to notice and hope it was isolated.
- Should a fact the agent inferred rather than one the user stated be written to memory?Only with the inference marked as such and a lower trust level. Inferred facts are the most common source of confidently wrong memories, because the extractor has no way to distinguish its own guess from a statement later on. Store the observation that supports it where you can, and prefer inferences that are cheap to re-derive over ones the agent will treat as ground truth.
saying these in an interview costs you the question
- Storing general world knowledge the model already has in its weights
- Treating memory recall as more important than memory precision
- Writing verbatim message text instead of normalized, self-contained facts
- Bundling several claims into one record so nothing can expire independently
- Extracting memories from untrusted retrieved content without any trust distinction