How do you stop an agent's memory store filling with near-duplicate facts?
answer
- read before you insert
- canonical key, then similarity
- merge, do not append
- three sources, one record
- over-merging loses facts silently
basics
~20 sLook before you insert. Normalize each candidate fact into a canonical key, search the store for that key and for semantically similar records, then merge instead of adding: one record, one claim, with a link to every source that corroborated it.
solid answer
~50 sDeduplication is a read-before-write step. Normalize the extracted fact into a canonical form — resolve the entity, standardize the predicate, convert relative dates to absolute ones — and use that as a lookup key. Search the store by that key and, because paraphrases will not collide on it, by embedding similarity within the same scope. Hand the candidate plus the small set of matches to a resolver, which decides: insert a new record, merge into an existing one, or do nothing. Merging is the important case. Three call notes that all say the account renews in Q3 should end as **one** record with three provenance links, not three records — corroboration raises confidence and the extra sources stay auditable. Keeping the duplicates instead lets one fact crowd out others at recall time, and it means an update has to be applied three times or the store contradicts itself.
code
json · 13 lines{
"id": "acme:renewal_date",
"scope": {"account": "acme"},
"key": "account:acme/renewal_date",
"value": "2026-09-30",
"first_seen": "2026-02-11T10:04:00Z",
"last_confirmed": "2026-06-03T15:22:00Z",
"provenance": [
{"session": "s-8821", "message": 42, "stated_at": "2026-02-11T10:04:00Z"},
{"session": "s-9134", "message": 7, "stated_at": "2026-04-02T09:15:00Z"},
{"session": "s-9770", "message": 19, "stated_at": "2026-06-03T15:22:00Z"}
]
}go deeper
Know that inserting a memory should start with a lookup for something equivalent already in the store, and that identical facts should end up as one record rather than several.
Explain the pipeline: normalize into a canonical key, search by key and by semantic similarity within the same scope, then decide insert versus merge — and merge by keeping one record with all its sources.
Show the operational angle: idempotent writes so retries do not duplicate, batching the resolver call to control cost, and a conservative similarity threshold because over-merging destroys facts with no trace.
Own where the line sits: how much write-path latency and inference cost deduplication may spend per fact, whether the domain can support a controlled predicate vocabulary at all, and which error — a bloated store or a silently merged pair of distinct facts — your product can better afford.
## Why duplicates accumulate Any extraction pipeline that runs repeatedly over overlapping material will re-derive the same facts. A per-turn extractor sees the same standing preference mentioned twice. An end-of-session pass re-reads material a previous session already covered. Retries after a failure replay the same batch. Without a dedup step, the store grows roughly with conversation volume rather than with knowledge. ## Why duplicates are actually harmful It is tempting to shrug — storage is cheap. The costs are elsewhere: - **Recall crowding.** If a fact occupies three of the slots returned to the model, other relevant facts are squeezed out. The store's redundancy becomes the context's redundancy. - **Update fan-out.** When the fact changes, an update must find every copy. Miss one and the store now contradicts itself, and which version surfaces becomes luck. - **False corroboration.** Three copies of the same claim, derived from a single conversation, look like three independent confirmations to anything that scores by count. - **Audit noise.** Nobody can tell whether the fact is well-supported or merely re-extracted. ## The write path A practical insert looks like: normalize, look up, resolve, write. **Normalize.** Turn the extracted sentence into a canonical fact key. That usually means resolving the entity to a stable identifier rather than a name string, standardizing the predicate into a small controlled vocabulary where your domain allows it, and rendering values in a canonical form — absolute dates rather than "next quarter", currency and units normalized. Two extractions that mean the same thing must produce the same key as often as possible. **Look up.** Query the store for that key within the same scope — the same user, account or workspace, never across tenants. Because paraphrases and unusual phrasings will miss an exact key, add a semantic similarity search over recent and related records, and take the top handful above a threshold. **Resolve.** With the candidate and its matches in hand, decide. Exact key match on the same value is a merge. No match is an insert. A match on the key with a *different* value is not a duplicate at all — that is a contradiction, and it goes down the update path instead. Anything ambiguous is worth an LLM call: the model sees the candidate and the two or three near-matches and returns a decision. That call is cheap because the candidate set is tiny. **Write.** On a merge, keep the single record and append the new source to its provenance list, updating the last-confirmed timestamp. On an insert, write the record with its key, scope, value, timestamps and source. ## The corroboration case, concretely Three separate notes across three calls all state that an account renews in Q3. The right end state is one record — `account:acme / renewal_date / 2026-09-30` — carrying three provenance links, each naming the session, the message and the date it was stated. The fact is now more trustworthy *and* more compact, and when the renewal date moves there is exactly one record to update. The opposite policy — keep all three, let recall sort it out — is where stores rot. It looks harmless at ten records and is unmanageable at ten thousand. ## Thresholds and their failure modes Similarity-based dedup has two ways to be wrong and you must choose which you prefer. - **Too aggressive:** distinct facts get merged. "Prefers morning meetings" and "prefers Monday meetings" are semantically close and mean different things; collapsing them destroys information invisibly. - **Too lax:** near-duplicates survive and the store bloats. Over-merging is the more dangerous error because it is unrecoverable — the discarded fact leaves no trace — so most teams set a conservative similarity threshold and let the resolver model make the close calls rather than a bare distance cut-off. ## Cost, idempotency and batching Every write now costs a lookup, possibly an embedding, and sometimes a resolver call. On a high-volume system that is a real budget item, and it is another argument for extracting once per session rather than once per turn. Two practical mitigations: batch the resolution — hand the whole session's candidate facts to one resolver call along with their matches — and make the write idempotent, so a retried extraction after a timeout does not produce the very duplicates you are working to avoid. Deriving the record identifier from the normalized key and scope gives you that idempotency almost for free. ## What interviewers look for That you treat insert as read-before-write; that you distinguish a duplicate (same claim) from a contradiction (same subject, different value); that merging preserves every source rather than discarding them; and that you know over-merging silently destroys facts while under-merging merely wastes space.
- Same subject and predicate, different value — is that a duplicate?No. A duplicate is the same claim restated; the same key with a different value is a contradiction and belongs on the update path, where you decide whether the world changed or the earlier extraction was wrong. Routing it through the dedup path is how stores end up silently overwriting good facts or, worse, merging two incompatible values into one record.
- Why not simply store everything and let retrieval rank the duplicates down?Because duplicates crowd the small set of records that reach the model, they make count-based scoring mistake re-extraction for corroboration, and every future update has to find all the copies or the store contradicts itself. Retrieval cannot repair a store whose redundancy it cannot distinguish from evidence.
- How do you keep an extraction retry from creating the duplicates you are trying to prevent?Make the write idempotent. Derive the record identifier deterministically from the normalized fact key plus its scope, so replaying the same extraction lands on the same record and becomes a merge rather than an insert. Also record the source session and message on the provenance entry, so a replayed source is recognized and not counted twice as corroboration.
saying these in an interview costs you the question
- Assuming storage is cheap so duplicates are harmless
- Deduplicating on exact string equality of the raw sentence
- Merging by similarity alone with an aggressive threshold
- Counting repeated extractions of one statement as independent corroboration
- Discarding the older records' sources when merging