skip to content

What does a stored memory entry's provenance prove when an extractor wrote it during the user's own session?

level: middleimportance: should knowfreq 54%

answer

  1. who wrote it, not who meant it
  2. the stamp is truthful and useless
  3. the session really was the user's
  4. the missing field is the source span
  5. nothing anomalous to find

basics

~20 s

It proves which component wrote the entry and in which session, and nothing else. Both stamps are truthful for a claim lifted from a stranger's email, so provenance cannot separate attacker-seeded facts from ones the user actually stated.

solid answer

~50 s

The provenance on an auto-extracted memory entry names the writer and the session, and in this construction both are genuine: the extractor really did write it, and the session really was the legitimate user's, because the write happened while that user was triaging their own mail. What the record does not carry is the *source text* - which message in that session the claim was lifted from, and whether that message was authored by the account owner or by an unknown sender. So the entry a stranger seeded and the entry the user dictated are byte-for-byte indistinguishable in every field that exists. That is why removal is harder than injection here: there is no session to revoke, no attacker identity in the record, and no evidence either way about any individual claim. Provenance answers 'which session wrote this', never 'who meant it'.

code

json · 11 lines
json
{
  "memory_id": "m_4471",
  "claim": "[extracted standing-preference claim - text elided]",
  "written_by": "session-summary-extractor",
  "session_id": "s_20a3f9",
  "session_owner": "user_1182",
  "written_at": "2026-03-04T17:22:08Z",
  "source_message_id": null,
  "confidence": 0.71
}
... 3,411 further entries, identical in shape

go deeper

for a junior

Know that a memory entry records who wrote it and when, and that this is not the same as who came up with the claim. Be able to say why a truthful stamp can still hide an untrusted origin.

for a middle

Explain field by field what an auto-extracted entry records, and identify the one that is missing - a pointer to the source span. Be ready to say why extraction as paraphrase makes that pointer hard to keep.

for a senior

Demonstrate what you can and cannot establish during triage from such a store, and resist the reflex to 'check the provenance'. Show that a fully consistent record can leave the authorship question completely open.

for a principal

Be prepared to say what an attribution record in a personalising product is really for, and what the organisation has implicitly accepted by shipping one that answers write questions only.

### What a provenance field is actually recording On an assistant that fills long-term memory automatically, each entry typically carries some subset of: an id, the claim text, the component that wrote it, a session id, an owning account, a timestamp, maybe a confidence score. Every one of those is a fact about the *write*, not about the *claim*. That distinction is the whole question. The write genuinely happened inside a session belonging to the legitimate user, performed by a legitimate component, at a real time. Nothing in the record is forged, spoofed or anomalous. And yet the claim itself may have been authored by an unknown sender whose email the assistant read for ten seconds while triaging and never acted on. ### Why this is worse than a forged record A forged stamp is detectable in principle - signatures fail, ids do not resolve, timestamps disagree. A *truthful* stamp that answers the wrong question is not detectable at all, because there is no anomaly to find. The construction does not defeat provenance; it satisfies provenance. Whoever inherits the store is looking at thousands of records that all pass every consistency check they can run. ### The field that is missing, and why it usually is The one field that would decide the question is a pointer back to the exact source span - the message id, the offset, the sender. Auto-extraction usually does not keep it, for the ordinary reason that an extracted claim is a *paraphrase* of material spread across a transcript rather than a quotation of one span, and the transcript itself is discarded at teardown. So even where a message id survives, the text it points at may be gone, and the claim as stored may not appear verbatim anywhere. The consequence is directional and worth stating precisely: a memory entry's provenance proves which session wrote it, not that a person meant it. ### What that does to triage Someone investigating an odd assistant behaviour a month later can usually establish: | Question | Answerable from the store? | | --- | --- | | Which entry produced this behaviour | Often, if recall is logged | | When it was written, and by which component | Yes | | Which session it was written in | Yes | | Which message in that session it came from | Usually no | | Whether the account owner authored the claim | No | The first three look like a complete audit trail and are useless for the actual question. That gap is what turns a single planted claim into an open-ended adjudication problem rather than a cleanup task. ### Naming the obstacle without prescribing the fix The obstacle here is not a filter and not a screen. It is the *record itself*, which everyone reads as attribution and which attributes nothing about authorship. Recognising that is what an interviewer is scoring - candidates who say 'check the provenance' have assumed the record answers a question it was never structured to answer. ### Distinguishing this from adjacent ideas - It is not about whether the store should have a provenance field at all - that is a design conversation about write policy. - It is not about scoping reads to the right user; the entry is in the correct user's scope, legitimately. - It is not model behaviour. The model is downstream and simply reads what recall hands it. The subject is narrower and sharper: an attribution record that is completely accurate and completely unhelpful, because the write it accurately attributes is not the event anybody cares about.

  • If the entry did carry a source message id, would that settle it?
    It would narrow it, not settle it. A message id tells you which item in the mailbox the claim came from, which is enough to separate 'the user said this' from 'an unknown sender said this' - provided the message still exists and the claim is genuinely traceable to it. Extraction is paraphrase over a whole transcript, so a single pointer is often approximate, and it is absent far more often than not.
  • Does a recall log help, given the write record does not?
    A recall log tells you which entry was in context when the odd behaviour happened, which is genuinely useful for reproducing the effect. It says nothing about authorship - it moves you from 'something in memory did this' to 'this specific claim did this', and then you are back at the same unanswerable question about where the claim came from.

saying these in an interview costs you the question

  • Says provenance identifies the author of a claim
  • Assumes a planted entry looks anomalous in the record
  • Treats session ownership as evidence of user intent
  • Expects a forged or spoofed field to be findable
  • Confuses which write happened with which claim was meant

context