An assistant's memory store holds thousands of extracted claims with no source text - what do you tell the owner about trusting it?
answer
- no test exists, only a cost decision
- identical fields for both kinds of claim
- sampling finds the obvious, not the ordinary
- reviewed is not clean
- name who owns the residual
basics
~20 sNo claim in an auto-extracted memory store can be attributed, so the choice is what to accept, not what to verify: keep it and carry claims of unknown authorship, discard it and lose personalisation, or re-derive what surviving data supports.
solid answer
~50 sStart by making the limit explicit: every entry is truthfully stamped with a real component and a real session, the source text is gone, and a claim seeded by an inbound message is indistinguishable from one the user stated. So there is no verification available, only a decision about cost. The three honest positions are keep everything and accept that some unknown fraction has an author nobody can name; discard the store and pay in lost personalisation and user-visible regression; or re-derive the subset that still-existing data supports and drop the rest, which is the expensive middle. A human spot-check estimates how common obviously-wrong claims are and tells you nothing about ordinary-looking ones, which is exactly the shape a planted claim has. Whatever is chosen, the assurance language has to match: 'we reviewed a sample' is not 'the store is clean', and somebody has to own the residual in writing.
go deeper
Know that entries in an auto-filled memory store cannot be traced back to the text they came from, so 'check where it came from' is not an available step.
Be able to explain why sampling a corpus of extracted claims finds implausible entries but not ordinary-looking ones, and why that is the population that matters here.
Show you can lay out keep, discard and re-derive with the cost of each, and that you would state findings as an exposure bound rather than a count of affected entries.
Own the call and the language around it: which position is taken, who signs for the residual, what a review may be claimed to have bought, and whether the conversation is about one incident or about a design property that will keep producing them.
### The position you are actually in A store of automatically extracted claims about users, filled over months, with no retained source text. Some unknown number of entries originated in content the assistants merely read - inbound messages from senders nobody vetted - rather than in anything the account owners said. Every entry's provenance is truthful and useless: it names a legitimate extractor and a legitimate session, because the writes did happen there. This is a judgment call, not an investigation. The investigation has no available evidence, and pretending otherwise is the failure mode. ### Say the limit out loud first The most valuable thing to tell an owner is the shape of the ignorance: **there is no test that separates a seeded claim from a genuine one**, because the two are identical in every field that exists. Owners routinely assume that a store with an audit trail can be audited. Correcting that assumption early is what makes the rest of the conversation honest, and it reframes the ask from 'clean it' to 'decide what we are willing to carry'. ### The three positions, and who pays for each | Position | What it costs | What it leaves | | --- | --- | --- | | Keep the store as it is | Nothing up front | An unknown fraction of claims with unattributable authorship, read back indefinitely | | Discard it wholesale | Visible regression in personalisation; user-perceived amnesia; support load | No unattributable claims, and no memory | | Re-derive from surviving data | Engineering time, and it only covers claims whose evidence still exists | A smaller store you can say something about, plus a remainder you still have to decide on | None of these is a security control and none of them is free. The lead's job is to name which one is being taken, name who signed for it, and make sure the residual is written down rather than assumed away. ### What a spot-check can honestly be claimed to buy Sampling a corpus of extracted claims gives you a prevalence estimate for entries that look wrong on inspection - contradictory, oddly specific, out of character. It gives you nothing on the population that matters, because a claim seeded through an extractor has to look like an ordinary durable preference in order to have been kept at all. The selection pressure that got it into the store is the same pressure that makes it survive human reading. So the defensible claim is narrow: 'we sampled N entries and found X that were implausible on their face'. The indefensible one is 'we reviewed the store and it is clean'. If the assurance language cannot be narrowed to the first, the review should not be cited at all. ### Bug or design limit Expect the owner to ask whether this is a defect. The useful answer distinguishes the property from the incident: a product that remembers people without keeping sources was a design choice, made for good reasons, and its consequence - unattributable durable state - follows directly from it. A specific planted claim is an incident. Conflating the two either escalates a design property into a crisis or dismisses a real incident as inherent. Say which one is on the table, and say that the design property will keep producing incidents of this shape until somebody decides to change it, which is a roadmap conversation with an owner and a budget rather than a finding to be closed. ### What to expect to be pushed on - *How many entries are affected?* Unknown, and any number offered is fabricated. What can be bounded is exposure: how many accounts had assistants reading untrusted inbound content during the window. - *Can we just watch for it in future?* That is a different, forward-looking conversation with a different owner; it does not touch the entries already stored. - *Can we tell affected users?* Only at the population level, because per-entry attribution does not exist - and that materially shapes any disclosure decision. ### The mark of a strong answer It refuses to promise verification, prices the options rather than ranking them, is explicit about what a review buys, and ends with a named owner for the residual. A weak answer proposes a scan of the store for suspicious content and moves on - which is precisely the claim the shape of the problem does not support.
- The owner wants a number for how many entries are affected. What do you give them?Not a count of affected entries - that number does not exist and inventing it is worse than refusing. What can be bounded is exposure: how many accounts had assistants reading untrusted inbound content over the window, and how many entries were written in those sessions. That is an upper bound on the population at risk, stated as such, not an estimate of how many were actually seeded.
- Would you disclose to users, given per-entry attribution is impossible?Any disclosure has to be at the population level, because you cannot tell an individual whether their store is affected. That is a real constraint on the decision, not a reason to skip it: a notice that says 'remembered preferences may include content the assistant read rather than content you stated' is honest and actionable, where 'your account was affected' would be unsupportable in either direction.
- Is this a bug to fix or a property to accept?Both, at different altitudes. A specific planted claim is an incident with an owner. Unattributable durable state is a consequence of choosing to remember people without retaining sources, and it will keep producing incidents of this shape until that choice changes. Say which of the two is on the table, because treating a design property as a closeable ticket guarantees the ticket reopens.
saying these in an interview costs you the question
- Promises to scan the store and identify planted claims
- Cites a sample review as evidence the store is clean
- Gives an invented count of affected entries
- Treats wholesale deletion as free
- Leaves the residual risk with no named owner