skip to content

What does an attacker give up by aiming at an assistant's end-of-turn memory extractor instead of the live answer?

level: seniorimportance: should knowfreq 36%

answer

  1. you trade control for persistence
  2. no reply comes back to tell you
  3. the extractor rewrites what it keeps
  4. fires later, or never
  5. one success is not a rate

basics

~20 s

Control and feedback. Whether the claim is kept, how it is reworded and when it is read back are decided by components the sender never sees, with nothing echoed back. What is bought is persistence with no expiry.

solid answer

~50 s

Going after the live turn is immediate and observable: the response either shows the effect or it does not, so a construction can be iterated. Going after the durable store gives up all three of those. Selection is out of the sender's hands - an extractor keeps what looks like a lasting fact about the user, so a long procedural instruction is either dropped or compressed into something that no longer has the effect. Extraction is paraphrase, so only claims that survive being restated in another component's words persist at all. And firing is deferred: the entry does something only when a later, unrelated request happens to pull it into context, which may be days later or never. In exchange the effect has no session boundary, no expiry, and no visible author - and it survives changes to the turn that seeded it, because the seed is long gone.

go deeper

for a junior

Know that a claim only persists if some component decides to keep it, and that nothing tells the sender whether it did. Be able to contrast an effect that lands now with one that lands later.

for a middle

Explain the pipeline steps that stand between untrusted text and a stored fact - selection at write time, paraphrase, then selection again at recall - and why each of them can silently drop the construction.

for a senior

Price the trade out loud: lower per-attempt reliability, no observability and constrained content, against persistence with no expiry and no author. Say where the construction stops working rather than claiming a general capability.

for a principal

Be ready to say what this route means for how a red-team programme reports results, given that success rates here are unobservable from outside and one reproduction against one deployment is not a rate.

### Two routes, and the trade between them Untrusted text that reaches an assistant's context can aim at the current answer or at what the assistant will remember. These are different constructions with different economics, and a candidate who can price both is well ahead of one who only knows that memory poisoning exists. **The live route.** The span influences the turn it lands in. It is observable - the reply either shows the effect or it does not - and therefore iterable. The cost is that it dies with the context, and its effect is bounded by what that one turn could do. **The durable route.** The span aims to become a stored fact. Everything about it is deferred and unobservable. ### What is given up, concretely 1. **Selection.** An automatic extractor keeps claims shaped like standing facts about the user - preferences, routines, relationships, settled decisions. Anything shaped like an instruction for right now is exactly what it discards. So the space of things that *can* persist is narrow and not chosen by the sender. 2. **Wording.** Extraction is paraphrase, not quotation. Whatever is stored has been restated by another component, in its own words, usually shortened. Any construction whose force depends on precise phrasing does not survive the round trip - it either loses the property that made it work or is not kept at all. 3. **Timing.** The entry does nothing until a later request happens to bring it into context. That could be the next day, next month, or never, and it depends on requests nobody controls. 4. **Feedback.** Nothing is returned. The live route confirms itself in the reply; the durable route confirms itself only when the effect eventually shows up somewhere the sender can see, which in most deployments is nowhere. Iterating on a construction with no signal is close to iterating blind. 5. **Reliability per attempt.** Because of 1 through 4, any single attempt has a low and unmeasurable success rate. That matters for how a finding gets reported, because reproducing it once against one deployment does not establish that it works. ### What is bought - **No session boundary.** The effect is not scoped to the conversation it arrived in, and applies in contexts unrelated to the original message. - **No expiry by default.** A store built to remember a person does not naturally forget. - **No visible author.** Every field on the entry attributes the write truthfully to a legitimate component in a legitimate session. - **Survivability.** The seeding turn, the message, and the transcript can all be gone. Whatever is done about the path the message came in on does not reach the claim already sitting in the store. ### Where it stops working The honest answer names limits rather than claiming a general capability: - The claim is never selected at write time, because it does not read like a durable fact about the user. - It is selected, but the paraphrase drops the property that gave it force. - It is stored, but recall never surfaces it in the contexts where it would matter. - A later, contradicting claim supersedes it. - The store is rebuilt or migrated, and unreproducible entries do not come across. ### How to talk about it in an interview The strong version of this answer is comparative and quantified in direction, not magnitude: lower per-attempt success, no observability, constrained content - in exchange for an effect with no end date and no attributable author. The weak version treats the durable route as strictly better because it lasts longer. It is not strictly better; it is a different bet, and which one is worth making depends entirely on whether persistence or control is the scarce resource.

  • How would you report the reliability of something you cannot observe from outside?
    By measuring it where you can - in a deployment you control, across many trials, reporting the fraction of seeding attempts that produce a stored claim and the fraction of later sessions that surface it. Two separate rates, not one. A finding that says 'this worked once' is describing an existence proof against one deployment, which is not the same as a rate and should not be written as one.
  • Why does a long, precisely-worded construction do badly on this route?
    Because it has to pass through a paraphrase. The extractor keeps a short claim in its own words, so anything whose effect depends on exact phrasing, structure or length is either not kept or is restated into something inert. What travels well is a short, ordinary-sounding, durable-looking statement - which also constrains how much can be smuggled at once.
  • Does the effect being deferred make it easier or harder to notice?
    Harder, in both directions. The person who seeded it gets no confirmation, and whoever eventually sees the odd behaviour is looking at a session with no connection to the original message - different day, different topic, nothing in the current turn to explain it. The delay removes the correlation that would normally point at a cause.

saying these in an interview costs you the question

  • Treats the durable route as strictly better than the live one
  • Assumes exact wording survives an extraction step
  • Expects observable feedback that the write happened
  • Reports one successful reproduction as a success rate
  • Ignores that recall must also select the entry later

context