skip to content

Memory and Plan

Persistence turns one obeyed turn into a standing position: a remembered fact nobody can trace to its author, and a corrupted objective that step checks confirm. Interviewers ask what teardown ends.

on this pageshow

explore

questions

8

Why doesn't ending the session clear an injected fact an email assistant wrote to long-term memory?

level: juniorimportance: must knowfreq 72%

answer

  1. two kinds of state, not one
  2. the transcript goes, the store stays
  3. summarising is a write
  4. read back as the user's own preference

basics

~10 s

Session teardown discards the conversation context, not the durable store. Anything the assistant extracted and saved during that session outlives it by design, so a planted claim is read back into later, unrelated conversations.

solid answer

~50 s

An assistant that keeps long-term memory holds two different kinds of state. The ephemeral context - system prompt, turn history, retrieved text, tool results - is rebuilt per request and thrown away at teardown. The durable store is whatever a component decided was worth remembering, and in products that fill it automatically an extractor reads the transcript at the end of a turn or session and lifts out short claims that look like standing facts about the user. Teardown is defined not to touch that store; that is the feature. So an inbound email from an unknown sender that the assistant merely read while triaging - never replied to, never acted on - can leave a claim behind, because being in context was enough for the extractor to see it. On a later, unrelated request the entry is recalled as the user's own remembered preference, with no trace of the stranger who wrote it.

go deeper

for a junior

Recall that an assistant has short-lived context and a separate long-term store, and that ending a conversation only clears the first. Be able to say why text the assistant merely read can end up in the second.

for a middle

Explain the mechanics: an extractor reads the transcript at end of turn, keeps claims shaped like standing facts, and the stored text loses the framing that marked its source as untrusted. Name that recall reinjects it into unrelated later contexts.

for a senior

Show you know the practical consequence - the exposure window is not a session, and cleanup means adjudicating stored claims rather than revoking anything. Be ready to say what evidence you would still have after teardown, which is usually very little.

for a principal

Own the framing that anything a product remembers on the user's behalf is state with no expiry and no author. Be ready to argue where that sits between a defect and an accepted property of a personalising product.

### The claim being tested 'The session ended, so the injection is gone' is the single most common wrong answer about assistants that keep memory. It is true of the transcript and false of everything the product deliberately kept. ### Two pieces of state, not one A tool-using assistant - say one over a user's mail, calendar and contacts - holds at least two kinds of state: - **Ephemeral context.** The system prompt, the turn history, fetched or retrieved text, tool return values. This is assembled per request and discarded when the session tears down. - **Durable store.** Whatever the product decided is worth remembering across sessions. In consumer assistants this is usually filled *automatically*: at end of turn or end of session, a component reads the transcript and extracts short claims that read like standing facts about the user - a preference, a routine, a relationship, a decision that supposedly already happened. Teardown ends the first. It is defined not to touch the second. ### Where the untrusted text gets in An inbound email from an unknown sender is untrusted text, and a triage assistant reads it. Nothing is acted on in that session: no reply, no calendar write, no tool call. The user skims a summary and moves on. That is already enough, because the message body sat in the context the extractor reads. If a span in it is shaped like a durable statement about the user rather than an instruction for right now, the extractor may lift it into the store as a fact. Notice what did *not* have to happen. The model did not have to be talked past its refusal training - nothing refused, because nothing harmful was asked in that turn. No tool fired. The user did not have to click anything. The construction is aimed at the write path, and the write path is fed by whatever was in context. ### Why the origin does not travel with the fact Once the claim is stored, its text is what survives. On a later request - different day, different topic, possibly a different device - a memory read pulls it back into context, and it arrives there formatted as the assistant's own recollection *about the user*, not as a quoted message from a stranger. The model cannot discount it as untrusted, because the framing that would have marked it untrusted was dropped at extraction time. Recall renders a fact; it does not re-render the conversation the fact came from. This is also why 'the assistant never obeyed anything' misses. The model obeying a retrieved span would prove the span reached the context. Here the interesting event is earlier and quieter: a span reached the *store*. ### What changes for the person who planted it The exposure window stops being a session. No expiry runs by default on a store built to remember people, and nothing in the entry links back to the message that seeded it. A single-turn injection dies when its context dies. This one is re-read on every later request whose recall happens to select it, in contexts that have nothing to do with the original mail. Two things still have to go right for it: the extractor has to keep the claim, and a later recall has to select it. Neither is under the sender's control, and neither returns any signal to them. ### The honest form of the answer Say what teardown covers and what it structurally cannot: teardown discards context, and anything summarised, extracted or remembered has *by construction already left context*. Then say what that implies for cleanup. You are not revoking a session. You are deciding which of thousands of remembered claims were authored by somebody other than the user, usually with no surviving evidence either way. ### Ways candidates get it wrong - Treating 'never acted on' as 'never processed'. - Assuming a memory entry inherits the trust level of the account it sits under. - Assuming that a store which can be written automatically can also be attributed automatically. - Confusing the context window with the long-term store, and answering about the wrong one.

  • Does it matter that the assistant never acted on the email during the triage session?
    No. Acting is a separate path. The write path is fed by what was in context, and the message body was in context the moment the assistant read it for triage. A construction aimed at durable state does not need a tool call, a reply or a user click in the session it arrives in - it needs to be shaped like a fact worth keeping.
  • From the planter's side, what would actually stop the claim being read back later?
    Only the entry ceasing to exist or ceasing to be selected on recall - the store being rebuilt, the claim superseded by a later contradicting one, or recall never surfacing it in the contexts that matter. Closing the conversation touches none of that. That asymmetry is the whole point: the effect has no natural end date attached to it.
  • How is this different from a directive that arrives in a retrieved chunk mid-conversation?
    A retrieved chunk influences the turn it lands in and dies with the context. This influences turns nobody has had yet. The difference is not the wording of the span but which piece of state it reaches - transient context versus a store the product refills into every future context on the user's behalf.

A visitor writes a note and sticks it on your fridge. The visit ends; the note does not. Next month you read it as your own reminder and never think about who wrote it.

saying these in an interview costs you the question

  • Assumes teardown deletes everything the session touched
  • Says the email was never acted on, so nothing happened
  • Treats a remembered fact as user-authored by definition
  • Confuses the context window with the long-term store
  • Thinks the model learned the text by updating weights

context

open as a page

Why does per-step verification against an agent's recorded plan confirm a goal-substitution attack?

level: middleimportance: must knowfreq 66%

basics

~20 s

Because the check compares each action against the recorded objective, and the recorded objective is what was edited. It tests consistency, not authenticity, so after the edit every action genuinely serves the goal on file and the verifier signs each one off.

open as a page

In an LLM agent, how does editing the standing objective differ from getting one tool call obeyed?

level: juniorimportance: should knowfreq 58%

basics

~20 s

One obeyed call ends when that step ends. An edited objective persists: the agent re-reads it at the top of every later step and plans from it, so it keeps pursuing the substituted goal on its own initiative with no further injected text.

open as a page

What does a stored memory entry's provenance prove when an extractor wrote it during the user's own session?

level: middleimportance: should knowfreq 54%

basics

~20 s

It proves which component wrote the entry and in which session, and nothing else. Both stamps are truthful for a claim lifted from a stranger's email, so provenance cannot separate attacker-seeded facts from ones the user actually stated.

open as a page

In a long-running LLM agent, what must an obeyed span reach to change the standing objective rather than one step?

level: seniorimportance: should knowfreq 40%

basics

~20 s

It must reach the record the loop re-reads, through a step that writes into it — usually one consolidating the standing goal. The agent performs that write, so the span must read as task detail, not an instruction.

open as a page

What does an attacker give up by aiming at an assistant's end-of-turn memory extractor instead of the live answer?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Control and feedback. Whether the claim is kept, how it is reworded and when it is read back are decided by components the sender never sees, with nothing echoed back. What is bought is persistence with no expiry.

open as a page

A goal-substitution finding reproduces once in five runs and the step trace is green — how do you adjudicate it?

level: principalimportance: nice to knowfreq 31%

basics

~20 s

Adjudicate the structural property, not the outcome. That the objective is run-writable and the step check reads it reproduces every time, while the end-to-end result is one in five. Report the rate with its trials, and file it against the architecture owner.

open as a page

An assistant's memory store holds thousands of extracted claims with no source text - what do you tell the owner about trusting it?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

No claim in an auto-extracted memory store can be attributed, so the choice is what to accept, not what to verify: keep it and carry claims of unknown authorship, discard it and lose personalisation, or re-derive what surviving data supports.

open as a page