When should an agent write to long-term memory: per turn, at session end, or on an explicit remember tool?
answer
- three triggers, not one
- who pays the latency
- what happens if the session never ends
- full transcript beats single turn
- explicit writes: precise, low recall
basics
~20 sWrite triggers trade freshness for cost. Per-turn extraction catches everything but adds a model call to every turn. End-of-session extraction is cheaper and better informed. An explicit remember tool is precise but captures only what someone thought to flag.
solid answer
~50 sThere are three common triggers and most production systems combine them. **Per-turn** extraction runs after each exchange: nothing is lost if the session is abandoned, but you pay an extra model call and some latency on every turn, and the extractor judges each fact without knowing how the conversation ends. **End-of-session** extraction runs once over the whole transcript: it is far cheaper, the extractor sees the full arc so it can drop things that were later corrected, but a dropped connection loses the session. **Explicit** writes — a `remember`-style tool the user or the agent calls — are the highest-precision path and the only one a user can steer, but they only capture what somebody consciously flagged. A B2B sales assistant is the canonical shape: extract at end of call over the full transcript, plus an explicit remember tool the rep can fire mid-call for something they know matters.
go deeper
Know that an agent does not automatically remember anything — something has to trigger a write. Be able to name per-turn, end-of-session, and explicit remember-tool triggers and say one advantage of each.
Explain the mechanics: extraction is itself a model call, so the trigger sets your cost and latency profile, and an end-of-session pass sees the full conversation while a per-turn pass does not. Mention the idle timeout as the safety net.
Show you have operated this. Talk about async write queues, idempotent retries, abandoned sessions writing nothing silently, and the hybrid of explicit plus end-of-session that most production systems converge on.
Own the trade-off across the product: what write volume the business can afford at scale, whether users must be able to see and control what is remembered, and whether losing a session's memory is an acceptable failure or a compliance and trust problem.
## What a write trigger is An agent's long-term memory only changes when something fires a write. The **write trigger** is the rule that decides when the system stops and asks "is anything here worth keeping past this session?" It is a separate decision from *what* gets stored and *how* conflicts are resolved, and it is the first thing an interviewer probes because every downstream property — cost, latency, freshness, loss on crash — follows from it. A useful framing: writing is not free, because extraction is itself an LLM call. Deciding when to pay for that call is an engineering trade-off, not a detail. ## Per-turn extraction After each user (or assistant) turn, an extractor examines the exchange and proposes memory writes. - **Upside:** nothing is lost if the session ends abruptly. Facts become available to the *same* session's later steps, which matters for long-running agents that may exceed a single context window. - **Downside:** you add a model call and its latency to every turn. If the extractor runs inline, the user waits; most teams push it to an async job so the turn returns first, which reintroduces the risk of losing the last turns. - **Quality downside:** the extractor sees one turn without knowing how the conversation resolves. A user who says "let's move the kickoff to March" and then "actually, forget March" produces a write that must later be retracted. ## End-of-session extraction One extraction pass over the whole transcript when the session closes (or after an idle timeout). - **Upside:** one call instead of N. The extractor has the full arc, so it can ignore statements that were superseded within the session and can write a small number of well-formed facts instead of a stream of fragments. - **Downside:** a crashed process, a dropped connection or a session that never formally ends means nothing is written. This is why the trigger is usually "session closed **or** idle for N minutes **or** transcript exceeded N tokens", with the timeout as the real safety net. - **Latency profile:** the cost lands entirely off the user's critical path, which is why it is the default for chat products. ## Explicit writes A tool the user invokes ("remember that I always want dates in ISO format") or that the agent invokes when it judges something durable. Some systems make the agent's memory fully self-editing: the model issues its own add/replace calls against a memory store as part of its normal tool use, which is the design MemGPT and its successor Letta popularized. - **Upside:** highest precision, because a human or an explicitly reasoning model decided it matters. It is also the only trigger a user can *see* and control, which is important for trust and for consent. - **Downside:** recall is poor. People do not flag the facts that turn out to matter; they mention them in passing. - **Agent-initiated writes** sit between: they scale better than user-initiated ones but inherit the model's judgment errors, so they need the same dedup and contradiction handling as automatic extraction. ## The hybrid that most teams land on Take a sales assistant sitting on a 40-minute discovery call. The pipeline: an explicit `remember` tool the rep can fire mid-call for anything they know is load-bearing, plus one end-of-session extraction pass that reads the whole transcript and emits a handful of durable facts — budget authority, renewal timing, the competitor in the evaluation, the integration blocker. Four good facts from forty minutes is a *success*, not a shortfall. Per-turn extraction gets added on top only where the session is long-running and the agent itself must read back what it learned earlier — a multi-hour coding or research agent, for example. ## Cost and latency, concretely Estimate before choosing. Per-turn extraction on a 30-turn conversation is 30 extraction calls whose combined input tokens are roughly quadratic if each pass re-reads history, and linear if each pass sees only the new turn. End-of-session is one call over the full transcript. For high-volume consumer traffic that difference decides the architecture; for a handful of internal agents it is noise and you should optimize for not losing data instead. ## Failure modes to name - **Silent loss:** end-of-session only, with no idle timeout — abandoned sessions write nothing and nobody notices, because there is no error. - **Write amplification:** per-turn extraction on a chatty conversation producing dozens of near-identical records, which pushes the whole problem onto deduplication. - **Extraction on unvalidated turns:** writing a fact the same conversation later contradicts. - **Blocking the response:** running extraction inline and adding a second of latency to every turn for a store nobody reads back. ## What good answers do Name the three triggers, give the cost/latency/loss profile of each, then say plainly that real systems combine them and that the *timeout* is what makes end-of-session extraction safe.
- Your end-of-session extractor only fires when the client sends a close event. What goes wrong, and what do you add?Abandoned sessions, crashed clients and killed processes never emit close, so those conversations write nothing — and the failure is silent because no error is raised. Add an idle timeout that fires extraction after N minutes of inactivity, and a size trigger for sessions that run long. Treat the close event as an optimization on top of the timeout, not the primary trigger.
- Would you ever run extraction inline, on the user's critical path?Rarely, and only when the same session must read back what it just learned — a long-running agent that will exceed its context window, for instance. Otherwise queue the write asynchronously so the turn returns immediately. If you do queue it, make writes idempotent, because retries after a failure will replay the same extraction.
- How does the choice of trigger change what the extractor prompt should look like?A per-turn extractor sees one exchange, so it needs explicit instructions to skip anything provisional and to prefer no write over a speculative one. An end-of-session extractor sees the whole arc, so it can be told to resolve within-session corrections itself and emit only the final state of each fact. The same prompt used for both underperforms in both.
saying these in an interview costs you the question
- Assuming memory writes are free and running extraction every turn
- Treating a client close event as a reliable session-end signal
- Writing every raw message instead of extracting facts
- Believing a user-invoked remember tool alone gives adequate coverage
- Running extraction inline and blaming the model for turn latency