In LangChain, how do buffer, summary and token-buffer conversation memories differ?
answer
- what do you drop when it overflows
- verbatim, window, token budget, summary
- summary costs an extra call per turn
- turn counts are a bad size proxy
- hybrid keeps recent verbatim
basics
~20 sBuffer memory replays every turn verbatim, so fidelity is perfect and tokens grow without limit. Summary memory keeps a running LLM-written summary, so the prompt stays flat but each turn costs an extra model call and loses detail. Token-buffer memory keeps the newest messages that fit a token limit.
solid answer
~50 sThey differ in what they throw away. `ConversationBufferMemory` keeps the full transcript verbatim: perfect recall, unbounded prompt growth. `ConversationBufferWindowMemory` keeps only the last k exchanges, which bounds size but forgets abruptly. `ConversationSummaryMemory` folds history into a running summary produced by an LLM: the prompt stays roughly flat, but you pay an extra model call per turn, add latency, and lose specifics like numbers and names. `ConversationSummaryBufferMemory` is the hybrid — recent turns verbatim, older ones compressed once they exceed a token limit. `ConversationTokenBufferMemory` keeps as many recent messages as fit under a token limit, measuring with the model's tokenizer rather than counting turns. These names are the pre-1.0 `langchain.memory` API, deprecated during 0.3 and not part of LangChain 1.x; the same tradeoff now lives in an explicit trimming or summarizing step you compose into the chain.
go deeper
Be able to say plainly that buffer keeps everything, window or token buffer keeps only the recent part, and summary replaces old turns with a shorter written recap.
Explain each option's cost per turn and what it discards, and note that these class names are pre-1.0 and that current code composes the policy explicitly.
Argue the hybrid case: recent verbatim, older compressed, with exact values kept as structured state because summaries drop numbers and negations.
Own the economics — an extra summarization call per turn multiplies inference spend and adds latency on the request path, so tie the policy to session-length distribution rather than picking one globally.
## The single question all of them answer Every conversation memory strategy answers one question: what do you drop when the transcript no longer fits the context window? The classic LangChain classes are just named points on that curve, and interviewers use the names as a vocabulary for the tradeoff. ## Verbatim buffer `ConversationBufferMemory` replays every message exactly. Nothing is lost, the model can quote turn one at turn forty, and there is no extra inference cost. What you pay is linear growth: prompt tokens, cost per turn and time-to-first-token all rise with conversation length, and eventually the request is rejected outright. It is the right default for short, bounded interactions — a five-turn form filler — and the wrong one for an always-on assistant. ## Fixed window `ConversationBufferWindowMemory` keeps the last k exchanges. Cost is now constant and predictable, which is its whole appeal. The failure mode is a cliff rather than a slope: the fact the user gave in turn one disappears the moment it slides out of the window, and nothing warns you. Counting turns is also a poor proxy for size, since one pasted stack trace can dwarf twenty short exchanges. ## Token buffer `ConversationTokenBufferMemory` fixes that proxy problem. It keeps the most recent messages that fit under `max_token_limit`, measuring with the model's own tokenizer, and drops from the oldest end. You get a genuine budget rather than a turn count, at the same information cost as a window: what falls off is simply gone. ## Running summary `ConversationSummaryMemory` sends the old history plus the new turn to an LLM and stores the rewritten summary. The prompt stays roughly flat no matter how long the conversation runs, which is the point. Three costs are easy to forget. First, an extra model call on every turn: more money, more latency on the user's critical path, another dependency that can fail. Second, lossy compression — summaries reliably drop exact figures, identifiers and negations, and a wrong number reads as confidently as a right one. Third, drift: summarizing a summary repeatedly amplifies whatever the first pass got wrong. ## Hybrid `ConversationSummaryBufferMemory` keeps recent turns verbatim and summarizes only what spills past a token limit. This is usually the shape you actually want, because recency is where verbatim fidelity matters and old context compresses well. It inherits both cost profiles: summarization fires only on overflow rather than every turn. ## Version reality These classes belong to the pre-1.0 `langchain.memory` package with the `BaseMemory` protocol (`load_memory_variables`, `save_context`, `memory_key`, `return_messages`), typically paired with `ConversationChain`. They were deprecated during the 0.3 line and are not part of LangChain 1.x. Saying that plainly is part of a good answer. In current code you keep a chat message history and compose the policy yourself: `trim_messages` for the token-buffer behaviour, a list slice for the window, and an explicit summarization step for the compressed variant. The classes are gone; the tradeoff is identical, and it is the tradeoff the interviewer is testing. ## How to choose Ask how far back the task genuinely needs to see, and what an error costs. Short transactional flows take a verbatim buffer. Long assistants take recent-verbatim plus a compressed tail. Anything where exact values matter should carry those values as structured state rather than trusting a summary to preserve them.
- Why is a fixed turn window a poor proxy for a token budget?Because turns vary hugely in size. Twenty one-line exchanges may cost less than a single pasted log file, so a window of ten turns can be tiny one moment and blow the context window the next. Budgeting in tokens, measured with the model's tokenizer, gives predictable prompt size regardless of what the user pastes.
- What kinds of information do running summaries lose first?Exact figures, identifiers, dates and negations. A summary that says the user discussed pricing will not preserve that they quoted 4,200 units or that they explicitly ruled out option B. Carry values that must be exact as structured state alongside the summary, rather than trusting prose compression to keep them.
saying these in an interview costs you the question
- Calls summary memory free, ignoring the extra LLM call
- Thinks buffer memory truncates itself automatically
- Treats a turn window as equivalent to a token budget
- Presents these classes as current LangChain 1.x API
- Assumes summaries preserve exact numbers and identifiers