What is Spring AI's ChatMemory abstraction and why is it needed when talking to an LLM?
answer
- LLM calls are stateless
- ChatMemory = store + replay past messages
- keyed by conversationId
- used via an advisor, not directly
- default MessageWindowChatMemory keeps last N
basics
~20 sLLM calls are stateless — the model forgets previous turns. ChatMemory is a Spring AI interface that stores a conversation's past messages and replays them into each new request, so the model appears to remember the dialogue.
solid answer
~40 sEach call to a chat model is independent — the model has no server-side memory of what was said before, so a follow-up like "and cheaper?" is meaningless without context. Spring AI's ChatMemory interface fixes this: it stores the ordered Messages of a conversation (keyed by a conversationId) and, on the next request, the memory content is prepended to the prompt so the model sees the history. You rarely call ChatMemory directly; instead you attach an advisor (MessageChatMemoryAdvisor or PromptChatMemoryAdvisor) to the ChatClient, and it reads history before the call and writes the new user/assistant messages back after. The default implementation is MessageWindowChatMemory, which keeps only the last N messages so the prompt (and token cost) stays bounded.
code
java · 13 lines// MessageWindowChatMemory is the default ChatMemory implementation.
ChatMemory chatMemory = MessageWindowChatMemory.builder()
.maxMessages(20) // keep only the last 20 messages
.build(); // uses InMemoryChatMemoryRepository by default
ChatClient chatClient = ChatClient.builder(chatModel)
.defaultAdvisors(MessageChatMemoryAdvisor.builder(chatMemory).build())
.build();
// Turn 1
chatClient.prompt().user("What's the capital of France?").call().content();
// Turn 2 — the advisor replays turn 1, so "there" resolves to Paris
chatClient.prompt().user("How many people live there?").call().content();go deeper
Must nail the core idea: LLMs are stateless; ChatMemory stores and replays the past turns so the model 'remembers'.
Should explain the advisor mechanism and that history is re-sent as tokens each call.
Should mention boundedness (windowing), conversationId partitioning, and the memory-vs-RAG distinction.
Frames memory as a token-budget/cost and correctness tradeoff and knows where persistence and pruning belong.
## The problem: LLMs are stateless A large language model (LLM) does not remember anything between HTTP calls. Every request is self-contained: the model only "knows" what is in the prompt you send it right now. If a user asks "What's the capital of France?" and then "How many people live there?", the second request contains no trace of the first unless *you* put it there. Without help, the model cannot resolve "there". **Conversation state** is the accumulated back-and-forth (user messages, assistant replies, sometimes tool results and a system message) that you must re-send on every turn to create the illusion of memory. ## What ChatMemory is `ChatMemory` is a Spring AI interface (package `org.springframework.ai.chat.memory`) that abstracts *where and how* that conversation state is stored. Its shape in Spring AI 1.0 is roughly: ```java public interface ChatMemory { String DEFAULT_CONVERSATION_ID = "default"; void add(String conversationId, List<Message> messages); List<Message> get(String conversationId); void clear(String conversationId); } ``` - `add` appends messages to a conversation. - `get` returns the stored messages for a conversation (already trimmed to the memory's policy). - `clear` wipes a conversation. - `conversationId` is the partition key that keeps different users/sessions apart (default value `"default"`). A `Message` is Spring AI's model of one turn — it has a type/role (SYSTEM, USER, ASSISTANT, TOOL) and content. ## How you actually use it You almost never call `add`/`get` yourself. Instead you register an **advisor** on the `ChatClient`. An advisor is a Spring AI interceptor that wraps each model call. The two memory advisors are `MessageChatMemoryAdvisor` and `PromptChatMemoryAdvisor`. On each request the advisor: 1. reads the stored history via `ChatMemory.get(conversationId)`, 2. injects it into the outgoing prompt, 3. lets the model respond, 4. writes the new user message and the assistant reply back via `ChatMemory.add(...)`. ```java var chatClient = ChatClient.builder(chatModel) .defaultAdvisors(MessageChatMemoryAdvisor.builder(chatMemory).build()) .build(); ``` ## Default implementation and boundedness The out-of-the-box implementation is `MessageWindowChatMemory`, which keeps only the most recent *N* messages (default 20). This matters because prompts have a finite context window and you pay per token — unbounded history would eventually overflow the context and cost more every turn. ## ChatMemory vs. RAG Don't confuse conversational memory with retrieval-augmented generation. ChatMemory replays the *literal recent turns* of one conversation. RAG (e.g. a VectorStore) semantically retrieves relevant chunks from a large corpus. They solve different problems and are often combined. ## When to use it Any multi-turn assistant/chatbot needs ChatMemory. A single-shot classification or summarization endpoint (one request, no follow-up) does not.
- If ChatMemory replays history on every call, what runaway cost does that create and how does Spring AI bound it?Each turn re-sends all remembered messages, so token usage (and latency/$) grows with conversation length. MessageWindowChatMemory bounds it by keeping only the last N messages (default 20); older ones are evicted.
- Is ChatMemory the same as a VectorStore for RAG?No. ChatMemory replays the literal recent turns of one conversation to give short-term dialogue continuity. A VectorStore semantically retrieves relevant documents from a large corpus (RAG). Different mechanisms, often combined.
saying these in an interview costs you the question
- Thinking the LLM server itself remembers past turns without you resending them
- Believing you must call ChatMemory.add/get manually instead of using an advisor
- Confusing conversational memory with RAG / vector retrieval
- Assuming memory is unbounded rather than windowed