skip to content

Chat Memory & Conversation State

Chat memory keeps a conversation coherent across turns, with a windowed history, advisors that inject it, and repositories that persist it per conversation. Interviewers ask how you bound the window, because unbounded history means unbounded cost.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

explore

questions

5

What is Spring AI's ChatMemory abstraction and why is it needed when talking to an LLM?

level: juniorimportance: must knowfreq 70%

answer

  1. LLM calls are stateless
  2. ChatMemory = store + replay past messages
  3. keyed by conversationId
  4. used via an advisor, not directly
  5. default MessageWindowChatMemory keeps last N

basics

~20 s

LLM calls are stateless — the model forgets previous turns. ChatMemory is a Spring AI interface that stores a conversation's past messages and replays them into each new request, so the model appears to remember the dialogue.

solid answer

~40 s

Each call to a chat model is independent — the model has no server-side memory of what was said before, so a follow-up like "and cheaper?" is meaningless without context. Spring AI's ChatMemory interface fixes this: it stores the ordered Messages of a conversation (keyed by a conversationId) and, on the next request, the memory content is prepended to the prompt so the model sees the history. You rarely call ChatMemory directly; instead you attach an advisor (MessageChatMemoryAdvisor or PromptChatMemoryAdvisor) to the ChatClient, and it reads history before the call and writes the new user/assistant messages back after. The default implementation is MessageWindowChatMemory, which keeps only the last N messages so the prompt (and token cost) stays bounded.

code

java · 13 lines
java
// MessageWindowChatMemory is the default ChatMemory implementation.
ChatMemory chatMemory = MessageWindowChatMemory.builder()
        .maxMessages(20)          // keep only the last 20 messages
        .build();                 // uses InMemoryChatMemoryRepository by default

ChatClient chatClient = ChatClient.builder(chatModel)
        .defaultAdvisors(MessageChatMemoryAdvisor.builder(chatMemory).build())
        .build();

// Turn 1
chatClient.prompt().user("What's the capital of France?").call().content();
// Turn 2 — the advisor replays turn 1, so "there" resolves to Paris
chatClient.prompt().user("How many people live there?").call().content();

go deeper

for a junior

Must nail the core idea: LLMs are stateless; ChatMemory stores and replays the past turns so the model 'remembers'.

for a middle

Should explain the advisor mechanism and that history is re-sent as tokens each call.

for a senior

Should mention boundedness (windowing), conversationId partitioning, and the memory-vs-RAG distinction.

for a principal

Frames memory as a token-budget/cost and correctness tradeoff and knows where persistence and pruning belong.

## The problem: LLMs are stateless A large language model (LLM) does not remember anything between HTTP calls. Every request is self-contained: the model only "knows" what is in the prompt you send it right now. If a user asks "What's the capital of France?" and then "How many people live there?", the second request contains no trace of the first unless *you* put it there. Without help, the model cannot resolve "there". **Conversation state** is the accumulated back-and-forth (user messages, assistant replies, sometimes tool results and a system message) that you must re-send on every turn to create the illusion of memory. ## What ChatMemory is `ChatMemory` is a Spring AI interface (package `org.springframework.ai.chat.memory`) that abstracts *where and how* that conversation state is stored. Its shape in Spring AI 1.0 is roughly: ```java public interface ChatMemory { String DEFAULT_CONVERSATION_ID = "default"; void add(String conversationId, List<Message> messages); List<Message> get(String conversationId); void clear(String conversationId); } ``` - `add` appends messages to a conversation. - `get` returns the stored messages for a conversation (already trimmed to the memory's policy). - `clear` wipes a conversation. - `conversationId` is the partition key that keeps different users/sessions apart (default value `"default"`). A `Message` is Spring AI's model of one turn — it has a type/role (SYSTEM, USER, ASSISTANT, TOOL) and content. ## How you actually use it You almost never call `add`/`get` yourself. Instead you register an **advisor** on the `ChatClient`. An advisor is a Spring AI interceptor that wraps each model call. The two memory advisors are `MessageChatMemoryAdvisor` and `PromptChatMemoryAdvisor`. On each request the advisor: 1. reads the stored history via `ChatMemory.get(conversationId)`, 2. injects it into the outgoing prompt, 3. lets the model respond, 4. writes the new user message and the assistant reply back via `ChatMemory.add(...)`. ```java var chatClient = ChatClient.builder(chatModel) .defaultAdvisors(MessageChatMemoryAdvisor.builder(chatMemory).build()) .build(); ``` ## Default implementation and boundedness The out-of-the-box implementation is `MessageWindowChatMemory`, which keeps only the most recent *N* messages (default 20). This matters because prompts have a finite context window and you pay per token — unbounded history would eventually overflow the context and cost more every turn. ## ChatMemory vs. RAG Don't confuse conversational memory with retrieval-augmented generation. ChatMemory replays the *literal recent turns* of one conversation. RAG (e.g. a VectorStore) semantically retrieves relevant chunks from a large corpus. They solve different problems and are often combined. ## When to use it Any multi-turn assistant/chatbot needs ChatMemory. A single-shot classification or summarization endpoint (one request, no follow-up) does not.

  • If ChatMemory replays history on every call, what runaway cost does that create and how does Spring AI bound it?
    Each turn re-sends all remembered messages, so token usage (and latency/$) grows with conversation length. MessageWindowChatMemory bounds it by keeping only the last N messages (default 20); older ones are evicted.
  • Is ChatMemory the same as a VectorStore for RAG?
    No. ChatMemory replays the literal recent turns of one conversation to give short-term dialogue continuity. A VectorStore semantically retrieves relevant documents from a large corpus (RAG). Different mechanisms, often combined.

saying these in an interview costs you the question

  • Thinking the LLM server itself remembers past turns without you resending them
  • Believing you must call ChatMemory.add/get manually instead of using an advisor
  • Confusing conversational memory with RAG / vector retrieval
  • Assuming memory is unbounded rather than windowed

context

open as a page

How does conversationId work in Spring AI ChatMemory, and how do you keep two users' conversations separate?

level: middleimportance: must knowfreq 65%

basics

~20 s

conversationId is the key that partitions stored messages. Each user/session gets its own id; you pass it per request via the advisor param ChatMemory.CONVERSATION_ID. If you omit it, everything shares the default id "default" and users see each other's history.

open as a page

What is the difference between MessageChatMemoryAdvisor and PromptChatMemoryAdvisor?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Both replay conversation history, but differently. MessageChatMemoryAdvisor injects the history as a list of separate role-tagged messages (USER/ASSISTANT). PromptChatMemoryAdvisor flattens the history into text inside the system prompt. Message-based preserves structure; prompt-based is a single blended system message.

open as a page

Explain MessageWindowChatMemory's windowing behavior and how ChatMemory differs from ChatMemoryRepository.

level: seniorimportance: should knowfreq 45%

basics

~20 s

MessageWindowChatMemory keeps only the most recent N messages (default 20), evicting the oldest so the prompt stays bounded — but it retains system messages. ChatMemory is the policy layer (windowing); ChatMemoryRepository is the pure storage layer it delegates persistence to.

open as a page

How do you make Spring AI chat memory durable in production using JDBC or Cassandra ChatMemoryRepository, and what operational concerns arise?

level: principalimportance: should knowfreq 40%

basics

~20 s

Replace the default in-memory store with a persistent ChatMemoryRepository — JdbcChatMemoryRepository (rows in a relational table) or CassandraChatMemoryRepository (wide-column, supports TTL). Add the matching Spring Boot starter; it auto-configures the bean, and you plug it into MessageWindowChatMemory.

open as a page