Explain MessageWindowChatMemory's windowing behavior and how ChatMemory differs from ChatMemoryRepository.
answer
- window = last N messages, default 20, FIFO evict
- system messages retained across the window
- ChatMemory = policy; ChatMemoryRepository = storage
- MessageWindowChatMemory delegates persistence to a repository
- counts messages not tokens; in-memory default is volatile
basics
~20 sMessageWindowChatMemory keeps only the most recent N messages (default 20), evicting the oldest so the prompt stays bounded — but it retains system messages. ChatMemory is the policy layer (windowing); ChatMemoryRepository is the pure storage layer it delegates persistence to.
solid answer
~40 sMessageWindowChatMemory is the default ChatMemory. It enforces a sliding window of at most maxMessages (default 20): when adding new messages pushes the count over the limit, it evicts the oldest, but it preserves system messages so instructions aren't lost. This bounds token cost and keeps you inside the context window. Crucially there are two layers: ChatMemory is the *policy* — it decides what to keep and how to window — while ChatMemoryRepository is the *storage* — a simpler interface (findByConversationId / saveAll / deleteByConversationId) that just persists a conversation's messages. MessageWindowChatMemory holds a ChatMemoryRepository and delegates the actual read/write to it. Swap the repository (InMemory, JDBC, Cassandra) to change *where* messages live without changing the windowing policy. This separation is why you can have durable persistence and bounded windows independently.
code
java · 12 lines// Policy (windowing) + storage (repository) are independent knobs.
ChatMemoryRepository repository = new InMemoryChatMemoryRepository(); // swap for JDBC/Cassandra
ChatMemory chatMemory = MessageWindowChatMemory.builder()
.chatMemoryRepository(repository) // WHERE messages live
.maxMessages(20) // HOW MANY are kept (system msgs retained)
.build();
// The repository has no windowing; it just stores/returns messages:
// List<Message> findByConversationId(String id)
// void saveAll(String id, List<Message> messages)
// void deleteByConversationId(String id)go deeper
Should know memory keeps only the recent messages so it stays bounded.
Should state maxMessages default 20 and FIFO eviction.
Should articulate the ChatMemory (policy) vs ChatMemoryRepository (storage) split, system-message retention, and message-count-not-tokens.
Reasons about token budgets, durability/multi-instance, concurrency on a conversationId, and when windowing must be augmented by summarization/RAG.
## Two layers: policy vs. storage Spring AI splits conversational memory into two collaborating abstractions: 1. **`ChatMemory`** — the *policy* layer. It decides **what** to keep (windowing, trimming, system-message retention) and exposes `add/get/clear(conversationId)`. 2. **`ChatMemoryRepository`** — the *storage* layer. It only knows how to persist and fetch a conversation's messages, with a smaller CRUD-ish surface: ```java public interface ChatMemoryRepository { List<String> findConversationIds(); List<Message> findByConversationId(String conversationId); void saveAll(String conversationId, List<Message> messages); void deleteByConversationId(String conversationId); } ``` The repository has **no concept of windowing** — it stores and returns whatever it's given. The policy lives above it. ## MessageWindowChatMemory `MessageWindowChatMemory` is the default `ChatMemory` implementation (it replaced the older `InMemoryChatMemory`). It composes a `ChatMemoryRepository` and layers a **sliding window** on top: - **`maxMessages`** (default **20**): the maximum number of messages retained for a conversation. When `add(...)` would exceed it, the **oldest** messages are evicted first-in-first-out. - **System-message retention**: system messages are treated specially — they are **preserved** even as the window slides, so your standing instructions aren't silently dropped when the window fills with dialogue. (If a new system message arrives, older ones may be replaced, but system instructions aren't evicted merely because the window is full.) Build it explicitly to override defaults or plug in persistence: ```java ChatMemory chatMemory = MessageWindowChatMemory.builder() .chatMemoryRepository(jdbcChatMemoryRepository) // where messages live .maxMessages(30) // window policy .build(); ``` If you don't specify a repository, it defaults to `InMemoryChatMemoryRepository` (a `ConcurrentHashMap`-backed store — process-local and lost on restart). ## Why the split matters This separation is deliberate and powerful: - **Change storage without changing policy.** Move from in-memory to `JdbcChatMemoryRepository` or `CassandraChatMemoryRepository` by swapping the repository; the 20-message window still applies. - **Change policy without changing storage.** Adjust `maxMessages` or use a different `ChatMemory` implementation while keeping the same durable store. - **Testability.** Unit-test windowing with an in-memory repository; integration-test persistence separately. ## Gotchas - **Window is a *count*, not a token budget.** MessageWindowChatMemory counts messages, not tokens. Twenty very long messages can still blow the context window; if token cost is the concern you may need a smaller window or a token-aware strategy. - **In-memory default leaks/loses data.** The default repository is process-local: it grows unbounded across conversations (no TTL) and is wiped on restart. Production wants JDBC/Cassandra. - **Windowing loses old context.** Anything evicted is gone from the model's view — long-term recall needs RAG/summarization, not just a bigger window. - **Concurrency.** Two concurrent requests on the same conversationId can interleave add/get; design your id-per-request flow accordingly. ## When to tune Raise `maxMessages` when the assistant needs more recent context and you can afford the tokens; lower it to cut cost/latency. For durability and multi-instance deployments, always back it with a persistent `ChatMemoryRepository`.
- If maxMessages is 20, why can the prompt still overflow the model's context window?Because MessageWindowChatMemory counts messages, not tokens. Twenty long messages may exceed the context limit. For hard token budgets you need a smaller window or a token-aware/summarizing strategy.
- You swap InMemoryChatMemoryRepository for JdbcChatMemoryRepository. Does the windowing behavior change?No. Windowing lives in the ChatMemory (MessageWindowChatMemory), not the repository. Changing the repository only changes where messages are stored; the 20-message window still applies.
saying these in an interview costs you the question
- Thinking the repository, not MessageWindowChatMemory, does the windowing
- Believing maxMessages is a token limit rather than a message count
- Assuming system messages get evicted when the window fills
- Not knowing the default in-memory repository is volatile and unbounded across conversations