skip to content

How do you make Spring AI chat memory durable in production using JDBC or Cassandra ChatMemoryRepository, and what operational concerns arise?

level: principalimportance: should knowfreq 40%

answer

  1. in-memory default = volatile + not shared + unbounded
  2. JDBC starter -> JdbcChatMemoryRepository, SPRING_AI_CHAT_MEMORY table
  3. initialize-schema: always/embedded/never (prod = never + Liquibase)
  4. Cassandra repo = native TTL auto-expiry
  5. window caps replay not storage -> retention/PII/latency concerns

basics

~20 s

Replace the default in-memory store with a persistent ChatMemoryRepository — JdbcChatMemoryRepository (rows in a relational table) or CassandraChatMemoryRepository (wide-column, supports TTL). Add the matching Spring Boot starter; it auto-configures the bean, and you plug it into MessageWindowChatMemory.

solid answer

~40 s

By default chat memory lives in an in-memory map — lost on restart and not shared across instances. For production you back MessageWindowChatMemory with a persistent ChatMemoryRepository. Add the JDBC starter (spring-ai-starter-model-chat-memory-repository-jdbc) and Spring Boot auto-configures a JdbcChatMemoryRepository over your DataSource, persisting messages in a SPRING_AI_CHAT_MEMORY table; control DDL with spring.ai.chat.memory.repository.jdbc.initialize-schema. Or use the Cassandra starter for a CassandraChatMemoryRepository, which additionally supports a time-to-live so old conversations auto-expire. Either way the repository is injected into MessageWindowChatMemory, so windowing still applies. Operational concerns: unbounded row growth without TTL/pruning (JDBC has no built-in expiry), horizontal scale and stickiness (persistence lets any instance serve any conversation), schema initialization strategy across environments, PII/privacy of stored transcripts (encryption, retention policy, right-to-erasure via clear()), and read/write latency on the hot path.

code

java · 16 lines
java
// application.properties
// spring.ai.chat.memory.repository.jdbc.initialize-schema=never   // prod: own DDL via Liquibase
// spring.ai.chat.memory.repository.cassandra.time-to-live=PT24H   // Cassandra auto-expiry

@Configuration
class ChatMemoryConfig {
    // JdbcChatMemoryRepository (or CassandraChatMemoryRepository) is auto-configured
    // by the matching spring-ai-starter-model-chat-memory-repository-* starter.
    @Bean
    ChatMemory chatMemory(ChatMemoryRepository repository) {
        return MessageWindowChatMemory.builder()
                .chatMemoryRepository(repository) // durable, shared across instances
                .maxMessages(20)                  // window still bounds the prompt
                .build();
    }
}

go deeper

for a junior

Should know the default memory is lost on restart and a database can persist it.

for a middle

Should name the JDBC/Cassandra repositories and that a starter auto-configures them into MessageWindowChatMemory.

for a senior

Should discuss schema initialization, Cassandra TTL, and why persistence enables multi-instance scaling.

for a principal

Reasons end-to-end: retention/growth, PII and erasure, hot-path latency, consistency, schema governance, and choosing JDBC vs Cassandra vs in-memory by workload.

## Why the default isn't production-grade Out of the box `MessageWindowChatMemory` uses `InMemoryChatMemoryRepository` — a `ConcurrentHashMap`. That means: - **Volatile**: every conversation is lost on restart/redeploy. - **Not shared**: instance A can't see a conversation started on instance B (breaks load-balanced/multi-instance deployments). - **Unbounded**: no cross-conversation cap or expiry; memory leaks as conversations accumulate. Production replaces the repository with a durable `ChatMemoryRepository`. ## JDBC persistence — `JdbcChatMemoryRepository` Add the Boot starter: ``` spring-ai-starter-model-chat-memory-repository-jdbc ``` Spring Boot auto-configuration builds a `JdbcChatMemoryRepository` on top of your existing `DataSource`/`JdbcTemplate`. Messages are stored in a relational table (default name **`SPRING_AI_CHAT_MEMORY`**) with columns for conversation id, message content, message type/role, and a timestamp for ordering. **Schema initialization** is controlled by: ```properties spring.ai.chat.memory.repository.jdbc.initialize-schema=always | embedded | never ``` - `always` — run the bundled DDL on startup (handy for dev). - `embedded` — only for embedded databases (H2, etc.). - `never` — you manage DDL yourself (typical for prod, where Liquibase/Flyway owns schema). Spring AI ships dialect-specific schema scripts. In a Liquibase/Flyway shop you'd set `never` and add the table to your migrations so schema is versioned and reviewed. ## Cassandra persistence — `CassandraChatMemoryRepository` Add: ``` spring-ai-starter-model-chat-memory-repository-cassandra ``` This auto-configures a `CassandraChatMemoryRepository`, a wide-column store well suited to append-heavy, high-write conversational workloads partitioned by conversation id. Its standout feature is native **TTL (time-to-live)**: ```properties spring.ai.chat.memory.repository.cassandra.time-to-live=PT24H ``` Rows automatically expire after the TTL, so abandoned conversations self-clean — no batch pruning job needed. This is a real operational advantage over JDBC. ## Wiring it in Whichever repository you pick, it plugs into the same policy layer: ```java ChatMemory chatMemory = MessageWindowChatMemory.builder() .chatMemoryRepository(jdbcChatMemoryRepository) // or Cassandra .maxMessages(20) .build(); ``` Windowing (policy) and persistence (storage) stay independent — you get durable, shared storage *and* a bounded prompt. ## Operational concerns (the principal-level meat) 1. **Unbounded growth / retention.** The window caps what's *replayed*, not what's *stored*. JDBC keeps every persisted message forever unless you prune. You need a retention policy: a scheduled delete, or use Cassandra TTL. Otherwise the table grows without bound. 2. **Horizontal scale.** Persistent, shared storage removes the need for sticky sessions — any instance can resume any conversation by id. This is often the primary reason to persist. 3. **Schema management across envs.** Decide `initialize-schema` per environment; in prod prefer versioned migrations (`never` + Liquibase/Flyway) over auto-DDL for auditability and rollback. 4. **Privacy / PII.** Transcripts often contain personal data. Consider encryption at rest, access controls, a retention limit, and honoring erasure requests via `ChatMemory.clear(conversationId)`. Don't log full transcripts. 5. **Latency on the hot path.** Every turn does a read (load history) and a write (persist new messages). A slow DB adds latency to each model call; size/connection-pool accordingly and keep the window modest. 6. **Concurrency & ordering.** Interleaved requests on one conversationId can race; the timestamp/order column and your id-per-request discipline matter. 7. **Consistency model.** Cassandra is eventually consistent by default — tune consistency level if a just-written turn must be readable immediately by another node. ## When to choose which - **JDBC**: you already run a relational DB, want transactional/queryable transcripts, and will manage retention yourself (or volumes are modest). - **Cassandra**: very high write throughput, many conversations, and you want automatic expiry via TTL and easy horizontal scale. - **In-memory**: only demos, tests, or single-instance ephemeral use.

  • The 20-message window is set — why can the JDBC table still grow without bound?
    The window limits what is replayed into the prompt, not what is persisted. JdbcChatMemoryRepository keeps every stored message indefinitely unless you prune. You need a retention job, or use Cassandra's TTL for automatic expiry.
  • What does moving from in-memory to a persistent repository buy you in a load-balanced deployment?
    Any instance can serve any conversation by id — no sticky sessions — and conversations survive restarts/redeploys, because the state now lives in shared durable storage instead of one process's heap.
  • In production, which initialize-schema setting is safest and why?
    never — let versioned migrations (Liquibase/Flyway) own the DDL so schema changes are reviewed, auditable, and reversible, rather than auto-applied at startup where they can surprise you or race across instances.

saying these in an interview costs you the question

  • Assuming maxMessages also limits how much is stored on disk
  • Relying on JDBC to auto-expire old conversations (it doesn't; only Cassandra has TTL)
  • Leaving initialize-schema=always in production instead of versioned migrations
  • Ignoring PII/retention/encryption of stored transcripts
  • Believing the in-memory default is fine for multi-instance deployments

context