Compare stateful vs stateless retry interceptors in Spring AMQP. When must you use stateful retry?
answer
- stateless = in-memory loop, one delivery, blocks thread
- stateful = broker redelivers, count by messageId in RetryContextCache
- stateful needed for fresh transaction per attempt
- MessageKeyGenerator default = messageId
- no messageId -> stateful loops forever
basics
~20 sStateless retry loops in memory on the same delivery without going back to the broker between attempts. Stateful retry rejects and lets the broker redeliver, tracking attempts by a message key across redeliveries. Use stateful when each attempt needs a fresh transaction/redelivery.
solid answer
~50 sA stateless RetryInterceptor keeps the message in the consumer and retries the listener call in a loop in memory; the broker sees one delivery, and backoff blocks that consumer thread. It's simpler and the default, but it can't cross a transaction boundary — if the message must be redelivered by the broker for each attempt (e.g. an external transaction that rolls back and must restart), stateless won't work. A stateful interceptor instead rejects the message so the broker redelivers it; it identifies the same message across redeliveries via a MessageKeyGenerator (default: the messageId) stored in a RetryContextCache, counting attempts until exhaustion, then invoking the recoverer. Use stateful when you need a new transaction per attempt or true broker redelivery. The catch: stateful requires a reliable unique message id — without a messageId or a custom NewMessageIdentifier, it can't correlate redeliveries.
code
java · 28 lines// Stateful retry: broker redelivers; state keyed by messageId
@Bean
SimpleRabbitListenerContainerFactory rabbitListenerContainerFactory(
ConnectionFactory cf,
SimpleRabbitListenerContainerFactoryConfigurer configurer,
MessageRecoverer recoverer) {
var factory = new SimpleRabbitListenerContainerFactory();
configurer.configure(factory, cf);
factory.setAdviceChain(
RetryInterceptorBuilder.stateful() // <-- stateful
.maxAttempts(4)
.backOffOptions(2000, 2.0, 30_000)
.messageKeyGenerator(msg -> // correlate redeliveries
msg.getMessageProperties().getMessageId())
.recoverer(recoverer)
.build());
return factory;
}
// Producer must generate ids or stateful retry can't count attempts:
@Bean
RabbitTemplate rabbitTemplate(ConnectionFactory cf) {
var t = new RabbitTemplate(cf);
t.setMessageConverter(new Jackson2JsonMessageConverter());
// ensure every message has a unique messageId
// (via a MessagePostProcessor or setBeforePublishPostProcessors)
return t;
}go deeper
Know there are two retry modes; stateless loops in memory, stateful uses broker redelivery.
Know stateless blocks the thread and loses state on restart; stateful counts attempts by messageId.
Articulate the transaction-boundary reason for stateful and the messageId requirement/gotcha.
Weigh in-thread retry vs a TTL-based delayed-retry topology for throughput and durability at scale.
**The retry interceptor:** Spring AMQP wraps your listener with a retry advice (`RetryOperationsInterceptor`, built via `RetryInterceptorBuilder`) backed by Spring Retry's `RetryTemplate`. It comes in two flavors — **stateless** and **stateful** — that differ in *where the retry state lives* and *whether the broker is involved between attempts*. **Stateless retry (`RetryInterceptorBuilder.stateless()`):** - The message is delivered **once**. The interceptor catches the listener exception and **loops in memory**, re-invoking the listener up to `maxAttempts`, sleeping for the backoff between tries **on the consumer thread**. - The broker is *not* contacted between attempts — no redelivery, the delivery tag stays open. - On exhaustion, the `MessageRecoverer` is called (ack/republish/reject). - **Pros:** simple, no dependency on message ids, no broker round-trips. It's the common default. - **Cons:** (1) Backoff **blocks the consumer thread** — long backoffs reduce throughput and hold prefetch slots. (2) Retry state is **in JVM memory** — a restart loses the count. (3) It **cannot span a transaction boundary correctly**: if the listener runs inside a transaction that the container commits/rolls back per delivery, an in-memory retry re-runs *within the same failed transactional context*, which is wrong when the resource (DB/JMS) needs a fresh transaction per attempt. **Stateful retry (`RetryInterceptorBuilder.stateful()`):** - Each attempt is a **separate broker delivery**. On failure, the interceptor **rethrows** so the container rejects the message with requeue, and the **broker redelivers** it as a new delivery. - To know 'this redelivery is attempt N of the *same* logical message', it needs to **identify the message across redeliveries**. It uses a **`MessageKeyGenerator`** (default: the message's `messageId` property) and stores retry state in a **`RetryContextCache`** (in-memory map by default) keyed by that id. Each redelivery looks up the count, increments, and either retries or, on exhaustion, invokes the recoverer. - **When you MUST use stateful:** any time an attempt requires a **brand-new transaction** — e.g. a `@Transactional` listener writing to a database, where a rollback must fully unwind and the next attempt must start a clean transaction. Stateless can't do this because it retries inside the same broken unit of work. Also when you deliberately want the broker to own redelivery (survives some failure modes better, spreads attempts). - **The hard requirement:** every message needs a **unique, stable id**. If producers don't set `messageId`, correlation breaks and *every redelivery looks like attempt 1* → infinite retries. Fix by enabling `RabbitTemplate.setCreateMessageIds(true)` on the producer, or supply a custom **`NewMessageIdentifier`** / `MessageKeyGenerator` (e.g. derive the key from a business key/header). There's also a `NewMessageIdentifier` to tell the interceptor whether a delivery is genuinely new. - **Cons:** more moving parts, needs message ids, the `RetryContextCache` can grow (bounded `MapRetryContextCache` by default, capacity ~2^11); a poison message that exhausts still needs a recoverer or DLX. **Backoff & recovery are shared:** Both build on the same `BackOffPolicy` (fixed/exponential via `backOffOptions(initial, multiplier, max)`) and both end in a `MessageRecoverer` on exhaustion. **Choosing:** - Default to **stateless** for simple, non-transactional or transaction-tolerant handlers with short backoffs. - Use **stateful** when the listener is transactional and each retry needs a fresh transaction, or when you want broker-driven redelivery — and ensure message ids exist. - For *long* backoffs, prefer neither in-thread approach — use a **delayed-retry pattern**: dead-letter to a TTL queue whose DLX routes back to the main queue (a 'retry queue' / 'wait queue'), so you don't block consumer threads. This is often the production-grade choice. **Gotcha recap:** The single most common stateful-retry bug is missing `messageId` → attempts never accumulate → endless redelivery loop. Always verify id generation before enabling stateful retry.
- Your stateful retry never gives up and floods the DLQ never fires — messages loop forever. What's the most likely cause?Messages have no messageId (or the MessageKeyGenerator returns null/duplicate keys), so every redelivery is treated as a brand-new first attempt and the retry count never accumulates. Enable message-id generation or provide a key generator based on a stable business key.
- Why is a long exponential backoff a poor fit for stateless (in-thread) retry, and what's the alternative?Stateless backoff Thread.sleep()s on the consumer thread, so long waits idle the consumer and hold prefetch/QoS slots, throttling throughput. The alternative is a delayed-retry topology: dead-letter to a TTL 'wait' queue whose DLX routes the message back after the delay, freeing the consumer thread.
saying these in an interview costs you the question
- Saying stateless retry re-delivers via the broker between attempts
- Claiming stateful retry works without message ids
- Believing stateless retry can restart an external transaction per attempt