skip to content

Design a production retry-and-dead-letter topology that gives delayed (backoff) retries without blocking consumer threads, and caps attempts before parking to a DLQ.

level: principalimportance: should knowfreq 40%

answer

  1. wait queue: TTL + no consumer + DLX back to main
  2. TTL expiry = broker-side backoff, no thread blocked
  3. cap via x-death count or custom x-retry-count header
  4. terminal parking DLQ has NO dlx
  5. per-message TTL only expires at queue head -> use tiered queues

basics

~20 s

Send failures to a 'wait' queue that has a TTL and no consumer; its dead-letter exchange routes expired messages back to the main queue after the delay. Count redeliveries via the x-death header and, once a cap is hit, route to a terminal parking DLQ instead of retrying.

solid answer

~50 s

The idea is to move backoff off the consumer thread and onto the broker. Main queue's failures reject-without-requeue and dead-letter to a retry (wait) queue that has an x-message-ttl and no consumer; that queue's own x-dead-letter-exchange points back at the main exchange, so when the TTL expires the broker republishes the message to the main queue — a delayed redelivery, with zero thread blocking. To cap attempts, read the x-death header's count (or a custom retry-count header) in the listener; once it exceeds the max, route the message to a terminal parking DLQ instead of the retry queue. For variable backoff you can use several wait queues with increasing TTLs (or RabbitMQ's delayed-message exchange plugin). This scales far better than in-thread stateless/stateful retry for long delays, and the parking DLQ preserves poison messages with full x-death history for triage and replay.

code

java · 60 lines
java
@Configuration
class DelayedRetryTopology {

    // 1. Main queue dead-letters to the retry exchange on reject-no-requeue
    @Bean Queue orders() {
        return QueueBuilder.durable("orders")
            .withArgument("x-dead-letter-exchange", "orders.retry.dlx")
            .build();
    }

    // 2. Wait queue: TTL, no consumer, dead-letters BACK to the main exchange
    @Bean Queue ordersRetry() {
        return QueueBuilder.durable("orders.retry")
            .withArgument("x-message-ttl", 30_000)               // 30s backoff
            .withArgument("x-dead-letter-exchange", "orders.exchange")
            .withArgument("x-dead-letter-routing-key", "orders") // -> back to main
            .build();
    }

    // 3. Terminal parking DLQ: NO dlx, holds poison messages for triage
    @Bean Queue ordersParked() {
        return QueueBuilder.durable("orders.parked").build();
    }

    @Bean DirectExchange retryDlx() { return new DirectExchange("orders.retry.dlx"); }
    @Bean Binding retryBinding() {
        return BindingBuilder.bind(ordersRetry()).to(retryDlx()).with("orders");
    }
}

@Component
class OrderListener {
    private static final int MAX_ATTEMPTS = 5;
    private final RabbitTemplate rabbit;
    OrderListener(RabbitTemplate rabbit) { this.rabbit = rabbit; }

    @RabbitListener(queues = "orders")
    void handle(Message message) {
        try {
            process(message); // must be idempotent
        } catch (RetryableException ex) {
            if (deathCount(message) >= MAX_ATTEMPTS) {
                rabbit.send("", "orders.parked", message); // park terminally
            } else {
                throw new AmqpRejectAndDontRequeueException(ex); // -> retry loop
            }
        }
    }

    @SuppressWarnings("unchecked")
    private long deathCount(Message m) {
        var deaths = (List<Map<String, ?>>) m.getMessageProperties()
            .getHeaders().get("x-death");
        if (deaths == null || deaths.isEmpty()) return 0;
        Object c = deaths.get(0).get("count");
        return c == null ? 0 : ((Number) c).longValue();
    }

    private void process(Message m) { /* ... */ }
}

go deeper

for a junior

Understand the concept: a TTL wait queue delays messages then sends them back.

for a middle

Wire the TTL+DLX loop and know it moves backoff off the consumer thread.

for a senior

Add attempt capping via x-death/custom header and a terminal parking DLQ; ensure idempotency.

for a principal

Own the full topology: tiered backoff vs delayed-exchange plugin, per-message TTL head-of-line pitfall, durability/confirms, DLQ alerting and replay tooling, and the decision matrix vs in-thread retry.

**Why not just use the retry interceptor?** In-thread retry (stateless or stateful) does its backoff by sleeping the **consumer thread**. For short delays that's fine; for realistic backoffs (seconds→minutes) it idles consumers, consumes prefetch slots, and throttles the whole queue. The production-grade pattern pushes the *waiting* onto the broker so consumer threads stay free. **Core mechanism — TTL + DLX 'wait queue' (a.k.a. delayed retry / dead-letter loop):** 1. **Main queue** `orders` is declared with `x-dead-letter-exchange = orders.retry.dlx`. 2. On a *retryable* failure, the listener rejects-without-requeue (throws `AmqpRejectAndDontRequeueException` or the recoverer does), so the broker dead-letters the message to `orders.retry.dlx`. 3. **Wait queue** `orders.retry` is bound to `orders.retry.dlx`, has **`x-message-ttl`** (say 30s), and — critically — **no consumer**. Messages just sit there ticking down. 4. `orders.retry` itself declares **`x-dead-letter-exchange = orders.exchange`** (the *main* exchange) with a routing key that lands back on `orders`. When the TTL expires, the broker dead-letters the message **back to the main queue** — a **delayed redelivery** with no thread blocked. That loop = automatic backoff. Each pass through the wait queue adds an `x-death` entry. **Capping attempts (avoiding an infinite delayed loop):** The above alone would retry forever. To bound it: - Read the **`x-death`** header in the listener. It's a list; the relevant entry's **`count`** tells you how many times the message was dead-lettered from a given queue. Once `count >= maxAttempts`, **don't** let it re-enter the retry loop — instead publish/route it to a **terminal parking DLQ** (`orders.parked`) that has *no* DLX, where it rests for human triage/replay. - Alternatively maintain your **own `x-retry-count` header**, incrementing on each pass (more explicit, immune to x-death edge cases like the header being reset when routing changes). **Variable / exponential backoff:** A single fixed-TTL wait queue gives constant delay. Options: - **Tiered wait queues:** `orders.retry.5s`, `orders.retry.30s`, `orders.retry.5m` — route to a longer one as the retry count grows, approximating exponential backoff. - **RabbitMQ Delayed Message Exchange plugin** (`x-delayed-message` exchange type): set a per-message `x-delay` header for arbitrary delays without multiple queues. Trade-off: it's a plugin (not core), delays are held in the exchange (mnesia), and very large numbers of delayed messages can pressure the broker. **Per-message TTL caveat (important gotcha):** RabbitMQ only expires a message when it reaches the **head** of the queue (queues are FIFO for expiry checks). If you set *per-message* TTLs of different values in one queue, a message with a short TTL stuck behind one with a long TTL won't expire until the head one does — head-of-line blocking of expiry. That's why **tiered queues each with a uniform `x-message-ttl`** are safer than mixed per-message TTLs in a single wait queue. **Idempotency is mandatory:** Any redelivery scheme means the listener *will* occasionally see a message more than once (at-least-once delivery). Handlers must be **idempotent** (dedupe by business key / processed-id table), or delayed retry will double-apply side effects. **Durability:** Make the main, retry, and parked queues **durable** and messages **persistent**; consider `RepublishMessageRecovererWithConfirms` or publisher confirms on any explicit republish so a lost publish to the parking DLQ is detected. **Observability & ops:** Alert on parking-DLQ depth (it means something is genuinely broken), expose retry counts as metrics, and build a **replay** path (shovel messages from `orders.parked` back to `orders` after a fix). Keep the `x-exception-*` headers (via RepublishMessageRecoverer) so parked messages self-document. **Putting it together — decision guide:** - Short, few retries, non-transactional → in-thread **stateless** interceptor. - Fresh transaction per attempt → **stateful** interceptor (needs messageId). - Long/backoff delays, high throughput, must not block consumers → **TTL+DLX wait-queue topology** with an attempt cap and a terminal parking DLQ (this design). **Common mistakes:** (1) Forgetting the cap → infinite delayed loop. (2) Mixed per-message TTLs in one wait queue → head-of-line expiry blocking. (3) Non-idempotent handlers → duplicate side effects. (4) Parking DLQ with a DLX still attached → parked messages leak back into the loop. (5) Relying on x-death count without realizing it can be reset/split across queues — prefer an explicit retry-count header if precision matters.

  • Why do you set the TTL on the wait queue rather than per-message when using a single retry queue?
    RabbitMQ only expires a message when it reaches the head of the queue. With differing per-message TTLs in one queue, a short-TTL message behind a long-TTL one is blocked from expiring until the head expires — head-of-line expiry blocking. A uniform queue-level x-message-ttl avoids this; for varied delays use separate tiered wait queues.
  • What must the terminal parking DLQ NOT have, and why?
    It must not have an x-dead-letter-exchange. If it did, any reject/TTL there would route parked poison messages back into the retry loop, defeating the whole point of a terminal resting place for manual triage and replay.
  • Why is handler idempotency non-negotiable in this design?
    The delayed-retry loop plus at-least-once delivery guarantees a message can be processed more than once. Without idempotency (dedupe by business key / processed-ids table), retries double-apply side effects like charging a card or sending an email.

saying these in an interview costs you the question

  • Blocking the consumer thread with long Thread.sleep backoff at scale
  • No attempt cap, creating an infinite delayed-retry loop
  • Attaching a DLX to the terminal parking queue
  • Assuming per-message TTLs expire independently regardless of queue position
  • Ignoring idempotency for redelivered messages

context