What is @RetryableTopic non-blocking retry, and how does it differ from DefaultErrorHandler's blocking seek retries?
answer
- failed record → retry topics (orders-retry-0/1/2) → dlt
- main consumer NOT blocked (throughput preserved)
- breaks per-partition ordering
- auto-creates topics + consumers, uses DLPR internally
- @DltHandler, include/exclude, dltStrategy, @Backoff
basics
~20 s@RetryableTopic makes retries non-blocking: a failed record is forwarded to separate retry topics (with delays), so the main consumer keeps processing new records instead of being stuck. After the configured attempts, it goes to a DLT. DefaultErrorHandler instead retries in place by seeking, which blocks the partition.
solid answer
~40 s`@RetryableTopic` (on a `@KafkaListener`, Spring Kafka 2.7+) implements **non-blocking retries**. On failure the record isn't re-polled in place; it's **published to a dedicated retry topic** (e.g. `orders-retry-0`, `orders-retry-1`, …), each consumed by an auto-created listener with a time delay. The main topic's consumer immediately moves on to the next record — no partition blocking, so throughput is preserved. After the configured `attempts` are exhausted the record lands in a `<topic>-dlt`. Spring auto-creates the retry/DLT topics and consumers, using a `DeadLetterPublishingRecoverer` internally to forward between them. The big trade-off: it **breaks per-partition ordering**, because a failed record is retried later, out of band, while newer records proceed. DefaultErrorHandler (blocking, seek-based) preserves ordering but stalls the partition during retries. You can even combine them: a short blocking retry then hand off to non-blocking topics.
code
java · 18 lines@RetryableTopic(
attempts = "4",
backoff = @Backoff(delay = 1_000, multiplier = 2.0, maxDelay = 10_000),
autoCreateTopics = "true",
dltStrategy = DltStrategy.FAIL_ON_ERROR,
exclude = { IllegalArgumentException.class }) // fatal → straight to DLT
@KafkaListener(topics = "orders", groupId = "orders-svc")
public void handle(Order order) {
inventory.reserve(order); // may throw a transient exception
}
// Handles records that exhausted all retry topics
@DltHandler
public void onDlt(Order order,
@Header(KafkaHeaders.ORIGINAL_TOPIC) String topic,
@Header(KafkaHeaders.EXCEPTION_MESSAGE) String msg) {
log.error("Order {} dead-lettered from {}: {}", order.id(), topic, msg);
}go deeper
Know it retries on separate topics so the main flow keeps moving, then dead-letters.
Explain topic auto-creation (retry-0/1, dlt), attempts/@Backoff, and @DltHandler.
Articulate the ordering trade-off vs blocking seek retries, delay realization, and include/exclude classification.
Decide blocking vs non-blocking (or hybrid) per stream from ordering/throughput/SLA constraints and own the retry-topic topology and its operational cost.
**The core idea.** Blocking retries (DefaultErrorHandler) keep re-delivering the *same* record on the *same* partition until it succeeds or is recovered — nothing behind it moves. That's great for ordering, terrible for throughput and for slow back-offs. **Non-blocking retries** decouple the retry from the main consumption path: the failed record is *moved* to another topic and retried there later, so the main consumer never stalls. **How @RetryableTopic works.** Annotate the listener method: ```java @RetryableTopic(attempts = "4", backoff = @Backoff(delay = 1000, multiplier = 2.0)) @KafkaListener(topics = "orders") void handle(Order o) { ... } ``` Behind the scenes Spring (via `RetryTopicConfigurer`) creates: - **Retry topics** — by default one per retry level: `orders-retry-0`, `orders-retry-1`, `orders-retry-2` (with `attempts=4` you get the main attempt + 3 retry topics). Each has its own auto-generated consumer that waits the configured delay before processing (the delay is realized by the consumer holding/pausing until the record's timestamp + delay is reached). - **A DLT** — `orders-dlt` — where the record goes after the last retry fails. A `DeadLetterPublishingRecoverer` is used internally to forward the record from one topic to the next (and finally to the DLT), stamping the usual failure headers plus an attempt counter. **Topic strategies.** - `fixedDelayTopicStrategy` — with a fixed (non-exponential) delay you can choose `SINGLE_TOPIC` (reuse one retry topic) instead of one-per-attempt. - `SameIntervalTopicReuseStrategy` — reuse a single topic for the trailing equal-interval retries. - `dltStrategy` — `FAIL_ON_ERROR` (default), `NO_DLT` (no dead-letter topic), or `ALWAYS_RETRY_ON_ERROR`. - `@DltHandler` — a method in the same class to handle records that reach the DLT. - `include` / `exclude` (+ `traversingCauses`) — which exceptions trigger non-blocking retry vs go straight to DLT (the non-blocking analogue of retryable/fatal classification). - `timeout` / `autoCreateTopics` / partition count / `concurrency` per retry topic are configurable. **Delay realization.** Because Kafka has no native per-message delay, the retry-topic consumer checks the record's original timestamp; if the delay hasn't elapsed it **pauses** the partition and re-seeks so the record is re-read once the time passes — this keeps the delay accurate without spinning. **Ordering caveat (the headline trade-off).** Records that fail are retried *later and elsewhere*, while subsequent records on the main topic are processed immediately. So **per-key/per-partition ordering is not preserved**. If your domain requires strict ordering (e.g. state-machine events for one aggregate), non-blocking retries are unsafe — use blocking retries, or partition so that ordering-sensitive keys tolerate reordering-on-failure, or dedupe/reconcile downstream. **Combining blocking + non-blocking.** You can configure a small *blocking* retry set (via `ListenerContainerFactoryConfigurer` / `setBlockingRetryableExceptions` + a back-off) that runs in place first, then falls through to the non-blocking topics — useful to absorb very short transient blips without a topic hop, while still not blocking for long outages. **Global vs annotation config.** Instead of the annotation you can declare a `RetryTopicConfiguration` bean (`RetryTopicConfigurationBuilder`) to apply the same policy across many listeners and customize topic naming, partitions, serializers, etc. **When to use non-blocking.** High-throughput streams where head-of-line blocking is unacceptable and strict ordering isn't required; long back-offs (minutes/hours) that would otherwise stall the partition or risk rebalance. **When to prefer blocking:** strict ordering, simple topology, short retries. **Gotchas.** (1) Ordering loss is the number-one surprise. (2) It creates real topics — operational sprawl and ACLs/retention to manage. (3) Retried records arrive on a *different* topic name, so your listener/metrics must account for that (`@KafkaListener` on the retry topics is auto-added). (4) Delays are approximate and consume partitions while waiting. (5) Exactly-once/transaction semantics get more complex across the topic hops.
- What ordering guarantee do you lose with @RetryableTopic, and why?Per-partition (and thus per-key) ordering. A failed record is moved to a retry topic and processed later/out-of-band, while newer records on the main topic are processed immediately — so a later event can be handled before an earlier failed one. If strict ordering matters, use blocking retries instead.
- How does a retry-topic consumer implement a delay when Kafka has no per-message delay?The consumer checks the record's timestamp; if the configured delay hasn't elapsed it pauses the partition and re-seeks, re-reading the record once the delay passes — so it waits without busy-spinning and without committing the record early.
- Can you use blocking and non-blocking retries together?Yes. You can configure a small in-place blocking retry set (short back-off) that runs first to absorb brief blips, then falls through to the non-blocking retry topics for longer outages — combining fast local retries with non-stalling behavior.
saying these in an interview costs you the question
- Saying @RetryableTopic preserves ordering (it does not — that's its main trade-off)
- Thinking it seeks in place like DefaultErrorHandler (it forwards to separate topics)
- Believing you must manually create the retry/DLT topics and their consumers (Spring auto-creates them)
- Assuming the delay blocks the main partition (the wait happens on the retry-topic consumer, not the main one)