skip to content

Contrast blocking retries (DefaultErrorHandler) with non-blocking retries (@RetryableTopic). When would you choose one over the other?

level: seniorimportance: must knowfreq 65%

answer

  1. Blocking = seek + same thread + ordering kept
  2. Non-blocking = retry topics + main flows + reordering
  3. Head-of-line blocking vs throughput
  4. max.poll.interval.ms risk on long blocking backoff
  5. Can combine: few blocking then hand off

basics

~20 s

Blocking retries re-process the failed record on the same consumer thread, holding up the partition until it succeeds. Non-blocking retries (@RetryableTopic) forward the record to separate retry topics with delays, so the main topic keeps flowing while failed records are retried independently.

solid answer

~50 s

Blocking retries (DefaultErrorHandler) seek the consumer back and re-poll the same record in-memory; the partition is stalled — nothing after the failed record is processed until it succeeds or is recovered. This preserves strict per-partition ordering but a single slow/poison record blocks everyone behind it, and long backoffs risk exceeding max.poll.interval.ms. Non-blocking retries (@RetryableTopic) immediately publish the failed record to a dedicated retry topic (e.g. *-retry-0, -retry-1 with growing delays) and commit the original, so the main partition keeps flowing; a separate consumer handles the delayed retry topic. This trades ordering for throughput and isolation. Choose blocking when strict ordering matters and failures are rare/transient and quick to clear; choose non-blocking when head-of-line blocking is unacceptable and you can tolerate reordering. They can be combined: a few fast blocking retries, then hand off to non-blocking topics.

go deeper

for a junior

Know that blocking holds the partition while retrying; non-blocking moves the record to other topics so the main flow continues.

for a middle

Explain the retry-topic mechanics of @RetryableTopic, immediate commit of the original, and the ordering tradeoff.

for a senior

Weigh ordering vs throughput, poll-interval risk, idempotency needs, and articulate a clear decision rule plus the combined pattern.

for a principal

Set org-wide guidance on retry strategy per workload class, model duplicate/ordering semantics, and design topic/infra topology and observability for retries.

## The core tension: ordering vs. progress Kafka guarantees **ordering within a partition**. A consumer reads a partition sequentially, tracking one offset. This creates the fundamental tradeoff in retry design. ### Blocking retries — `DefaultErrorHandler` When a record fails, `DefaultErrorHandler` **seeks the consumer back** to that record's offset and re-polls it on the **same thread**. Consequences: - **Ordering preserved**: nothing past the failed record is processed until it resolves, so per-partition order is intact. - **Head-of-line blocking**: a single record that keeps failing (a 'poison pill') or that needs a long backoff stalls **every** record behind it in that partition. - **`max.poll.interval.ms` risk**: the consumer must return to `poll()` within this window (default 5 min) or the broker considers it dead and triggers a **rebalance**. Long cumulative backoffs can blow this. Spring mitigates by pausing/resuming the consumer between attempts, but the ceiling still exists. - **No extra topics or infrastructure** needed. ### Non-blocking retries — `@RetryableTopic` `@RetryableTopic` (Spring Kafka 2.7+) flips the model. On failure, the record is **published to a separate retry topic** and the original offset is **committed immediately**, so the main partition keeps moving. Mechanics: - Spring auto-creates topics like `myTopic-retry-0`, `myTopic-retry-1`, … (or a single retry topic per attempt with `SingleTopic` strategy) and finally `myTopic-dlt`. - Each retry topic has an associated **delay**; a record arriving 'early' is held (the listener for the retry topic pauses/sleeps until the timestamp + backoff has elapsed) before reprocessing. - A failed retry republishes to the **next** retry topic; after the last, it goes to the **DLT**. - **Ordering is sacrificed**: the failed record is now reprocessed later, out of order relative to the main stream. - Requires producer access and extra topics (operational/storage cost). ## Decision guide | Factor | Prefer blocking | Prefer non-blocking | |---|---|---| | Strict per-partition ordering required | Yes | No (reordering occurs) | | Failures rare + transient + clear quickly | Yes | — | | Head-of-line blocking unacceptable / high throughput | No | Yes | | Long backoffs (minutes/hours) | No (poll-interval risk) | Yes (delay lives in topics) | | Want minimal infra | Yes | No (extra topics + producer) | ## Combining them A common pattern: configure a **few quick blocking retries** in `DefaultErrorHandler` (cheap recovery from a momentary blip) and then have its recoverer hand off to **non-blocking retry topics** for longer, isolated retrying — getting fast-path recovery without long head-of-line blocking. ## Edge cases - Non-blocking retries can cause **duplicate side effects** if processing isn't idempotent (a record might be processed in main + retry topics). - With `@RetryableTopic`, exceptions can still be classified fatal/non-retryable so they go straight to DLT, skipping retry topics. - Non-blocking delays are approximate: the retry consumer wakes, checks the record's original timestamp + backoff, and re-seeks/pauses if the delay hasn't elapsed.

  • Why does non-blocking retry break ordering guarantees?
    Because the failed record is committed and republished to a separate retry topic, then reprocessed later by a different consumer. Records that came after it in the main partition are already processed, so the failed record's effects now apply out of original order.
  • How can a long blocking backoff cause a rebalance?
    Blocking retries keep the consumer away from poll(). If cumulative backoff exceeds max.poll.interval.ms (default 5 min), the broker assumes the consumer died and triggers a rebalance, reassigning partitions and causing duplicate processing.

saying these in an interview costs you the question

  • Saying non-blocking retries preserve ordering — they don't; the record is reprocessed out of order.
  • Saying blocking retries scale better for long delays — long delays risk max.poll.interval.ms rebalances.
  • Claiming @RetryableTopic needs no extra topics — it auto-creates retry topics and a DLT.
  • Assuming non-blocking retries are automatically safe — duplicates require idempotent processing.

context