skip to content

In a Spring Batch chunk-oriented step, what does retry(Exception.class).retryLimit(n) do, and when would you use it?

level: juniorimportance: must knowfreq 70%

answer

  1. faultTolerant() first
  2. retry(type) + retryLimit(n)
  3. transient not deterministic
  4. builds SimpleRetryPolicy
  5. rolls back chunk and re-attempts

basics

~20 s

On a faultTolerant() step, retry(SomeException.class) marks that exception as retryable and retryLimit(n) caps the number of attempts. Spring Batch re-runs the failing work instead of failing the job. Use it for transient errors like a brief network or DB blip.

solid answer

~40 s

You enable it with faultTolerant() on the chunk step, then retry(TransientException.class) tells Spring Batch that exception type is worth retrying, and retryLimit(n) sets the maximum attempts. When an item throws a retryable exception during processing or writing, Batch rolls back the chunk transaction and re-attempts the work up to the limit; if it eventually succeeds the step continues, and only if all attempts fail does the exception propagate (or fall through to skip when skip is also configured). It targets transient, self-healing failures: a lock timeout, a flaky HTTP call, a deadlock. You should NOT retry deterministic failures like a validation or parse error, because they fail identically every attempt and just waste time. Configuration lives on FaultTolerantStepBuilder via retry(), retryLimit(), retryPolicy().

code

java · 15 lines
java
@Bean
Step step(JobRepository jobRepository, PlatformTransactionManager tx,
          ItemReader<Order> reader, ItemProcessor<Order, Invoice> processor,
          ItemWriter<Invoice> writer) {
    return new StepBuilder("invoiceStep", jobRepository)
            .<Order, Invoice>chunk(10, tx)
            .reader(reader)
            .processor(processor)
            .writer(writer)
            .faultTolerant()
            .retry(TransientDataAccessException.class)   // retryable type
            .retry(HttpServerErrorException.class)
            .retryLimit(3)                                // up to 3 attempts total
            .build();
}

go deeper

for a junior

Know that retry() marks an exception retryable and retryLimit caps attempts, and that it is for transient failures.

for a middle

Add that it requires faultTolerant(), rolls back the chunk transaction, and applies to processor/writer phases.

for a senior

Discuss idempotency risks on retried writes and choosing retryable vs skippable exception classification.

for a principal

Frame retry as one layer of a resilience strategy (retry -> skip -> restart) and weigh it against downstream backpressure and non-idempotent side effects.

### What retry is for Spring Batch processes data in **chunks**: it reads N items, processes them, and writes them all in one transaction. If one item throws an exception, by default the whole step **fails**. Fault tolerance handles failures more gracefully. **Retry** targets *transient* failures — errors likely to succeed if you simply try again a moment later (a network hiccup, a DB deadlock, a lock/timeout, a momentarily-unavailable downstream service). ### How you configure it Retry only works after you call `faultTolerant()` on the chunk-step builder, which switches the step to the fault-tolerant implementation. Then: - `retry(Class<? extends Throwable>)` — registers an exception type as **retryable**. - `retryLimit(int)` — the maximum number of attempts (see the maxAttempts gotcha below). Under the hood these two calls build a `SimpleRetryPolicy` (from Spring Retry) whose `retryableExceptions` map contains your class and whose `maxAttempts` equals your `retryLimit`. ### What happens at runtime When a retryable exception is thrown, Spring Batch **rolls back the current chunk's transaction** and re-attempts. Retry is applied around the **processor** and **writer** phases (not the reader by default). If a later attempt succeeds, the step proceeds as if nothing happened. If every attempt fails, the exception is rethrown — which either fails the step, or, if the same exception type is also registered with `skip(...)`, the item is skipped instead. ### When to use / not use - **Use** for idempotent, transient operations: remote calls, optimistic-lock conflicts (`OptimisticLockingFailureException`), deadlocks (`DeadlockLoserDataAccessException`). - **Avoid** for deterministic errors (`FlatFileParseException`, bean validation): retrying just repeats the same failure. Skip those instead. - Be careful retrying **non-idempotent writes** — a retried write can double-apply side effects. ### Key gotcha `retryLimit(n)` is the **total attempts including the first**, not n *extra* retries. `retryLimit(1)` means one attempt and effectively no retry.

  • Does retry work on the reader phase too?
    By default retry is applied around the processor and writer, not the reader. A read failure is handled by skip/restart semantics, not chunk retry, because re-reading is driven by the input stream's cursor/position.
  • What kinds of exceptions are bad candidates for retry?
    Deterministic ones — parse errors, validation failures, constraint violations on truly bad data. They will fail identically every attempt. Use skip for those; reserve retry for transient/self-healing failures.

saying these in an interview costs you the question

  • Thinking retry re-reads or retries the reader by default
  • Retrying deterministic errors like validation/parse failures
  • Believing retryLimit(n) means n retries after the first attempt
  • Assuming retry works without calling faultTolerant()

context