In a Spring Batch chunk-oriented step, what does retry(Exception.class).retryLimit(n) do, and when would you use it?
answer
- faultTolerant() first
- retry(type) + retryLimit(n)
- transient not deterministic
- builds SimpleRetryPolicy
- rolls back chunk and re-attempts
basics
~20 sOn a faultTolerant() step, retry(SomeException.class) marks that exception as retryable and retryLimit(n) caps the number of attempts. Spring Batch re-runs the failing work instead of failing the job. Use it for transient errors like a brief network or DB blip.
solid answer
~40 sYou enable it with faultTolerant() on the chunk step, then retry(TransientException.class) tells Spring Batch that exception type is worth retrying, and retryLimit(n) sets the maximum attempts. When an item throws a retryable exception during processing or writing, Batch rolls back the chunk transaction and re-attempts the work up to the limit; if it eventually succeeds the step continues, and only if all attempts fail does the exception propagate (or fall through to skip when skip is also configured). It targets transient, self-healing failures: a lock timeout, a flaky HTTP call, a deadlock. You should NOT retry deterministic failures like a validation or parse error, because they fail identically every attempt and just waste time. Configuration lives on FaultTolerantStepBuilder via retry(), retryLimit(), retryPolicy().
code
java · 15 lines@Bean
Step step(JobRepository jobRepository, PlatformTransactionManager tx,
ItemReader<Order> reader, ItemProcessor<Order, Invoice> processor,
ItemWriter<Invoice> writer) {
return new StepBuilder("invoiceStep", jobRepository)
.<Order, Invoice>chunk(10, tx)
.reader(reader)
.processor(processor)
.writer(writer)
.faultTolerant()
.retry(TransientDataAccessException.class) // retryable type
.retry(HttpServerErrorException.class)
.retryLimit(3) // up to 3 attempts total
.build();
}go deeper
Know that retry() marks an exception retryable and retryLimit caps attempts, and that it is for transient failures.
Add that it requires faultTolerant(), rolls back the chunk transaction, and applies to processor/writer phases.
Discuss idempotency risks on retried writes and choosing retryable vs skippable exception classification.
Frame retry as one layer of a resilience strategy (retry -> skip -> restart) and weigh it against downstream backpressure and non-idempotent side effects.
### What retry is for Spring Batch processes data in **chunks**: it reads N items, processes them, and writes them all in one transaction. If one item throws an exception, by default the whole step **fails**. Fault tolerance handles failures more gracefully. **Retry** targets *transient* failures — errors likely to succeed if you simply try again a moment later (a network hiccup, a DB deadlock, a lock/timeout, a momentarily-unavailable downstream service). ### How you configure it Retry only works after you call `faultTolerant()` on the chunk-step builder, which switches the step to the fault-tolerant implementation. Then: - `retry(Class<? extends Throwable>)` — registers an exception type as **retryable**. - `retryLimit(int)` — the maximum number of attempts (see the maxAttempts gotcha below). Under the hood these two calls build a `SimpleRetryPolicy` (from Spring Retry) whose `retryableExceptions` map contains your class and whose `maxAttempts` equals your `retryLimit`. ### What happens at runtime When a retryable exception is thrown, Spring Batch **rolls back the current chunk's transaction** and re-attempts. Retry is applied around the **processor** and **writer** phases (not the reader by default). If a later attempt succeeds, the step proceeds as if nothing happened. If every attempt fails, the exception is rethrown — which either fails the step, or, if the same exception type is also registered with `skip(...)`, the item is skipped instead. ### When to use / not use - **Use** for idempotent, transient operations: remote calls, optimistic-lock conflicts (`OptimisticLockingFailureException`), deadlocks (`DeadlockLoserDataAccessException`). - **Avoid** for deterministic errors (`FlatFileParseException`, bean validation): retrying just repeats the same failure. Skip those instead. - Be careful retrying **non-idempotent writes** — a retried write can double-apply side effects. ### Key gotcha `retryLimit(n)` is the **total attempts including the first**, not n *extra* retries. `retryLimit(1)` means one attempt and effectively no retry.
- Does retry work on the reader phase too?By default retry is applied around the processor and writer, not the reader. A read failure is handled by skip/restart semantics, not chunk retry, because re-reading is driven by the input stream's cursor/position.
- What kinds of exceptions are bad candidates for retry?Deterministic ones — parse errors, validation failures, constraint violations on truly bad data. They will fail identically every attempt. Use skip for those; reserve retry for transient/self-healing failures.
saying these in an interview costs you the question
- Thinking retry re-reads or retries the reader by default
- Retrying deterministic errors like validation/parse failures
- Believing retryLimit(n) means n retries after the first attempt
- Assuming retry works without calling faultTolerant()