How exactly does Spring Batch re-run work when a retryable exception is thrown mid-chunk, and does retryLimit(3) mean three retries or three total attempts?
answer
- maxAttempts includes first try
- retryLimit(3) = 1 + 2 retries
- writer failure -> single-item scan
- per-item RetryContextCache
- processor/writer may run many times -> idempotency
basics
~20 sretryLimit(3) means three total attempts (the first try plus two retries), because it maps to SimpleRetryPolicy.maxAttempts, which counts the initial call. On a retryable failure Batch rolls back the chunk transaction and reprocesses the items again, up to that limit.
solid answer
~40 sretryLimit(n) sets SimpleRetryPolicy.maxAttempts, and maxAttempts includes the very first invocation — so retryLimit(3) is one initial attempt plus two retries, and retryLimit(1) disables retry. At runtime, when a retryable exception surfaces during processing or writing, Batch rolls back the chunk's transaction and re-drives the chunk. During the write phase specifically, once retries are exhausted (or immediately, depending on config) Batch may switch to single-item scanning: it re-writes the chunk one item at a time to isolate which item is failing, so good items in the same chunk are not lost. This item-by-item reprocessing means your processor and writer can be invoked multiple times for the same item, which is why idempotency and cache-safe processors matter. The retry state is keyed per item via a RetryContextCache so counts survive across the rollback/re-drive.
code
java · 10 lines// retryLimit(3) => SimpleRetryPolicy(maxAttempts=3): 1 initial call + 2 retries
return new StepBuilder("step", jobRepository)
.<In, Out>chunk(50, tx)
.reader(reader).processor(processor).writer(writer)
.faultTolerant()
.retry(OptimisticLockingFailureException.class)
.retryLimit(3)
// processor is expensive + idempotent-unsafe to re-run during scan:
.processorNonTransactional()
.build();go deeper
Just remember retryLimit(3) is three total attempts, not three extra ones.
Explain the rollback-and-re-drive cycle and the writer scan-mode single-item reprocessing.
Connect scan mode to idempotency requirements and processorNonTransactional, and to performance under large chunks.
Reason about worst-case throughput collapse to item-at-a-time writes and design chunk sizing plus idempotent writers accordingly.
### maxAttempts counts the first try `retryLimit(n)` configures a `SimpleRetryPolicy` whose `maxAttempts = n`. In Spring Retry, **maxAttempts is the total number of attempts including the initial one**. So: - `retryLimit(1)` → 1 attempt, no retry. - `retryLimit(3)` → 1 initial + 2 retries = 3 invocations max. This off-by-one is a classic interview trap. ### The rollback-and-re-drive cycle Because a chunk executes inside a single transaction, a failure means the transaction must roll back. Spring Batch then **re-drives the chunk**. What re-runs depends on where the exception was thrown: - **Processor failure:** the items are re-processed (and re-written) on the next attempt. - **Writer failure:** Batch cannot know which item in the batch write caused the failure, so after rollback it enters **scan mode** — it re-processes and re-writes the chunk **one item at a time** with chunk size effectively 1, so it can pinpoint the offending item. Successful items commit; the failing item continues to retry (and then may be skipped if skip is configured). ### Per-item retry state Retry counts are tracked per item, not per chunk, using a **RetryContextCache** and an item **key generator** (identity by default). This is how the count persists across the transaction rollback, and why you can override `keyGenerator(...)` if your items need value-based identity. ### Idempotency consequences Because processors and writers can be called **multiple times for the same item**, your processor should be side-effect-free and your writer should tolerate re-writes (upserts, dedupe keys). A processor that mutates external state or increments a counter will misbehave under retry. If your processor is expensive and safe to run outside the rollback boundary, `processorNonTransactional()` avoids re-processing already-processed items during scan. ### Gotchas - Setting both `retryLimit()` and a custom `retryPolicy()` conflicts — use one. - Retry with a large chunk plus writer failures degrades to single-item writes (scan), which can be much slower; size chunks with that worst case in mind. - There is **no delay between retries by default** — retries are immediate unless you add backoff via Spring Retry (covered separately).
- Why does a writer failure cause single-item processing?A batched write fails as a unit, so Batch cannot tell which item broke it. After rollback it re-drives the chunk one item at a time (scan mode) to isolate the culprit, letting the good items commit and the bad one retry or skip.
- How does the retry count survive a transaction rollback?Retry state is stored per item in a RetryContextCache keyed by an item identity (default identity, overridable via keyGenerator), independent of the DB transaction, so counts persist across re-drives.
saying these in an interview costs you the question
- Saying retryLimit(3) allows three retries after the first attempt
- Assuming the whole chunk is discarded and never re-processed item-by-item
- Ignoring that processors/writers may execute multiple times per item
- Expecting a built-in delay between retries