skip to content

When a fault-tolerant Spring Batch step rolls back a chunk, why does it re-read the chunk item-by-item, and how does that isolate the bad item?

level: middleimportance: must knowfreq 50%

answer

  1. Bulk mode hides which item failed
  2. Rollback -> replay cached items one-by-one
  3. Chunk size effectively becomes 1 during scan
  4. Reader not re-hit; items buffered
  5. Failing chunk = N transactions, slow

basics

~20 s

After a rollback Batch can't tell which item failed, since all N were in one transaction. So it re-processes the chunk one item at a time (chunk size effectively 1). The item that throws again is identified as the bad one and gets skipped or retried; the others commit.

solid answer

~40 s

In normal (bulk) mode the whole chunk is processed and written together, so when it fails Batch only knows 'something in this chunk failed', not which item. On a fault-tolerant step it caches the items it read, rolls back, then switches into a single-item 'scan' mode: it replays the cached items one at a time, each in its own transaction. The item that throws again is the culprit — Batch applies the configured skip or retry policy to just that item, while the good items are written and committed individually. Once the chunk is drained it returns to bulk mode. This is why the ItemProcessor must be idempotent: items in a rolled-back chunk get processed more than once. It also makes a failing chunk much slower than a healthy one.

go deeper

for a junior

Aware that after a failure the chunk is retried in a slower one-by-one mode.

for a middle

Explains bulk-vs-scan mode, buffering of read items, and why the processor must be idempotent.

for a senior

Ties scanning to skip/retry policy application and the throughput cost of high skip rates.

for a principal

Reasons about chunk sizing, idempotent writers (upserts), and upstream data quality to bound failure cost at scale.

## The problem: bulk mode hides the culprit In steady state a chunk step operates in **bulk mode**: it reads N items, processes all N, and hands the whole list to the `ItemWriter` in one call inside one transaction. That's efficient (one batch INSERT, etc.). But when the write (or a processing step) throws, the transaction rolls back and Batch is left with a list of N items and the knowledge that *one of them* is poison — it has **no idea which**. ## The solution: single-item scanning When the step is `.faultTolerant()` and a rollback happens, Batch does **not** just fail. It has **cached** the items it read for this chunk (fault-tolerant steps buffer read items so they can be replayed without touching the reader again). It then enters a **one-item-at-a-time** processing mode — effectively a chunk size of 1 — and replays the cached items: 1. Take item 1, process it, write it, commit. If it succeeds, move on. 2. Take item 2, process/write/commit. Continue. 3. When it reaches the poison item, it throws again — but now Batch **knows exactly which item** failed because it's the only one in the transaction. 4. Batch applies the configured policy to that one item: `skip(...)` (record it as skipped, roll back just that single-item transaction, continue) or `retry(...)` (attempt it again up to the retry limit). 5. Remaining items after the bad one are likewise processed individually until the chunk is drained, then the step returns to bulk mode for the next chunk. ## Why the reader isn't re-invoked The items were already read and **buffered** by the fault-tolerant infrastructure, so scanning replays from that buffer rather than pulling fresh from the `ItemReader`. This matters because many readers (file cursors, JDBC cursors) can't cheaply re-read. ## Consequences you must know - **Idempotency**: because a rolled-back chunk is re-processed item-by-item, the **`ItemProcessor` (and any side effects it triggers) can run multiple times for the same item**. Processors must be pure/idempotent, or you double-charge, double-email, etc. - **Performance**: a chunk that hits an error degrades from one bulk write to N individual transactions. A high skip rate can tank throughput; that's a signal to fix data upstream or shrink chunk size. - **Retry + rollback interaction**: with `.retry(...)`, the item is retried (each retry can roll back) up to `retryLimit`; if still failing and also skippable, it may then be skipped. - **Skip on read vs process/write**: a skippable exception thrown during **reading** does not need a rollback (nothing was written) — Batch just skips that item and reads the next. Rollback + rescan is specifically the process/write path. ## When it matters in design Choose chunk size knowing that failures are re-scanned: big chunks = great happy-path throughput but expensive failures; small chunks = cheaper failures, more commits. If your writer isn't idempotent, keep chunks small or make the writer upsert-safe.

  • During the single-item scan, does Batch call the ItemReader again for those items?
    No. Fault-tolerant steps buffer the items already read for the chunk and replay from that buffer, so the reader isn't re-invoked. This lets non-rewindable readers (cursors, files) still support skip/retry.
  • What property must the ItemProcessor have because of this rescanning, and why?
    It must be idempotent / free of external side effects, because a rolled-back chunk is re-processed item-by-item, so the same item may pass through the processor multiple times. Non-idempotent side effects (charging a card, sending email) would be duplicated.

saying these in an interview costs you the question

  • Claiming Batch magically knows which item failed without rescanning
  • Thinking the reader is called again for each item during the scan
  • Ignoring that the processor may run multiple times per item

context