skip to content

Rollback Semantics

When a chunk fails, its transaction rolls back and the framework re-reads the items one at a time to find the offending one, unless you exclude the exception from rollback. Understanding this scan explains why a fault-tolerant step is slower after an error.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

explore

questions

5

In Spring Batch chunk-oriented processing, what happens to the current chunk's transaction when an exception is thrown while processing or writing an item?

level: juniorimportance: must knowfreq 55%

answer

  1. Chunk = transaction boundary
  2. One bad item rolls back the whole chunk
  3. commit-interval == chunk size
  4. Reader non-transactional
  5. faultTolerant() -> re-read one-by-one

basics

~20 s

The whole chunk's transaction rolls back, so none of the items in that chunk get written. Spring Batch wraps each chunk (read-process-write of N items) in a single transaction, so one bad item undoes the entire chunk.

solid answer

~40 s

A chunk-oriented Step reads N items, processes them, then writes them all in one transaction (the commit interval = chunk size). If any item throws during processing or writing, that surrounding transaction rolls back, so every item in the chunk is discarded, not just the failing one. The reader is non-transactional, so reads aren't 'undone', but no items are persisted. Without fault tolerance the exception propagates and the Step fails. With faultTolerant() plus skip/retry configured, Batch rolls back and then re-reads the chunk item-by-item to isolate which item is bad, so the good items still succeed. The key mental model: the transaction boundary is the chunk, not the individual item.

code

java · 15 lines
java
@Bean
public Step chunkStep(JobRepository jobRepository,
                       PlatformTransactionManager txManager,
                       ItemReader<Order> reader,
                       ItemProcessor<Order, Order> processor,
                       ItemWriter<Order> writer) {
    return new StepBuilder("chunkStep", jobRepository)
            // chunk(10) => read/process/write 10 items per transaction.
            // If any of the 10 fails, all 10 roll back.
            .<Order, Order>chunk(10, txManager)
            .reader(reader)
            .processor(processor)
            .writer(writer)
            .build();
}

go deeper

for a junior

Must know that the chunk (not the item) is the transaction unit and that one failure discards the whole chunk.

for a middle

Should connect chunk size = commit interval and know fault tolerance changes recovery to single-item scanning.

for a senior

Explains reader non-transactionality, restart semantics, and the throughput-vs-failure-cost trade-off of chunk size.

for a principal

Frames chunk sizing and rollback cost as an architectural knob affecting idempotency requirements and SLA under partial failure.

## Chunk-oriented processing and its transaction Spring Batch's most common Step type is **chunk-oriented**: it repeatedly (1) **reads** items one at a time via an `ItemReader` until it has accumulated a *chunk* of `commit-interval` items (the number passed to `.chunk(N, tx)`), (2) optionally **processes** each item via an `ItemProcessor`, then (3) **writes** the whole chunk at once via an `ItemWriter`. The critical fact: **the entire read-process-write cycle for one chunk runs inside a single Spring-managed transaction**, driven by the `PlatformTransactionManager` you pass to the step. So the transaction boundary is the *chunk*, not the individual item. The chunk size and the commit interval are the same number. ## What rollback means here If **any** item throws an exception during processing or writing, the surrounding transaction is marked rollback-only and rolled back. Consequences: - **None** of the items in that chunk are committed to the target resource (DB, etc.) — not just the failing one. - The `ItemReader` is **not** transactional (reading from a file/DB cursor doesn't 'un-read'), so the read position is a separate concern handled by Batch's state management, but no *output* is persisted. - Batch also rolls back its own bookkeeping for that chunk so the `StepExecution` write count isn't advanced for the failed chunk. ## Without fault tolerance By default a plain chunk step is **not** fault tolerant. The first exception propagates out, the transaction rolls back, the Step ends in `FAILED`, and the Job stops. On restart it resumes from the last committed chunk. ## With fault tolerance Calling `.faultTolerant()` on the `StepBuilder` and configuring `.skip(...)` and/or `.retry(...)` changes the recovery: after the rollback, Batch **re-reads the cached chunk and re-processes/re-writes it one item at a time** to pinpoint the offending item, so the healthy items in the chunk still commit and only the bad one is skipped or retried. That single-item scanning is the subject of the deeper questions on this topic. ## Gotchas - A larger chunk size means more items are thrown away per rollback and re-scanned — a trade-off between throughput and the cost of a failure. - Because a rolled-back chunk is re-processed, an `ItemProcessor` may run more than once for the same item; it should be **idempotent / side-effect-free**. - Rollback is per-chunk, so partial DB writes from a mid-chunk failure never become visible.

  • If chunk size is 10 and item #7 fails during write, how many items are committed for that chunk?
    Zero. The transaction spans all 10 items, so the whole chunk rolls back. Only with fault tolerance (skip/retry) will Batch then re-scan the chunk one-by-one so the good items eventually commit and #7 is skipped/retried.
  • Is the ItemReader part of the rolled-back transaction?
    The read itself is not transactional — you can't 'un-read' a file line or cursor row. Batch manages restart state separately. Rollback concerns the processed/written output, none of which is persisted for the failed chunk.

saying these in an interview costs you the question

  • Thinking only the single failing item is rolled back while the rest of the chunk commits
  • Believing each item has its own transaction by default
  • Assuming the reader participates in / gets rolled back by the chunk transaction

context

open as a page

When a fault-tolerant Spring Batch step rolls back a chunk, why does it re-read the chunk item-by-item, and how does that isolate the bad item?

level: middleimportance: must knowfreq 50%

basics

~20 s

After a rollback Batch can't tell which item failed, since all N were in one transaction. So it re-processes the chunk one item at a time (chunk size effectively 1). The item that throws again is identified as the bad one and gets skipped or retried; the others commit.

open as a page

What does .noRollback(Exception.class) do on a fault-tolerant Spring Batch step, and when would you use it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

noRollback tells the step that a given exception type should NOT trigger a transaction rollback. Use it when a processor throws a business/validation exception you want to handle (e.g. skip that item) without discarding the whole chunk's transaction. It applies to processor exceptions, not writer exceptions.

open as a page

Given Spring Batch re-processes a rolled-back chunk item-by-item, what does that imply for an ItemProcessor that has side effects, and how do you handle it?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Because a rolled-back chunk is replayed one item at a time, the same item can pass through the ItemProcessor more than once. Any side effect there (charging a card, sending email, calling an external API) would happen repeatedly. Keep the processor idempotent or pure and push non-idempotent effects to an idempotent writer.

open as a page

As a principal engineer, how would you decide when to use noRollback and how to size chunks, given the rollback-and-rescan cost model of fault-tolerant Spring Batch steps?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Use noRollback only for processor-phase exceptions where nothing transactional was written and re-scanning is wasted work — typically validation skips. Size chunks by balancing happy-path throughput against the cost of rolling back and re-scanning a whole chunk on failure, given your data's error rate and writer idempotency.

open as a page