skip to content

In Spring Batch chunk-oriented processing, what happens to the current chunk's transaction when an exception is thrown while processing or writing an item?

level: juniorimportance: must knowfreq 55%

answer

  1. Chunk = transaction boundary
  2. One bad item rolls back the whole chunk
  3. commit-interval == chunk size
  4. Reader non-transactional
  5. faultTolerant() -> re-read one-by-one

basics

~20 s

The whole chunk's transaction rolls back, so none of the items in that chunk get written. Spring Batch wraps each chunk (read-process-write of N items) in a single transaction, so one bad item undoes the entire chunk.

solid answer

~40 s

A chunk-oriented Step reads N items, processes them, then writes them all in one transaction (the commit interval = chunk size). If any item throws during processing or writing, that surrounding transaction rolls back, so every item in the chunk is discarded, not just the failing one. The reader is non-transactional, so reads aren't 'undone', but no items are persisted. Without fault tolerance the exception propagates and the Step fails. With faultTolerant() plus skip/retry configured, Batch rolls back and then re-reads the chunk item-by-item to isolate which item is bad, so the good items still succeed. The key mental model: the transaction boundary is the chunk, not the individual item.

code

java · 15 lines
java
@Bean
public Step chunkStep(JobRepository jobRepository,
                       PlatformTransactionManager txManager,
                       ItemReader<Order> reader,
                       ItemProcessor<Order, Order> processor,
                       ItemWriter<Order> writer) {
    return new StepBuilder("chunkStep", jobRepository)
            // chunk(10) => read/process/write 10 items per transaction.
            // If any of the 10 fails, all 10 roll back.
            .<Order, Order>chunk(10, txManager)
            .reader(reader)
            .processor(processor)
            .writer(writer)
            .build();
}

go deeper

for a junior

Must know that the chunk (not the item) is the transaction unit and that one failure discards the whole chunk.

for a middle

Should connect chunk size = commit interval and know fault tolerance changes recovery to single-item scanning.

for a senior

Explains reader non-transactionality, restart semantics, and the throughput-vs-failure-cost trade-off of chunk size.

for a principal

Frames chunk sizing and rollback cost as an architectural knob affecting idempotency requirements and SLA under partial failure.

## Chunk-oriented processing and its transaction Spring Batch's most common Step type is **chunk-oriented**: it repeatedly (1) **reads** items one at a time via an `ItemReader` until it has accumulated a *chunk* of `commit-interval` items (the number passed to `.chunk(N, tx)`), (2) optionally **processes** each item via an `ItemProcessor`, then (3) **writes** the whole chunk at once via an `ItemWriter`. The critical fact: **the entire read-process-write cycle for one chunk runs inside a single Spring-managed transaction**, driven by the `PlatformTransactionManager` you pass to the step. So the transaction boundary is the *chunk*, not the individual item. The chunk size and the commit interval are the same number. ## What rollback means here If **any** item throws an exception during processing or writing, the surrounding transaction is marked rollback-only and rolled back. Consequences: - **None** of the items in that chunk are committed to the target resource (DB, etc.) — not just the failing one. - The `ItemReader` is **not** transactional (reading from a file/DB cursor doesn't 'un-read'), so the read position is a separate concern handled by Batch's state management, but no *output* is persisted. - Batch also rolls back its own bookkeeping for that chunk so the `StepExecution` write count isn't advanced for the failed chunk. ## Without fault tolerance By default a plain chunk step is **not** fault tolerant. The first exception propagates out, the transaction rolls back, the Step ends in `FAILED`, and the Job stops. On restart it resumes from the last committed chunk. ## With fault tolerance Calling `.faultTolerant()` on the `StepBuilder` and configuring `.skip(...)` and/or `.retry(...)` changes the recovery: after the rollback, Batch **re-reads the cached chunk and re-processes/re-writes it one item at a time** to pinpoint the offending item, so the healthy items in the chunk still commit and only the bad one is skipped or retried. That single-item scanning is the subject of the deeper questions on this topic. ## Gotchas - A larger chunk size means more items are thrown away per rollback and re-scanned — a trade-off between throughput and the cost of a failure. - Because a rolled-back chunk is re-processed, an `ItemProcessor` may run more than once for the same item; it should be **idempotent / side-effect-free**. - Rollback is per-chunk, so partial DB writes from a mid-chunk failure never become visible.

  • If chunk size is 10 and item #7 fails during write, how many items are committed for that chunk?
    Zero. The transaction spans all 10 items, so the whole chunk rolls back. Only with fault tolerance (skip/retry) will Batch then re-scan the chunk one-by-one so the good items eventually commit and #7 is skipped/retried.
  • Is the ItemReader part of the rolled-back transaction?
    The read itself is not transactional — you can't 'un-read' a file line or cursor row. Batch manages restart state separately. Rollback concerns the processed/written output, none of which is persisted for the failed chunk.

saying these in an interview costs you the question

  • Thinking only the single failing item is rolled back while the rest of the chunk commits
  • Believing each item has its own transaction by default
  • Assuming the reader participates in / gets rolled back by the chunk transaction

context