skip to content

Read-Process-Write Transaction Boundary

Each chunk is one transaction, so a failure rolls back the whole chunk rather than the whole job, and jobs without a database can use the resourceless manager. Interviewers connect chunk size directly to how much work you lose on a failure.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

explore

questions

5

In a chunk-oriented Spring Batch step, what is the unit of work wrapped in a single transaction?

level: juniorimportance: must knowfreq 70%

answer

  1. one chunk = one transaction
  2. read N, process each, write list once, commit
  3. chunk-size == commit-interval
  4. 1000 items / chunk 100 = 10 commits
  5. amortize commit cost

basics

~10 s

One chunk. Spring Batch reads a group of items (the chunk size), processes them, then writes them all inside a single transaction. Commit happens once per chunk, not per item.

solid answer

~40 s

A chunk-oriented step processes data in read-process-write cycles. It reads items one at a time via an ItemReader up to the configured chunk size (commit-interval), optionally transforms each with an ItemProcessor, then hands the whole batch to the ItemWriter in one call. That entire read-process-write cycle for one chunk runs inside a single transaction managed by the step's PlatformTransactionManager. The transaction commits after the chunk is written and repeats for the next chunk. So if chunk size is 100 and you have 1000 items, you get 10 transactions and 10 commits. This batching is why Batch is efficient: you amortize commit cost across many rows instead of committing per row.

code

java · 15 lines
java
@Bean
public Step importStep(JobRepository jobRepository,
                       PlatformTransactionManager txManager,
                       ItemReader<Person> reader,
                       ItemProcessor<Person, Person> processor,
                       ItemWriter<Person> writer) {
    return new StepBuilder("importStep", jobRepository)
            // chunk size 100: read 100, process each, write the list,
            // then ONE commit on txManager -- repeat for the next 100.
            .<Person, Person>chunk(100, txManager)
            .reader(reader)
            .processor(processor)
            .writer(writer)
            .build();
}

go deeper

for a junior

Must know: one chunk = one transaction, commit per chunk not per item.

for a middle

Should explain read-process-write cycle and that the writer receives the whole list once.

for a senior

Frames chunk size as the commit-interval tuning knob and ties it to commit-cost amortization.

for a principal

Connects chunk boundary to JobRepository metadata participating in the same transaction for restartability consistency.

## The chunk model Spring Batch's primary processing model is **chunk-oriented processing**. A `Step` built with `StepBuilder.chunk(size, transactionManager)` works in cycles. Each cycle: 1. **Read phase** — the framework calls `ItemReader.read()` repeatedly, once per item, until it has accumulated `chunk-size` items (or the reader returns `null`, signalling end of input). 2. **Process phase** — if an `ItemProcessor` is configured, each read item is passed through `process()`. A processor may transform the item, or return `null` to filter it out. 3. **Write phase** — all surviving processed items are passed **as a single list** to `ItemWriter.write(Chunk<T>)` in one call. ## The transaction boundary The **entire cycle for one chunk is wrapped in a single transaction**. Before the read phase begins, Spring Batch starts a transaction via the `PlatformTransactionManager` you supplied to the step. After the write phase succeeds, it **commits**. Then it starts a fresh transaction for the next chunk. Concretely, with `chunk-size = 100`: - 100 reads, up to 100 processes, 1 write call, 1 commit — repeat. - 1000 input items ⇒ 10 transactions, 10 commits. This is the core efficiency of Batch: **commit cost is amortized** over the whole chunk instead of paid per row. ## Chunk size vs commit-interval 'Chunk size' and 'commit-interval' are the same knob — older XML config called it `commit-interval`; the Java `StepBuilder` calls it the `chunk` size. It is the number of items processed per transaction/commit. ## Why the boundary matters Because the boundary is the whole chunk, a failure anywhere in the cycle rolls back **all** items in that chunk together — you never get a partially-written chunk committed to the database. That atomicity per chunk is the guarantee the model gives you. ## Key classes - `ItemReader<T>` / `ItemProcessor<I,O>` / `ItemWriter<T>` — the three phases. - `Chunk<T>` — the list passed to the writer (Spring Batch 5+ replaced the raw `List` with `Chunk`). - `PlatformTransactionManager` — supplied to `StepBuilder.chunk(size, txManager)`; owns the boundary. - `JobRepository` — its own metadata writes (step execution counts) participate in the same transaction so metadata stays consistent with data.

  • How many commits happen for 250 items with a chunk size of 100?
    Three: chunks of 100, 100, and 50. The last chunk is short because the reader returned null before filling 100, and it still gets its own transaction and commit.
  • Does the ItemProcessor returning null cause a rollback?
    No. Returning null from the processor filters the item out of the chunk (it is not written and counts as a filter). It is a normal control-flow signal, not an error, so the transaction proceeds and commits.

saying these in an interview costs you the question

  • Saying each item gets its own transaction/commit
  • Thinking the writer is called once per item rather than once per chunk with a list
  • Confusing processor returning null (a filter) with an exception (a rollback)

context

open as a page

If writing the 50th item of a 100-item chunk throws, what happens to the other 99 items in that chunk?

level: middleimportance: must knowfreq 68%

basics

~10 s

The whole chunk's transaction rolls back, so none of the 100 items are committed. By default the step fails and stops. All 100 are undone together — the chunk is atomic.

open as a page

You have a Spring Batch step that reads a CSV and calls a REST API with no database at all. Which PlatformTransactionManager should the chunk use, and why?

level: seniorimportance: should knowfreq 40%

basics

~10 s

Use ResourcelessTransactionManager. Chunk steps require a PlatformTransactionManager, but with no database there is no real transactional resource to manage, so this no-op manager satisfies the API without touching a datasource.

open as a page

What is the transactionAttribute on a chunk step, and what can you tune with it?

level: seniorimportance: should knowfreq 45%

basics

~10 s

It configures the chunk transaction's properties — propagation, isolation level, timeout, and read-only flag — via a DefaultTransactionAttribute. You set it on the StepBuilder to override defaults like isolation or a per-chunk timeout.

open as a page

How do you keep non-transactional side effects (emails, REST calls, external queues) from being duplicated or corrupted when a chunk transaction rolls back?

level: principalimportance: should knowfreq 33%

basics

~20 s

Rollback only undoes transactional resources, not emails or REST calls. Make those side effects idempotent, or defer them until after commit, or stage them in the same transactional store and process later. Never assume rollback un-does external effects.

open as a page