skip to content

In a chunk-oriented Spring Batch step, what is the unit of work wrapped in a single transaction?

level: juniorimportance: must knowfreq 70%

answer

  1. one chunk = one transaction
  2. read N, process each, write list once, commit
  3. chunk-size == commit-interval
  4. 1000 items / chunk 100 = 10 commits
  5. amortize commit cost

basics

~10 s

One chunk. Spring Batch reads a group of items (the chunk size), processes them, then writes them all inside a single transaction. Commit happens once per chunk, not per item.

solid answer

~40 s

A chunk-oriented step processes data in read-process-write cycles. It reads items one at a time via an ItemReader up to the configured chunk size (commit-interval), optionally transforms each with an ItemProcessor, then hands the whole batch to the ItemWriter in one call. That entire read-process-write cycle for one chunk runs inside a single transaction managed by the step's PlatformTransactionManager. The transaction commits after the chunk is written and repeats for the next chunk. So if chunk size is 100 and you have 1000 items, you get 10 transactions and 10 commits. This batching is why Batch is efficient: you amortize commit cost across many rows instead of committing per row.

code

java · 15 lines
java
@Bean
public Step importStep(JobRepository jobRepository,
                       PlatformTransactionManager txManager,
                       ItemReader<Person> reader,
                       ItemProcessor<Person, Person> processor,
                       ItemWriter<Person> writer) {
    return new StepBuilder("importStep", jobRepository)
            // chunk size 100: read 100, process each, write the list,
            // then ONE commit on txManager -- repeat for the next 100.
            .<Person, Person>chunk(100, txManager)
            .reader(reader)
            .processor(processor)
            .writer(writer)
            .build();
}

go deeper

for a junior

Must know: one chunk = one transaction, commit per chunk not per item.

for a middle

Should explain read-process-write cycle and that the writer receives the whole list once.

for a senior

Frames chunk size as the commit-interval tuning knob and ties it to commit-cost amortization.

for a principal

Connects chunk boundary to JobRepository metadata participating in the same transaction for restartability consistency.

## The chunk model Spring Batch's primary processing model is **chunk-oriented processing**. A `Step` built with `StepBuilder.chunk(size, transactionManager)` works in cycles. Each cycle: 1. **Read phase** — the framework calls `ItemReader.read()` repeatedly, once per item, until it has accumulated `chunk-size` items (or the reader returns `null`, signalling end of input). 2. **Process phase** — if an `ItemProcessor` is configured, each read item is passed through `process()`. A processor may transform the item, or return `null` to filter it out. 3. **Write phase** — all surviving processed items are passed **as a single list** to `ItemWriter.write(Chunk<T>)` in one call. ## The transaction boundary The **entire cycle for one chunk is wrapped in a single transaction**. Before the read phase begins, Spring Batch starts a transaction via the `PlatformTransactionManager` you supplied to the step. After the write phase succeeds, it **commits**. Then it starts a fresh transaction for the next chunk. Concretely, with `chunk-size = 100`: - 100 reads, up to 100 processes, 1 write call, 1 commit — repeat. - 1000 input items ⇒ 10 transactions, 10 commits. This is the core efficiency of Batch: **commit cost is amortized** over the whole chunk instead of paid per row. ## Chunk size vs commit-interval 'Chunk size' and 'commit-interval' are the same knob — older XML config called it `commit-interval`; the Java `StepBuilder` calls it the `chunk` size. It is the number of items processed per transaction/commit. ## Why the boundary matters Because the boundary is the whole chunk, a failure anywhere in the cycle rolls back **all** items in that chunk together — you never get a partially-written chunk committed to the database. That atomicity per chunk is the guarantee the model gives you. ## Key classes - `ItemReader<T>` / `ItemProcessor<I,O>` / `ItemWriter<T>` — the three phases. - `Chunk<T>` — the list passed to the writer (Spring Batch 5+ replaced the raw `List` with `Chunk`). - `PlatformTransactionManager` — supplied to `StepBuilder.chunk(size, txManager)`; owns the boundary. - `JobRepository` — its own metadata writes (step execution counts) participate in the same transaction so metadata stays consistent with data.

  • How many commits happen for 250 items with a chunk size of 100?
    Three: chunks of 100, 100, and 50. The last chunk is short because the reader returned null before filling 100, and it still gets its own transaction and commit.
  • Does the ItemProcessor returning null cause a rollback?
    No. Returning null from the processor filters the item out of the chunk (it is not written and counts as a filter). It is a normal control-flow signal, not an error, so the transaction proceeds and commits.

saying these in an interview costs you the question

  • Saying each item gets its own transaction/commit
  • Thinking the writer is called once per item rather than once per chunk with a list
  • Confusing processor returning null (a filter) with an exception (a rollback)

context