What is chunk-oriented processing in a Spring Batch step, and what does the commit-interval control?
answer
- read-N, process-N, write-once, commit-once
- commit-interval = chunk size = n
- one transaction per chunk
- writer gets the whole list
- processor null = filter
basics
~20 sA Spring Batch step reads items one by one, optionally processes each, and writes them in batches. The commit-interval (the chunk size) is how many items are grouped and written together in one transaction before committing.
solid answer
~40 sChunk-oriented processing is Spring Batch's default step model for high-volume data. The step reads items one at a time via an ItemReader, optionally transforms each via an ItemProcessor, and buffers them until it reaches the chunk size (the commit-interval). Then the whole buffered list is handed to the ItemWriter in a single write call, and the surrounding transaction commits. You configure it with StepBuilder.chunk(n, transactionManager). So chunk(10) means: read 10, process 10, write once (a list of 10), commit — then repeat. Writing in batches lets the writer do bulk operations (e.g. JDBC batch inserts) and keeps one transaction per chunk instead of one per row, which is what makes it scale.
code
java · 13 lines@Bean
public Step importStep(JobRepository jobRepository,
PlatformTransactionManager transactionManager,
ItemReader<Person> reader,
ItemProcessor<Person, PersonDto> processor,
ItemWriter<PersonDto> writer) {
return new StepBuilder("importStep", jobRepository)
.<Person, PersonDto>chunk(100, transactionManager) // commit-interval = 100
.reader(reader) // read() called up to 100 times
.processor(processor) // process(item) called per item
.writer(writer) // write(chunk) called ONCE per 100
.build();
}go deeper
Must be able to state read-many/write-once and that commit-interval is the chunk size.
Should mention one transaction per chunk and that the writer receives a list for bulk writes.
Explains filtered-item counting, partial last chunk, and throughput reasoning behind batching the write.
Frames chunk model as the scalability/restartability tradeoff vs per-row or whole-job transactions.
**The problem it solves.** Batch jobs move large volumes of data (millions of rows). Doing one transaction per row is catastrophically slow (commit overhead per item); doing one transaction for the whole job risks huge rollbacks, memory blowups, and no restartability. Spring Batch's *chunk-oriented processing* is the middle ground: process data in fixed-size groups called **chunks**, one transaction per chunk. **The three collaborators.** - **ItemReader<I>** — returns items one at a time via `read()`; returns `null` to signal end of input (e.g. `JdbcCursorItemReader`, `FlatFileItemReader`). - **ItemProcessor<I,O>** — optional; transforms/validates/filters one item at a time via `process(item)`. Returning `null` *filters* the item (it is skipped, not written). - **ItemWriter<O>** — writes a whole list at once via `write(Chunk<O> items)` (in Spring Batch 5; `List` in 4.x). Called **once per chunk**, enabling bulk/batch writes. **The loop.** For a commit-interval of `n`, one chunk iteration is: call `read()` up to `n` times (or until it returns `null`); for each item read, call `process(item)`; collect the non-null processed results; call `write(...)` once with that collected list; commit the transaction. Then start the next chunk. Concisely: **read-N, process-N, write-once, commit-once, repeat.** **commit-interval = chunk size = the `n`.** Historically the XML attribute was literally `commit-interval`. In the Java DSL you pass it to `chunk(n, transactionManager)`. It is the number of *items read* that trigger a commit — not the number written (filtered items still count toward the interval). **Why batch the write.** A single `write(list)` lets the `ItemWriter` issue a JDBC batch (`addBatch`/`executeBatch`), a single bulk HTTP call, etc. That amortizes per-round-trip cost across `n` items and is the main throughput lever. **Configuration (Spring Batch 5).** ```java new StepBuilder("step", jobRepository) .<Input, Output>chunk(100, transactionManager) .reader(reader).processor(processor).writer(writer).build(); ``` The `<Input, Output>` generics are the reader's input type and the writer's output type; the processor bridges them. **Edge cases / gotchas.** - **Last chunk is partial.** If the reader exhausts mid-chunk (read returns `null`), the accumulated items are written and committed even though fewer than `n`. - **Filtered items.** A processor returning `null` removes the item from the written list but it still counts toward the commit-interval read count. Filtered counts show up in step statistics as `filterCount`. - **Reader/processor are per-item; writer is per-chunk.** A common mistake is expecting the writer to be called once per item. - **Chunk size is a tuning knob.** Too small = commit overhead dominates; too large = big transactions, high memory, larger rollback/replay on failure. **When to use.** Chunk-oriented is the default and right choice for read→transform→write pipelines. The alternative, a **tasklet step** (`Tasklet.execute`), is for single non-item operations (run a stored proc, move a file, send a notification).
- If chunk size is 100 but the file has 250 rows, how many times is the writer called and how big is each write?Three times: two writes of 100, then one final partial write of 50 when the reader returns null. Three transactions/commits total.
- What happens to an item when the ItemProcessor returns null?It is filtered — excluded from the list passed to the writer — but it still counts toward the commit-interval. It is recorded as filterCount in the step execution statistics.
saying these in an interview costs you the question
- Saying the writer is invoked once per item rather than once per chunk
- Thinking filtered (processor-null) items don't count toward the commit-interval
- Confusing commit-interval with a time interval