skip to content

What is an ItemWriter in Spring Batch, and why does its write method take a whole Chunk instead of a single item?

level: juniorimportance: must knowfreq 75%

answer

  1. write(Chunk) once per chunk, not per item
  2. commit interval = chunk = batch size
  3. batched I/O amortizes round-trips
  4. whole chunk in one transaction
  5. Spring Batch 5: Chunk replaced List

basics

~20 s

An ItemWriter is the step component that outputs processed data. Its write method receives a Chunk (a batch of items) so it can write many items in one operation, like a single batched SQL insert, which is faster than one call per item.

solid answer

~30 s

ItemWriter is the third stage of a chunk-oriented step (read, process, write). Unlike ItemReader which is called once per item, ItemWriter.write(Chunk<? extends T>) is called once per chunk with all the items the step accumulated (up to the commit interval). Writing a whole chunk lets the writer batch its I/O — e.g. JdbcBatchItemWriter issues a single JDBC batch, JpaItemWriter persists then flushes once. This amortizes network round-trips and is the main performance lever in batch. The whole chunk is written inside one transaction; if write throws, the transaction rolls back and the chunk may be retried or skipped depending on configuration.

code

java · 15 lines
java
public interface ItemWriter<T> {
    // Called ONCE per chunk with all buffered items
    void write(Chunk<? extends T> chunk) throws Exception;
}

// A trivial custom writer that batches a log line
public class LoggingItemWriter implements ItemWriter<Trade> {
    @Override
    public void write(Chunk<? extends Trade> chunk) {
        // chunk.getItems() -> List<Trade>, written together
        for (Trade t : chunk) {
            log.info("writing {}", t.getId());
        }
    }
}

go deeper

for a junior

Should know write handles a batch of items and that this is for efficiency.

for a middle

Should connect commit interval to chunk size and mention batched I/O per writer type.

for a senior

Should discuss transaction boundary, retry/skip interaction, idempotency and tuning the commit interval.

for a principal

Should reason about throughput vs memory vs rollback cost trade-offs and exactly-once side-effect concerns on retry.

**Chunk-oriented processing.** A Spring Batch `Step` typically runs in a read-process-write loop. The framework calls `ItemReader.read()` repeatedly to pull single items, optionally passes each through an `ItemProcessor`, and buffers the results until it reaches the *commit interval* (the chunk size). Then it hands the buffered items to the `ItemWriter` **all at once** and commits the transaction. This grouping is the 'chunk'. **The interface.** In modern Spring Batch (5.x) the contract is: ```java public interface ItemWriter<T> { void write(Chunk<? extends T> chunk) throws Exception; } ``` `Chunk<T>` is a lightweight container introduced in Spring Batch 5 (before that the method took `List<? extends T> items`). It wraps the items plus metadata (e.g. items skipped mid-chunk). **Why a whole chunk?** The reason is *batched I/O*. Databases, files, and message brokers are far more efficient when you send many rows/lines in one round-trip than one-at-a-time. `JdbcBatchItemWriter` turns the chunk into a single JDBC `addBatch()`/`executeBatch()`. `FlatFileItemWriter` writes all lines then flushes/forces the buffer once. `JpaItemWriter` calls `persist`/`merge` for each entity then a single `flush()`. So the commit interval directly controls the batch size and is the primary throughput/memory trade-off knob. **Transactionality.** Each chunk is processed inside one transaction managed by the step. The sequence is roughly: read N items, process them, `write(chunk)`, then commit. If `write` throws, the transaction rolls back. Depending on your skip/retry policy the chunk may be retried (often item-by-item to isolate the offender) or the failing item skipped. This is why writers should be idempotent-friendly and why side effects belong inside `write`, not scattered around. **Ordering & guarantees.** Items are written in read order within the chunk. The writer must write *all* items it is given; partial writes that aren't rolled back can cause duplicates on retry. **When to use which writer.** Use `FlatFileItemWriter` for CSV/fixed-width files, `JdbcBatchItemWriter` for raw SQL inserts/updates, `JpaItemWriter`/`HibernateItemWriter` when you already have JPA entities, `CompositeItemWriter` to send each item to several writers, and `ClassifierCompositeItemWriter` to route different items to different writers. **Gotcha.** Because the writer sees the whole chunk, per-item logic that must run exactly once (auditing, external calls) needs care on retry. Also, a huge commit interval means large in-memory chunks and bigger rollback cost — tune it.

  • What controls how many items end up in a single chunk passed to write?
    The commit interval configured on the step (e.g. .chunk(100, txManager)). The step reads/processes up to that many items, then invokes write once with that many. A smaller reader stream at end-of-data yields a final smaller chunk.
  • What happens if write() throws an exception?
    The chunk transaction rolls back. If a retry/skip policy is configured, the step may retry the chunk (often scanning item-by-item to isolate the failure) or skip the offending item; otherwise the step fails.

saying these in an interview costs you the question

  • Thinking write is called once per item like read()
  • Believing each item is committed independently
  • Not knowing the commit interval equals the chunk/batch size
  • Confusing ItemWriter with ItemProcessor's role

context