skip to content

Walk through the exact read/process/write sequence Spring Batch executes for one chunk, and where the transaction boundary sits.

level: middleimportance: must knowfreq 65%

answer

  1. ChunkProvider reads, ChunkProcessor processes+writes
  2. tx opens before first read, commits after write
  3. JobRepository checkpoint per chunk = restart point
  4. readCount/writeCount/filterCount/commitCount
  5. empty first read = no writer call

basics

~20 s

A transaction starts, then the reader is called up to N times, the processor runs on each read item, the collected results are written in one write() call, and the transaction commits. Then the next chunk starts a new transaction.

solid answer

~40 s

Each chunk runs inside its own transaction opened by the step's transaction manager. Spring Batch's ChunkOrientedTasklet drives it: the ChunkProvider loops calling ItemReader.read() until it has N items or read() returns null; then the ChunkProcessor calls ItemProcessor.process() on each item, dropping any that return null (filtered), and finally calls ItemWriter.write() once with the surviving list. After the write returns, the transaction commits and the JobRepository is updated with the new read/write/commit counts so the step can restart from that point. Reads are per-item, processing is per-item, writing and committing are per-chunk. If anything throws before commit, the whole chunk's transaction rolls back — nothing from that chunk is persisted (barring skip/retry configuration).

go deeper

for a junior

Can state read-many then write-once, roughly.

for a middle

Must give the ordered read/process/write/commit sequence and place the transaction boundary correctly.

for a senior

Adds the ChunkProvider/ChunkProcessor split, JobRepository checkpointing, and count semantics.

for a principal

Connects per-chunk checkpointing to restartability and reasons about replay behavior under fault tolerance.

**Who drives it.** A chunk step is implemented internally as a `ChunkOrientedTasklet`, whose `execute` is invoked repeatedly inside a Spring `TransactionTemplate` managed by the step. Each invocation processes exactly one chunk. Two collaborators split the work: a **ChunkProvider** (does the reading) and a **ChunkProcessor** (does the processing + writing). **Step-by-step for commit-interval = N:** 1. **Transaction begins.** The step's `PlatformTransactionManager` opens a transaction for this chunk. 2. **Read phase (ChunkProvider).** In a loop, call `ItemReader.read()`. Each non-null item is buffered. The loop stops when either N items have been read *or* the **CompletionPolicy** says the chunk is complete *or* `read()` returns `null` (end of data). (The default policy is `SimpleCompletionPolicy(N)`.) 3. **Process phase (ChunkProcessor).** For each buffered input item, call `ItemProcessor.process(item)`. If it returns `null`, the item is *filtered* (dropped) and counted in `filterCount`. Surviving outputs are collected into the output chunk. 4. **Write phase.** Call `ItemWriter.write(chunk)` **once** with the whole surviving list. 5. **Commit.** The transaction commits. Spring Batch then updates the `JobRepository`: `readCount`, `writeCount`, `filterCount`, `commitCount` in `BATCH_STEP_EXECUTION`. This checkpoint is what enables **restart** — a restarted job resumes after the last committed chunk. 6. **Repeat** until `read()` returns `null`, i.e. input is exhausted. The final chunk may be partial (fewer than N items). **Where the transaction boundary is.** Exactly one transaction wraps one chunk: it opens *before* the first read of the chunk and commits *after* the write. Reads, processing, and the write of a given chunk all happen inside that one transaction. (Detailed commit/rollback semantics are covered by the transaction-boundary leaf; here the key fact is simply *one transaction per chunk, boundary around read+process+write*.) **Counts / statistics.** `readCount` increments per successful read; `writeCount` per item actually written; `filterCount` per processor-null; `commitCount` per chunk committed. These are visible per step execution and are the basis for monitoring throughput. **Gotchas.** - **Reader state must survive across chunks** — the same `ItemReader` instance is reused, so cursor/position is preserved between chunk transactions. Stateful readers should be `@StepScope` for restart correctness. - **Don't do per-item DB writes in the processor.** The processor is per-item and inside the transaction, but batching belongs in the writer. Writing in the processor defeats the chunk model. - **Empty chunk short-circuits.** If the very first read returns null, the chunk is empty, the writer is not called, and the step completes. - **Reader exceptions vs writer exceptions differ on replay.** On a rollback with fault tolerance, Spring Batch may re-read or replay items one-at-a-time to isolate the bad one — but that's skip/retry territory. **When this matters in interviews.** Being able to say precisely 'read is per-item, process is per-item, write is per-chunk, commit is per-chunk, all inside one transaction' distinguishes someone who has run batch jobs from someone who has only read the docs.

  • Why does Spring Batch update the JobRepository after each chunk commit?
    To checkpoint progress. On restart, the job resumes from the last committed chunk instead of reprocessing everything, using the persisted read/write/commit counts in BATCH_STEP_EXECUTION.
  • If the ItemWriter throws on chunk 3, what happens to the items in chunk 3?
    The chunk-3 transaction rolls back, so none of chunk 3's writes persist. Chunks 1 and 2 already committed and stay. Without skip/retry, the step fails; with fault tolerance it may replay items individually.

saying these in an interview costs you the question

  • Thinking the writer is called once per item
  • Believing all chunks share one big transaction
  • Assuming filtered (processor-null) items don't count toward the commit-interval
  • Doing the actual persistence inside the processor instead of the writer

context