If writing the 50th item of a 100-item chunk throws, what happens to the other 99 items in that chunk?
answer
- whole chunk rolls back, all-or-nothing
- 99 good items lost too
- default: step FAILED, job stops
- restart resumes at start of failed chunk
- non-transactional side effects NOT rolled back
basics
~10 sThe whole chunk's transaction rolls back, so none of the 100 items are committed. By default the step fails and stops. All 100 are undone together — the chunk is atomic.
solid answer
~40 sThe transaction wrapping the chunk rolls back, so all 100 items are discarded — nothing from that chunk is committed. Any database work done in the read/process/write phase of that chunk is undone. By default (no fault tolerance) the exception propagates, the step ends in FAILED state, and the job stops. On a restart, because chunks commit independently and the JobRepository tracked how many items were read/committed, the step resumes from the start of the failed chunk, not from item zero — already-committed earlier chunks are not reprocessed. Note this base behavior is all-or-nothing per chunk: without skip/retry configuration there is no partial success within a chunk. Whether the framework retries or skips items is a separate fault-tolerance concern layered on top of this boundary.
go deeper
Knows the whole chunk rolls back — nothing partial commits.
Explains default FAILED outcome and that good items in the chunk are also lost.
Adds restart-from-failed-chunk semantics via JobRepository counts committed in the same transaction.
Highlights non-transactional side-effect leakage and prescribes idempotent/deferred external calls.
## Chunk atomicity The defining property of the chunk transaction boundary is **atomicity per chunk**: either the whole chunk commits or the whole chunk rolls back. There is no partial commit of some items within a chunk in the base model. So if item 50 of a 100-item chunk fails in the writer (or the processor, or even the reader), the `PlatformTransactionManager` **rolls back** the transaction that was started for that chunk. Every database mutation made while building that chunk is undone. The 99 'good' items are lost too — they were never committed. ## Default (non-fault-tolerant) outcome With a plain chunk step (no `.faultTolerant()`, no skip/retry): 1. The exception propagates out of the chunk. 2. The step transitions to `BatchStatus.FAILED`. 3. The job stops; the exit status reflects the failure. ## Restart semantics Batch metadata is stored in the `JobRepository` **inside the same transaction** as the chunk write. So the committed read/write counts always match what is actually in the target. On restart of a failed job instance: - Already-committed chunks are **not** reprocessed. - The step resumes at the beginning of the chunk that failed. That is why the boundary matters for correctness: the commit count and the data can never drift, because they commit or roll back together. ## Important gotcha: non-transactional side effects Rollback only undoes work on **transactional resources** enrolled in that `PlatformTransactionManager` (typically the JDBC datasource). If your writer also sent an email, called a REST API, or wrote to a non-transactional message queue, those side effects are **not** rolled back. When the chunk retries or the job restarts, those side effects can happen again — a duplicate. This is the classic reason to make writers idempotent or defer external calls. ## Relationship to skip/retry (out of scope here) How the framework behaves when you *do* enable fault tolerance — retrying the chunk, scanning item-by-item, skipping the offending item — is a separate mechanism built on top of this boundary. The base guarantee we are describing is simply: one exception anywhere in the cycle ⇒ the whole chunk rolls back. ## Key classes - `PlatformTransactionManager.rollback(...)` — invoked by the framework on any unhandled exception in the cycle. - `StepExecution` / `JobRepository` — track committed counts, enabling correct restart. - `BatchStatus.FAILED` — the terminal state of the step after an uncaught chunk exception.
- After the failure, will restarting the job reprocess the chunks that already committed?No. The JobRepository committed those chunks' counts in the same transaction as the data, so on restart the step skips completed chunks and resumes at the start of the chunk that failed.
- Your writer sends a confirmation email per item. On rollback, are the emails un-sent?No. Emails are a non-transactional side effect not enrolled in the PlatformTransactionManager, so they are not rolled back. On retry/restart they can be sent again. Make external effects idempotent or move them out of the transactional writer (e.g., stage them and send after commit).
saying these in an interview costs you the question
- Claiming only the failing item is discarded and the other 99 still commit
- Assuming rollback reverses emails / REST calls / queue sends
- Thinking restart replays from item zero rather than the failed chunk