How does Spring Batch handle a skip differently for the read, process, and write phases of a chunk?
answer
- read skip = no rollback, just next()
- process/write skip = rollback + replay
- write replay = single-item chunks to isolate
- replay => processor/writer must be idempotent
- onSkipInRead has no item; process/write have the item
basics
~20 sA read skip just drops the item and reads the next, with no rollback. A process or write skip forces the chunk transaction to roll back and re-run items one at a time to find and skip the single bad item, then commit the rest.
solid answer
~40 sIn a chunk-oriented step, reading is transactional-independent from writing. When a skippable exception occurs during read, the framework simply discards that item and calls read() again — no transaction rollback, cheap. During process or write it's different: the writer receives the whole chunk, so Spring Batch can't know which item threw. It rolls back the chunk transaction, then re-executes the chunk with each item in its own single-item chunk to isolate the culprit; the offending item is skipped and the surviving items are committed. This replay is why write skips are costly and why processors/writers should be idempotent — surviving items may be processed twice across the rollback/replay. onSkipInRead, onSkipInProcess, and onSkipInWrite callbacks fire for the respective phase, and read/process/write skip counts are tracked separately on the StepExecution.
code
java · 14 lines@Bean
public Step step(JobRepository repo, PlatformTransactionManager tx,
ItemReader<Rec> reader, ItemProcessor<Rec, Rec> proc,
ItemWriter<Rec> writer) {
return new StepBuilder("step", repo)
.<Rec, Rec>chunk(50, tx)
.reader(reader).processor(proc).writer(writer)
.faultTolerant()
.skip(ParseException.class) // read-phase failure
.skip(ValidationException.class) // process-phase failure
.skip(DataAccessException.class) // write-phase failure -> triggers single-item replay
.skipLimit(20)
.build();
}go deeper
Know the three phases exist and that skip counts are tracked separately.
Explain read = cheap/no rollback vs process/write = rollback and single-item replay.
Tie the replay to idempotency requirements and throughput degradation; mention buffered vs stateful writers.
Weigh chunk size, skip policy, and writer design together to bound worst-case single-item replay cost under bad data.
## Chunk mechanics recap A chunk-oriented step loops, all inside one transaction that commits per chunk: 1. repeatedly call `ItemReader.read()` until the chunk size is reached (buffering items in memory), 2. pass each through the `ItemProcessor`, 3. then hand the whole list to `ItemWriter.write()`. Skip behavior differs by phase because of *where in this loop* the exception happens and *how much state must be unwound*. ## Read skip `read()` is called item-by-item and its result isn't yet part of the write transaction's committed batch. When a skippable exception is thrown during read, Spring Batch: 1. discards the item, 2. invokes `SkipListener.onSkipInRead(Throwable)`, 3. increments the read skip count, 4. and calls `read()` again for the next item. **No rollback** occurs — this is the cheap, clean case. Typical trigger: `FlatFileParseException` on a malformed line. ## Process skip The processor runs per item, but the failure is discovered while building the chunk that will be written transactionally. To skip a process failure Spring Batch must **roll back** and re-drive the chunk so it can attribute the failure to one item. The bad item is skipped (`onSkipInProcess(item, throwable)`), the rest proceed. ## Write skip `write(chunk)` receives a *list*. If it throws, the framework cannot tell which element failed. So it **rolls back the chunk transaction** and then re-processes the chunk with an effective chunk size of 1 — each item in its own transaction — to isolate the failing item. That one item is skipped and `onSkipInWrite(item, throwable)` fires; the others are committed individually. This is the most expensive path: a chunk of 100 that has one bad write can degrade to ~100 single-item transactions for that chunk. ## Replay gotchas - **Idempotency gotcha.** Because process/write skips replay items after a rollback, an item can pass through the processor (and sometimes side-effecting code) more than once. If your processor or writer has side effects (sending emails, calling external APIs, incrementing counters), design them to be idempotent or defer side effects, or you'll get duplicates on the replay. - **Ordering/stateful writer gotcha.** A writer that batches into a single SQL statement or keeps internal state can misbehave when replayed one item at a time. Buffered writers (e.g. `FlatFileItemWriter`) are handled by the framework's transactional buffering, but custom writers must tolerate the single-item replay. ## Counts and listeners `StepExecution` exposes: - `getReadSkipCount()`, - `getProcessSkipCount()`, - `getWriteSkipCount()`, - and total `getSkipCount()`. The three `SkipListener` callbacks let you log or persist the rejected items (e.g. to a dead-letter table) per phase. Note the read callback only gets the `Throwable` (no item, since the item never fully materialized), while process/write callbacks also receive the offending item. ## When it matters in interviews Candidates who say "a skip is a skip" miss the performance cliff of write skips and the idempotency requirement — those are the real production pitfalls.
- Why must an ItemProcessor be idempotent when write skips are enabled?After a write failure the chunk is rolled back and replayed one item at a time, so items that already went through the processor get processed again. Non-idempotent side effects would run twice.
- Why does onSkipInRead not receive the failed item?The exception occurs inside read() before a complete item object is produced, so there is no item to hand back — only the Throwable.
saying these in an interview costs you the question
- Claiming a write skip only discards one item with no rollback
- Assuming processors are called exactly once per item under fault tolerance
- Thinking read, process and write skips share one counter only (they have separate counts)
- Believing the framework knows which list element failed a batch write without replay