Given Spring Batch re-processes a rolled-back chunk item-by-item, what does that imply for an ItemProcessor that has side effects, and how do you handle it?
answer
- Rollback replays items -> processor runs 2+ times
- Transaction only undoes the DB, not REST/email/Kafka
- Processor = pure transform, no side effects
- Effects -> idempotent writer (upsert/dedup key)
- retry() amplifies reprocessing
basics
~20 sBecause a rolled-back chunk is replayed one item at a time, the same item can pass through the ItemProcessor more than once. Any side effect there (charging a card, sending email, calling an external API) would happen repeatedly. Keep the processor idempotent or pure and push non-idempotent effects to an idempotent writer.
solid answer
~40 sFault tolerance recovers a failed chunk by rolling back and re-processing the cached items individually. The transaction only protects **transactional resources** (the DB write). It does NOT undo external, non-transactional side effects a processor performed before the failure — a REST call, an email, a message published. Since the item is re-processed on the rescan, those effects fire again, causing duplicates. Mitigations: make the `ItemProcessor` a pure transformation with no external calls; move side effects into the `ItemWriter` where they're applied once per successful commit and can be made idempotent (upserts, dedup keys, natural idempotency tokens); use `processorNonTransactional()` awareness and design external calls to be replay-safe. Also account for retry: with `.retry(...)`, re-attempts amplify the same replay concern. The rule of thumb: processors transform, writers commit, and anything non-transactional must be idempotent.
code
java · 18 lines// ANTI-PATTERN: side effect in the processor.
// On a chunk rollback + rescan this email can be sent twice.
class BadProcessor implements ItemProcessor<Order, Order> {
private final EmailClient email;
BadProcessor(EmailClient email) { this.email = email; }
public Order process(Order o) {
email.sendConfirmation(o); // non-transactional, replays!
return o.markProcessed();
}
}
// BETTER: processor is pure; the writer performs an idempotent upsert
// and side effects are deferred to an idempotent, replay-safe step.
class GoodProcessor implements ItemProcessor<Order, Order> {
public Order process(Order o) {
return o.markProcessed(); // pure transform, no I/O side effect
}
}go deeper
Aware that an item may be processed more than once after a failure.
Knows to keep processors pure and move effects to writers.
Explains transactional vs non-transactional resources and designs idempotent writers/outbox.
Sets team conventions (processors transform, writers idempotent, outbox for external effects) and evaluates transactional messaging trade-offs.
## Why an item can be processed more than once When a fault-tolerant chunk rolls back, Batch replays the **buffered** items one at a time to find the poison item. That replay re-invokes the `ItemProcessor` for items it had already processed in the failed bulk attempt. So for a chunk of N where item k is bad, items 1..k-1 (and k itself) may be processed **twice** — once in the failed bulk pass, once in the scan. With `.retry(...)`, an item can be processed even more times across retry attempts. ## Transactions only cover transactional resources The chunk transaction is managed by a `PlatformTransactionManager` bound to a **transactional resource** — typically the target database. A rollback undoes writes to *that* resource. It does **not** undo: - an HTTP/REST call the processor made, - an email or SMS sent, - a message published to Kafka/RabbitMQ (unless you've wired transactional messaging), - a mutation to an in-memory cache or third-party system. Those are **non-transactional side effects**. On replay they simply happen again → duplicates, double-charges, duplicate notifications. ## The design fix 1. **Keep the `ItemProcessor` pure**: input item → transformed output, no I/O with side effects. This is the single most important rule and why processors are conceptually 'transform' steps. 2. **Move effects to the `ItemWriter`**, which is invoked in the write phase closest to commit, and make the writer **idempotent**: upserts / `INSERT ... ON CONFLICT`, dedup on a natural key, or an idempotency token so a replayed write is a no-op. 3. **For unavoidable external calls**, use an idempotency key the downstream honors, or defer the call to a later idempotent step / outbox pattern rather than doing it mid-chunk. 4. **Be aware of retries**: `.retry()` multiplies attempts; the same purity/idempotency discipline covers it. ## Related knobs - `processorNonTransactional()` on the fault-tolerant builder controls whether processed results are **cached** across a rollback or re-run through the processor. By default a fault-tolerant step re-processes items on retry/scan; `processorNonTransactional()` changes caching behavior for processors whose output is expensive but this doesn't make external side effects safe — it's about reprocessing cost, not idempotency of side effects. ## Gotcha summary - 'The transaction will roll it back' is **false** for anything non-transactional. - Assume every item may be processed 2+ times under fault tolerance. - Idempotency is a *correctness* requirement, not an optimization, once skip/retry is enabled.
- Doesn't the transaction rollback undo whatever the processor did?Only for transactional resources bound to the transaction manager (usually the DB write). Non-transactional side effects — REST calls, emails, published messages — are not undone, and they re-fire when the chunk is replayed.
- Where should a non-idempotent external call ideally live in a fault-tolerant chunk step?Not in the processor. Prefer an idempotent ItemWriter (upsert/dedup/idempotency key) or an outbox + separate idempotent step, so replays don't duplicate the effect.
saying these in an interview costs you the question
- Assuming the chunk transaction rolls back external side effects like emails or REST calls
- Putting non-idempotent side effects in the ItemProcessor
- Believing an item is processed at most once under fault tolerance