How do you keep non-transactional side effects (emails, REST calls, external queues) from being duplicated or corrupted when a chunk transaction rolls back?
answer
- rollback only undoes enrolled resources
- emails/REST/queue = not rolled back => duplicates on retry
- idempotency is the robust fix
- defer via afterCommit synchronization
- transactional outbox = atomic intent
basics
~20 sRollback only undoes transactional resources, not emails or REST calls. Make those side effects idempotent, or defer them until after commit, or stage them in the same transactional store and process later. Never assume rollback un-does external effects.
solid answer
~50 sThe chunk transaction only rolls back resources enrolled in its PlatformTransactionManager — typically the JDBC datasource. External effects (emails, REST posts, non-transactional queue publishes) done inside the writer are not undone, so on rollback-then-retry or job restart they can duplicate. Strategies: (1) Make external calls idempotent — keys/dedup on the receiver side, so replays are harmless. (2) Defer effects until after commit — register a TransactionSynchronization (afterCommit) or an ItemWriteListener.afterWrite so the call fires only once the chunk is durable. (3) Stage-and-forward — the transactional writer inserts an outbox row in the same DB transaction, and a later step or poller performs the external call, giving transactional atomicity for the intent. Also consider the transactional-resource readers/writers Batch provides (e.g., a transactional file writer that buffers until commit) so file output participates correctly with the chunk boundary.
code
java · 20 lines// Outbox: stage the intent in the SAME transaction as business data.
@Component
class OrderApiWriter implements ItemWriter<Order> {
private final JdbcTemplate jdbc;
OrderApiWriter(JdbcTemplate jdbc) { this.jdbc = jdbc; }
@Override
public void write(Chunk<? extends Order> chunk) {
for (Order o : chunk) {
jdbc.update("insert into orders(id, total) values (?,?)",
o.id(), o.total());
// enqueue the external call transactionally -- commits/rolls back
// together with the order row; a relay sends it later, idempotently.
jdbc.update("insert into outbox(aggregate_id, type, payload) values (?,?,?)",
o.id(), "ORDER_CREATED", o.toJson());
}
// No REST call here: if this chunk rolls back, the outbox row rolls
// back too, so nothing is ever sent for uncommitted orders.
}
}go deeper
Aware that non-DB effects aren't rolled back.
Suggests idempotency and not calling external systems inside the writer.
Uses afterCommit synchronization / listeners and knows retries cause replays.
Designs a transactional outbox for atomic intent and reasons about exactly-effectively-once across the boundary and crash windows.
## The root problem The chunk boundary's rollback is scoped to resources **enrolled in the step's `PlatformTransactionManager`** — in practice the JDBC datasource (via `DataSourceTransactionManager`) or the JPA `EntityManager` (via `JpaTransactionManager`). Anything outside that enrollment is invisible to rollback: - Sending an email (SMTP). - Calling a REST API. - Publishing to a non-transactional queue (a plain Kafka/HTTP publish). - Writing to a file with a naive writer that flushes immediately. If the writer performs one of these and then the chunk rolls back (an exception later in the same chunk, or a retry, or a full job restart), the external effect **already happened** and is not reversed. The result is duplicated emails, double REST posts, phantom messages. ## Why retries/restarts amplify it Because chunks commit independently and the framework can retry a failed chunk or restart a failed job from the failed chunk, the read-process-write cycle for that chunk may run **more than once**. Every extra run repeats the non-transactional side effect. So this is not a rare edge — any fault-tolerant or restartable job invites duplicates. ## Mitigation strategies (principal-level toolkit) ### 1. Idempotency (make replay harmless) Design the external operation so doing it twice equals doing it once: send a natural/business key the receiver dedups on, use conditional writes, or an idempotency key header on REST. This is the most robust because it tolerates any replay cause. ### 2. Defer until after commit Only perform the effect once the chunk is durably committed: - Register a `TransactionSynchronization` and act in `afterCommit()`. - Or use an `ItemWriteListener.afterWrite(...)` / `ChunkListener.afterChunk(...)` — but note these still fire in proximity to the boundary; `afterCommit` synchronization is the precise hook for 'only if it actually committed'. This prevents the effect when the chunk rolls back, but a crash *between* commit and the deferred call can still lose it — so pair with idempotency or retries. ### 3. Transactional outbox (stage-and-forward) The writer, in the **same** DB transaction as the business data, inserts a row into an 'outbox' table describing the intended external call. Because it is the same transaction, the outbox row and the business data commit or roll back together — atomic *intent*. A separate step, poller, or messaging relay later reads the outbox and performs the actual external call (with retries + idempotency). This is the gold standard for exactly-effectively-once semantics across a non-transactional boundary. ### 4. Use transactional-resource writers For file output, Spring Batch's `FlatFileItemWriter` (and similar) buffer writes and flush on commit when configured transactional, so a rolled-back chunk does not leave partial lines. Prefer these over hand-rolled immediate-flush writers so file output honors the chunk boundary. ## Anti-patterns - Doing the REST/email call directly in the writer and assuming chunk atomicity covers it. - Turning off restart/retry to 'avoid duplicates' — that trades correctness for a much weaker system. - Assuming `ResourcelessTransactionManager` adds any protection — it is a no-op. ## Key classes / hooks - `TransactionSynchronizationManager.registerSynchronization(...)` + `TransactionSynchronization.afterCommit()`. - `ItemWriteListener`, `ChunkListener` — lifecycle hooks around the boundary. - `FlatFileItemWriter` (transactional file buffering) — resource that respects rollback. - Outbox pattern (application-level) — atomic intent inside the DB transaction.
- Why is an afterCommit TransactionSynchronization not, by itself, exactly-once?The commit succeeds, then the JVM can crash before the afterCommit callback runs, losing the effect. It prevents effects on rollback but not loss after commit. Combine it with a durable outbox and idempotent delivery for effectively-once.
- Where does the transactional outbox actually perform the external call?Not in the writer. A separate poller, message-relay, or a second batch step reads committed outbox rows and performs the call with retries and idempotency keys, marking rows done. The writer only records intent atomically with the data.
- Does making the writer @Transactional with a new propagation help?No — Spring Batch already owns the chunk transaction; nesting won't enroll SMTP/REST into it because those aren't transactional resources. The boundary problem is about resource enrollment, not propagation.
saying these in an interview costs you the question
- Believing chunk rollback undoes emails/REST/queue publishes
- Disabling retry/restart to dodge duplicates instead of designing idempotency
- Thinking @Transactional on the writer enrolls a REST client into the transaction