Even with restart support, why must a Spring Batch ItemWriter be idempotent, and how do you make one idempotent?
answer
- chunk replay on restart = at-least-once
- business write + metadata not same tx = danger
- upsert / MERGE / ON CONFLICT
- assigned-id save = update not insert
- external sink -> idempotency key / dedupe table
basics
~20 sOn restart after a failure, Spring Batch re-reads and re-writes the chunk that was in flight when the job crashed, so a writer can see the same items twice. Making writes upserts (insert-or-update on a key) keeps re-runs safe.
solid answer
~40 sChunk-oriented steps commit one chunk per transaction. If the process dies after a chunk's business writes commit but before the batch metadata commits — or mid-chunk — a restart replays that chunk, giving the writer at-least-once semantics. So a plain `INSERT` writer can create duplicate rows or throw unique-constraint violations on restart. You make the writer idempotent by keying on a natural/business key and using an upsert: `INSERT ... ON CONFLICT DO UPDATE` (Postgres), SQL `MERGE`, or JPA `save()` on an entity with an assigned id. Alternatives are delete-by-key-then-insert within the chunk transaction, or dedupe checks. The goal is that reprocessing an item produces the same end state, never a duplicate side effect. For non-database sinks (emails, payments) you need an external idempotency key or a dedupe store, since you can't upsert those.
code
java · 19 lines// Idempotent JdbcBatchItemWriter using Postgres UPSERT on a business key.
@Bean
JdbcBatchItemWriter<Order> orderWriter(DataSource ds) {
return new JdbcBatchItemWriterBuilder<Order>()
.dataSource(ds)
.sql("""
INSERT INTO orders (order_ref, customer_id, amount, status)
VALUES (:orderRef, :customerId, :amount, :status)
ON CONFLICT (order_ref) -- natural/business key
DO UPDATE SET customer_id = EXCLUDED.customer_id,
amount = EXCLUDED.amount,
status = EXCLUDED.status
""")
.beanMapped()
.build();
}
// Re-writing the same Order on a restart just re-applies the same row -> no dup,
// no unique-constraint failure. Compare to a plain INSERT, which would either
// duplicate the row or blow up on the unique index during recovery.go deeper
May not realize restart can replay a chunk; learns the at-least-once idea.
Knows to use upserts/MERGE for DB writers and why plain inserts break on restart.
Reasons about the transaction boundary between business write and metadata, and handles external sinks with idempotency keys.
Designs end-to-end idempotency: natural keys, atomic dedupe markers, restartable readers, and downstream contracts.
**Why duplicates happen even with restart.** A chunk-oriented Step processes items in chunks of size N inside a single transaction: read N, process N, write N, commit — and the same transaction updates the JobRepository metadata (`BATCH_STEP_EXECUTION` counts, the reader's `ExecutionContext`). Spring Batch aims for exactly-once *within* a clean run, but failures break that: - A crash **mid-chunk** (before commit) rolls back the whole chunk; on restart the reader resumes from the last committed position and re-reads those items — so they are written again from scratch (fine if nothing committed). - The dangerous window is when the **business write and the metadata are not in the same transaction** (e.g. writer targets a different datasource, a message queue, or a remote API). Then the business side can commit while the batch bookkeeping does not, and restart replays already-applied writes. This is **at-least-once** delivery. - Retry/skip logic can also re-invoke the writer with overlapping items after a rollback-and-retry of a chunk. Because of these, a robust writer must be **idempotent**: applying it twice for the same item yields the same final state and no duplicate external effect. **Techniques for a database writer.** 1. **Upsert on a business key.** Postgres `INSERT ... ON CONFLICT (natural_key) DO UPDATE`; SQL Server / Oracle `MERGE`; MySQL `INSERT ... ON DUPLICATE KEY UPDATE`. A `JdbcBatchItemWriter` can carry such SQL. Re-writing the same item just re-sets the same values. 2. **Assigned-id JPA save.** With `JpaItemWriter`/`save()` and an entity whose id is the business key (not DB-generated), the second write is an update (merge), not a new insert. 3. **Delete-then-insert by key** inside the chunk transaction: `DELETE WHERE key IN (...)` then `INSERT`. Idempotent because the delete removes any partial prior write. 4. **Guard with a dedupe check** (exists-by-key) — weaker under concurrency, but sometimes used. **Non-transactional / external sinks.** Emails, SMS, payment charges, and non-transactional queues can't be upserted. Use an **idempotency key** understood by the downstream system (e.g. Stripe's `Idempotency-Key` header), or maintain a local 'already-sent' table keyed by item id that you check-and-insert transactionally before performing the effect. The item's stable business id is the anchor. **Design guidance.** - Prefer a **natural key** you can derive from the item; a random generated key each run defeats deduplication. - Keep the business write and, where possible, the dedupe record in **one transaction** so they commit atomically. - Order matters for non-transactional resources: consider writing a 'processed' marker before or atomically with the effect. - Make the whole step **restartable** (restartable reader saving position in `ExecutionContext`), so restart replays the minimum. **Gotchas.** - Auto-generated primary keys make inserts non-idempotent by construction — a restart yields new keys and duplicate rows. - A unique constraint without upsert turns a duplicate into a hard failure on restart, blocking recovery. - `saveAll` with new entities in JPA still inserts duplicates unless the id is assigned/known.
- If the ItemWriter targets the same datasource as the JobRepository and shares its transaction, do you still need idempotency?It reduces the risk because the business write and metadata commit atomically, so a mid-chunk crash rolls both back together. But retry/skip re-processing of a chunk and any non-shared resource can still replay writes, so keeping the writer idempotent is the safe default.
- How do you make sending a confirmation email idempotent, since you can't upsert an email?Maintain a local 'sent' table keyed by the item's business id; in one transaction check-and-insert the marker, then send. On restart the marker already exists so you skip re-sending. Alternatively use a provider idempotency key so the downstream deduplicates.
saying these in an interview costs you the question
- Assuming Spring Batch guarantees exactly-once so writers never see an item twice
- Relying on auto-generated primary keys and expecting no duplicates on restart
- Using a random per-run key that makes deduplication impossible