skip to content

Even with restart support, why must a Spring Batch ItemWriter be idempotent, and how do you make one idempotent?

level: seniorimportance: should knowfreq 50%

answer

  1. chunk replay on restart = at-least-once
  2. business write + metadata not same tx = danger
  3. upsert / MERGE / ON CONFLICT
  4. assigned-id save = update not insert
  5. external sink -> idempotency key / dedupe table

basics

~20 s

On restart after a failure, Spring Batch re-reads and re-writes the chunk that was in flight when the job crashed, so a writer can see the same items twice. Making writes upserts (insert-or-update on a key) keeps re-runs safe.

solid answer

~40 s

Chunk-oriented steps commit one chunk per transaction. If the process dies after a chunk's business writes commit but before the batch metadata commits — or mid-chunk — a restart replays that chunk, giving the writer at-least-once semantics. So a plain `INSERT` writer can create duplicate rows or throw unique-constraint violations on restart. You make the writer idempotent by keying on a natural/business key and using an upsert: `INSERT ... ON CONFLICT DO UPDATE` (Postgres), SQL `MERGE`, or JPA `save()` on an entity with an assigned id. Alternatives are delete-by-key-then-insert within the chunk transaction, or dedupe checks. The goal is that reprocessing an item produces the same end state, never a duplicate side effect. For non-database sinks (emails, payments) you need an external idempotency key or a dedupe store, since you can't upsert those.

code

java · 19 lines
java
// Idempotent JdbcBatchItemWriter using Postgres UPSERT on a business key.
@Bean
JdbcBatchItemWriter<Order> orderWriter(DataSource ds) {
    return new JdbcBatchItemWriterBuilder<Order>()
            .dataSource(ds)
            .sql("""
                INSERT INTO orders (order_ref, customer_id, amount, status)
                VALUES (:orderRef, :customerId, :amount, :status)
                ON CONFLICT (order_ref)          -- natural/business key
                DO UPDATE SET customer_id = EXCLUDED.customer_id,
                              amount      = EXCLUDED.amount,
                              status      = EXCLUDED.status
                """)
            .beanMapped()
            .build();
}
// Re-writing the same Order on a restart just re-applies the same row -> no dup,
// no unique-constraint failure. Compare to a plain INSERT, which would either
// duplicate the row or blow up on the unique index during recovery.

go deeper

for a junior

May not realize restart can replay a chunk; learns the at-least-once idea.

for a middle

Knows to use upserts/MERGE for DB writers and why plain inserts break on restart.

for a senior

Reasons about the transaction boundary between business write and metadata, and handles external sinks with idempotency keys.

for a principal

Designs end-to-end idempotency: natural keys, atomic dedupe markers, restartable readers, and downstream contracts.

**Why duplicates happen even with restart.** A chunk-oriented Step processes items in chunks of size N inside a single transaction: read N, process N, write N, commit — and the same transaction updates the JobRepository metadata (`BATCH_STEP_EXECUTION` counts, the reader's `ExecutionContext`). Spring Batch aims for exactly-once *within* a clean run, but failures break that: - A crash **mid-chunk** (before commit) rolls back the whole chunk; on restart the reader resumes from the last committed position and re-reads those items — so they are written again from scratch (fine if nothing committed). - The dangerous window is when the **business write and the metadata are not in the same transaction** (e.g. writer targets a different datasource, a message queue, or a remote API). Then the business side can commit while the batch bookkeeping does not, and restart replays already-applied writes. This is **at-least-once** delivery. - Retry/skip logic can also re-invoke the writer with overlapping items after a rollback-and-retry of a chunk. Because of these, a robust writer must be **idempotent**: applying it twice for the same item yields the same final state and no duplicate external effect. **Techniques for a database writer.** 1. **Upsert on a business key.** Postgres `INSERT ... ON CONFLICT (natural_key) DO UPDATE`; SQL Server / Oracle `MERGE`; MySQL `INSERT ... ON DUPLICATE KEY UPDATE`. A `JdbcBatchItemWriter` can carry such SQL. Re-writing the same item just re-sets the same values. 2. **Assigned-id JPA save.** With `JpaItemWriter`/`save()` and an entity whose id is the business key (not DB-generated), the second write is an update (merge), not a new insert. 3. **Delete-then-insert by key** inside the chunk transaction: `DELETE WHERE key IN (...)` then `INSERT`. Idempotent because the delete removes any partial prior write. 4. **Guard with a dedupe check** (exists-by-key) — weaker under concurrency, but sometimes used. **Non-transactional / external sinks.** Emails, SMS, payment charges, and non-transactional queues can't be upserted. Use an **idempotency key** understood by the downstream system (e.g. Stripe's `Idempotency-Key` header), or maintain a local 'already-sent' table keyed by item id that you check-and-insert transactionally before performing the effect. The item's stable business id is the anchor. **Design guidance.** - Prefer a **natural key** you can derive from the item; a random generated key each run defeats deduplication. - Keep the business write and, where possible, the dedupe record in **one transaction** so they commit atomically. - Order matters for non-transactional resources: consider writing a 'processed' marker before or atomically with the effect. - Make the whole step **restartable** (restartable reader saving position in `ExecutionContext`), so restart replays the minimum. **Gotchas.** - Auto-generated primary keys make inserts non-idempotent by construction — a restart yields new keys and duplicate rows. - A unique constraint without upsert turns a duplicate into a hard failure on restart, blocking recovery. - `saveAll` with new entities in JPA still inserts duplicates unless the id is assigned/known.

  • If the ItemWriter targets the same datasource as the JobRepository and shares its transaction, do you still need idempotency?
    It reduces the risk because the business write and metadata commit atomically, so a mid-chunk crash rolls both back together. But retry/skip re-processing of a chunk and any non-shared resource can still replay writes, so keeping the writer idempotent is the safe default.
  • How do you make sending a confirmation email idempotent, since you can't upsert an email?
    Maintain a local 'sent' table keyed by the item's business id; in one transaction check-and-insert the marker, then send. On restart the marker already exists so you skip re-sending. Alternatively use a provider idempotency key so the downstream deduplicates.

saying these in an interview costs you the question

  • Assuming Spring Batch guarantees exactly-once so writers never see an item twice
  • Relying on auto-generated primary keys and expecting no duplicates on restart
  • Using a random per-run key that makes deduplication impossible

context