You must write each chunk to an external REST API that has no batch endpoint and is only at-least-once reliable. How do you design the ItemWriter for correctness under Spring Batch retries/restarts?
answer
- HTTP can't roll back => at-least-once
- Idempotency-Key per item dedupes replays
- retry re-scans chunk item-by-item
- crash after HTTP before commit => restart replays
- exactly-once => transactional outbox
basics
~20 sImplement a custom ItemWriter that loops the chunk calling the API, but make each call idempotent (send a client-generated idempotency key per item) so retries and restarts don't create duplicates. Track progress so restart resumes cleanly.
solid answer
~40 sA REST sink breaks Batch's transactional model: HTTP calls can't roll back, and Spring Batch may retry a chunk item-by-item on failure and restart a failed job later. So design for at-least-once delivery plus idempotency. Give each item a stable business/idempotency key and pass it (e.g. Idempotency-Key header) so the server dedupes replays. In write(chunk), iterate items; wrap transient failures (5xx/timeouts) in a retryable exception so the step's retry/skip policy works, and classify 4xx as non-retryable. Because each item send is idempotent, item-by-item retry re-scanning is safe. Remember the DB metadata commit and the HTTP effect aren't atomic: a crash after HTTP but before commit replays the chunk on restart — idempotency keys are what save you. If exactly-once is mandatory and the server can't dedupe, use a transactional outbox.
code
java · 21 linespublic class RestApiItemWriter implements ItemWriter<Order> {
private final RestClient rest;
public RestApiItemWriter(RestClient rest) { this.rest = rest; }
@Override
public void write(Chunk<? extends Order> chunk) {
for (Order o : chunk) {
try {
rest.post().uri("/orders")
// server dedupes replays with the same key
.header("Idempotency-Key", o.idempotencyKey())
.body(o)
.retrieve().toBodilessEntity();
} catch (HttpServerErrorException | ResourceAccessException e) {
// retryable: let the step RetryPolicy handle it
throw new RetryableWriteException(o.id(), e);
}
// 4xx -> non-retryable; classify as skip in the step config
}
}
}go deeper
Likely only knows to loop the chunk and call the API.
Should recognize retries can cause duplicates and mention idempotency at a high level.
Should design idempotency keys, retry/skip classification, and know HTTP can't roll back with the chunk.
Should articulate the non-atomic commit window, at-least-once vs exactly-once, and propose a transactional outbox when dedup isn't available.
**Why this is hard.** Spring Batch's correctness model assumes the writer participates in a transaction that commits with the chunk's `ExecutionContext`/`StepExecution` metadata. An external REST API is **not transactional**: once the HTTP call returns 200, the effect is committed remotely and cannot be rolled back if the surrounding Batch transaction later fails. Two Batch behaviors compound this: 1. **Retry within a chunk.** With a retry policy, a failed `write(chunk)` may be retried; the default fault-tolerant behavior can re-scan the chunk item-by-item, so items already sent successfully can be sent **again**. 2. **Restart across job runs.** If the JVM crashes after the HTTP calls but *before* the chunk transaction commits the Batch metadata, on restart Spring Batch replays that chunk (it never recorded it as done) — resending everything. Both mean **at-least-once** delivery is the realistic guarantee. The fix is **idempotency**, not trying to force rollback. **Design.** - **Idempotency key per item.** Derive a stable key from the item's business identity (or a UUID stored on the item before writing) and send it, e.g. an `Idempotency-Key` HTTP header. The server must dedupe: a replay with the same key returns the original result without creating a duplicate. This makes resends safe under both retry and restart. - **Custom writer.** ```java public class RestApiItemWriter implements ItemWriter<Order> { private final RestClient rest; public void write(Chunk<? extends Order> chunk) { for (Order o : chunk) { rest.post().uri("/orders") .header("Idempotency-Key", o.idempotencyKey()) .body(o).retrieve().toBodilessEntity(); } } } ``` - **Let the step manage retry/skip.** Throw a *retryable* exception (e.g. wrap 5xx/timeouts) so the step's `RetryPolicy` backs off and retries; classify 4xx as non-retryable/skip. Because each call is idempotent, item-by-item retry re-scanning is safe. - **Bounded batching without a batch endpoint.** You still get the chunk, so you can parallelize the per-item calls (bounded thread pool) to regain throughput, but keep idempotency. - **ItemStream / progress.** The writer generally stays stateless per chunk. If you need to record externally-confirmed offsets, implement `ItemStream` and persist a cursor in the `ExecutionContext` in `update()`; but the strong guarantee still comes from server-side dedup, not local state, because local state and the HTTP effect aren't atomic either. - **Exactly-once via outbox.** If duplicates are truly unacceptable and the server can't dedupe, use a transactional **outbox**: the writer inserts rows into an outbox table in the same DB transaction as the Batch metadata (so it's atomic and restart-safe), and a separate relay publishes to the REST API idempotently. This converts the problem into a local-transaction problem. **Gotchas.** - Assuming `transactional(true)` on a custom writer buys you rollback — it only buffers/rolls back local buffered writes, never remote HTTP effects. - Non-idempotent POST + Batch retry = silent duplicates; classic production bug. - Skip listeners firing side effects that also aren't idempotent. - Very large chunks amplify duplicate blast radius on a mid-chunk failure. **When to use what.** Server supports idempotency keys -> simple custom writer + keys. Server doesn't and duplicates are fatal -> outbox + relay. Duplicates tolerable -> at-least-once with keys is usually enough.
- The Batch metadata commit and the HTTP call aren't atomic. What exact failure window causes duplicates, and how do you neutralize it?If the process crashes after successful HTTP calls but before the chunk's transaction commits the StepExecution/ExecutionContext, restart replays the whole chunk and resends. You can't close the window locally (two systems, no shared transaction); you neutralize its effect with server-side idempotency keys, or eliminate it with a transactional outbox where the send record commits atomically with Batch metadata.
- How does the step's fault-tolerant retry interact with a non-idempotent writer here?On a chunk failure the fault-tolerant step can re-process the chunk item-by-item to isolate the bad item, re-invoking write for items that already succeeded. Without idempotency that produces duplicate POSTs. Idempotency keys (or tracking sent items) make the re-scan safe.
saying these in an interview costs you the question
- Believing transactional(true) rolls back the HTTP side effect
- Assuming exactly-once is achievable without idempotency keys or an outbox
- Ignoring that retry re-scans already-sent items
- Storing 'sent' state only in memory and expecting it to survive restart