What are the restartability, ordering, and fault-tolerance limitations of a multithreaded step, and when should you avoid it?
answer
- out-of-order commit -> no ordering guarantee
- single read-count assumes contiguous completion -> holes break restart
- saveState(false) + idempotent / restart-from-scratch
- skip/retry + shared reader = fragile
- prefer partitioning for order/restart/scale-out
basics
~20 sChunks run concurrently, so items are processed out of order and commit out of order. On failure the single read-count can't cleanly mark a resume point, so restart may skip or reprocess items. Avoid it when order matters, restart must resume exactly, or you rely heavily on skip/retry.
solid answer
~40 sA multithreaded step trades correctness guarantees for throughput. Ordering: chunks are dispatched to a pool and commit independently, so output order is non-deterministic — never rely on input order downstream. Restartability: Spring Batch records progress as a single read-count in BATCH_STEP_EXECUTION, which assumes sequential completion; with concurrency, chunk 6 can commit while chunk 5 fails, leaving a gap the read-count can't represent, so a restart may reprocess or skip. Teams commonly set saveState(false) and treat the job as restart-from-scratch / idempotent. Fault tolerance: mixing skip/retry with a shared stateful reader amplifies races and makes error accounting fuzzy. Avoid multithreaded steps when output ordering matters, exact mid-step resume is required, the step is read-bound (synchronized reads dominate), or you need robust skip/retry — prefer partitioning, which isolates state per worker and restarts cleanly.
code
java · 11 lines// Design a multithreaded step for restart-from-scratch, idempotent output:
JdbcBatchItemWriter<Record> upsertWriter = new JdbcBatchItemWriterBuilder<Record>()
.dataSource(ds)
// idempotent: re-running the whole job won't duplicate rows
.sql("INSERT INTO target(id, val) VALUES (:id, :val) " +
"ON CONFLICT (id) DO UPDATE SET val = EXCLUDED.val")
.beanMapped()
.build();
// reader.setSaveState(false) so Batch doesn't try to resume a non-contiguous position
// step.taskExecutor(pool) + SynchronizedItemStreamReader as usualgo deeper
Know outputs come out of order and restart is not exact.
Explain saveState(false) + idempotent writes as the mitigation.
Articulate why the single read-count can't represent non-contiguous completion, and the skip/retry fragility.
Choose between multithreaded step and partitioning on ordering/restart/scale-out grounds; account for DB connection pressure and downstream limits.
## Why guarantees weaken under concurrency A single-threaded chunk step has a clean invariant: it reads → writes → commits chunk 1, then chunk 2, then chunk 3, strictly in order. Spring Batch leans on that to persist progress. A **multithreaded step** breaks the invariant by processing many chunks concurrently, which undermines three things. ### 1. Ordering Chunks are handed to a `TaskExecutor` and complete whenever their thread finishes. Commits happen per-chunk, per-thread, in **non-deterministic order**. So: - Output rows are not written in input order. - Any downstream step or consumer that assumes ordering (e.g. 'the last row wins,' sequence-dependent aggregation) is unsafe. - Even within tolerance, you cannot assume item i is committed before item i+1. ### 2. Restartability Spring Batch tracks step progress primarily via a **read count** (and write/commit counts) stored in the `BATCH_STEP_EXECUTION` table, and via the reader's `ExecutionContext` (`ItemStream.update`). This model implicitly assumes **contiguous, in-order completion** — 'we got through the first N items.' With concurrency, completion is **non-contiguous**: chunk 6 may commit successfully while chunk 5's thread fails. There is no single integer N that says 'everything before N is done and nothing after is,' because there are holes. Consequences: - On restart, the persisted count doesn't correspond to a safe resume boundary, so you risk **reprocessing** already-written items or **skipping** un-written ones. - Standard advice: set **`saveState(false)`** on the reader and design the job to **restart from the beginning** with **idempotent** writes (upserts, dedupe keys), or scope the input so a full re-run is acceptable. ### 3. Fault tolerance (skip / retry) Spring Batch's `faultTolerant()` skip/retry logic tracks skipped/retried items and sometimes **re-reads** a chunk. Combined with a **shared stateful reader** across threads, this compounds the race hazards and makes the skip/retry accounting harder to reason about. It's not strictly forbidden, but it's fragile; if you need serious fault tolerance, isolate state. ## Other limits - **Read-bound steps don't benefit.** Because the reader must be synchronized (`SynchronizedItemStreamReader`), reads serialize. If reading is the bottleneck, threads just queue on the read lock. - **Single JVM only.** It scales up (more threads on one machine), not out (more machines). Downstream resources (DB connection pool, target API) can become the new bottleneck or get overwhelmed. - **Transaction/resource pressure.** N concurrent chunks means N concurrent transactions and N DB connections — size the connection pool accordingly or you'll see contention/timeouts. ## When to avoid it / what to use instead Avoid a multithreaded step when: output ordering matters; exact mid-step restart/resume is required; the step is read-bound; or you depend on robust skip/retry. In those cases prefer **partitioning** (`Partitioner` + `PartitionHandler`, local `TaskExecutorPartitionHandler` or remote): each partition is an independent step execution with its **own reader** over a disjoint slice of data, so there's no shared mutable state, restart works **per partition**, and you can even distribute across JVMs. Use a multithreaded step for the easy, low-config win on **I/O-bound, order-insensitive, idempotent** work in a single JVM.
- Why can't Spring Batch cleanly resume a failed multithreaded step from where it stopped?Progress is tracked as a single read/commit count that assumes contiguous, in-order completion. Concurrency creates non-contiguous holes (a later chunk committed, an earlier one failed), so no single count marks a safe resume boundary.
- How do you make a multithreaded step's failures safe to re-run?Set saveState(false), treat it as restart-from-scratch, and make writes idempotent (upserts / dedupe keys) so reprocessing already-written items has no side effects.
saying these in an interview costs you the question
- Claiming a multithreaded step preserves input/output ordering
- Assuming restart resumes exactly where it failed like a single-threaded step
- Combining skip/retry with a shared stateful reader without caution
- Believing it scales across multiple JVMs