Why are most Spring Batch ItemReaders unsafe in a multithreaded step, and how do you fix it?
answer
- one reader shared by all threads
- cursor / line / page counter = mutable state
- ResultSet not thread-safe
- SynchronizedItemStreamReader wraps + synchronizes read()
- reads serialize -> gain is in process/write; restart weakened
basics
~20 sThere is one reader instance shared by all threads, and most readers hold mutable state (a cursor, current line, or page index). Concurrent read() calls corrupt that state. Wrap the reader in SynchronizedItemStreamReader so read() is synchronized.
solid answer
~40 sA multithreaded step shares a single ItemReader across all worker threads. Nearly all built-in readers are stateful: JdbcCursorItemReader advances a JDBC ResultSet cursor, FlatFileItemReader tracks the current line number and a shared BufferedReader, and paging readers like JpaPagingItemReader/JdbcPagingItemReader mutate a page counter and results buffer. None of these synchronize read(), so concurrent calls interleave and corrupt the state — you get duplicated rows, skipped rows, or exceptions like ArrayIndexOutOfBoundsException. The fix is SynchronizedItemStreamReader (or SynchronizedItemStreamWriter for writers), which wraps a delegate and makes read() synchronized while delegating open/update/close. This serializes reading — so the parallelism gain comes from concurrent processing and writing, not reading. Paging readers are safer than cursor readers but still need synchronization for their counters. It also weakens restart, since the shared read-count no longer maps cleanly to completed work.
code
java · 15 lines@Bean
@StepScope
public SynchronizedItemStreamReader<Customer> reader(DataSource ds) {
JdbcCursorItemReader<Customer> delegate = new JdbcCursorItemReaderBuilder<Customer>()
.name("customerCursor")
.dataSource(ds)
.sql("SELECT id, name FROM customer")
.rowMapper((rs, i) -> new Customer(rs.getLong("id"), rs.getString("name")))
.saveState(false) // restart-from-scratch: safer under concurrency
.build();
SynchronizedItemStreamReader<Customer> safe = new SynchronizedItemStreamReader<>();
safe.setDelegate(delegate); // read() is now synchronized across threads
return safe;
}go deeper
Know the reader is shared and must be wrapped in SynchronizedItemStreamReader.
Name specific stateful readers (cursor, flat-file, paging) and the wrapper for both reader and writer.
Explain that reads serialize so the benefit is in concurrent process/write, plus the restart/saveState(false) implication.
Compare against partitioning/remote chunking on correctness and scalability; decide when shared-state synchronization is the wrong tool.
## Root cause: one reader, many threads A chunk-oriented step has exactly **one** `ItemReader` bean. In single-threaded mode that's fine — `read()` is only ever called from one thread. In a **multithreaded step**, the `TaskExecutorRepeatTemplate` dispatches chunk iterations to a pool, and every worker calls `read()` on the **same instance**. `ItemReader.read()` is inherently mutating (it returns the *next* item and advances position), and the built-in readers implement that with **non-synchronized mutable state**. ## Concretely, what state gets corrupted - **`JdbcCursorItemReader` / `HibernateCursorItemReader`**: hold an open cursor (`ResultSet`). A `ResultSet` is **not thread-safe**; two threads calling `next()`/`getX()` concurrently corrupt row position and can throw driver-level exceptions. Cursor readers are the *most* dangerous here. - **`FlatFileItemReader`**: keeps a `BufferedReader` and a current line count; concurrent reads interleave lines, skip lines, or throw parsing/index errors. - **`JdbcPagingItemReader` / `JpaPagingItemReader`**: maintain a page number and an in-memory results list index. Concurrency can re-fetch or skip pages, or throw `ArrayIndexOutOfBoundsException` when the buffer index races. ## The fix: SynchronizedItemStreamReader `org.springframework.batch.item.support.SynchronizedItemStreamReader<T>` is a decorator: you set a delegate `ItemStreamReader`, and its `read()` is `synchronized`, so only one thread reads at a time. Lifecycle methods (`open`, `update`, `close` from `ItemStream`) are delegated. There's a symmetric `SynchronizedItemStreamWriter` for writers that aren't thread-safe. ```java SynchronizedItemStreamReader<Foo> reader = new SynchronizedItemStreamReader<>(); reader.setDelegate(jdbcCursorItemReader); ``` Spring Batch also offers the `@Bean`-friendly builder in some versions; but the decorator is the canonical approach. (Historically `@StepScope` was sometimes suggested but that does **not** help here — the reader is still shared within the step execution.) ## Important consequence: reading is serialized Because `read()` is now synchronized, the **reading phase is effectively single-threaded**. The throughput benefit of a multithreaded step therefore comes from the **processing and writing** running concurrently across threads, while reads happen one at a time under the lock. If your bottleneck is the *read* itself (a slow query), a multithreaded step won't help much — partitioning (each partition has its own reader over a data range) is the better tool. ## Restart / state-tracking caveat Stateful readers persist an `ItemStream` execution context (e.g. `read.count`) so a restart can resume. Under concurrency this bookkeeping no longer maps to a clean 'resume point,' so many teams set `saveState(false)` on the reader and design the job to **restart from scratch / be idempotent** rather than resume mid-way. ## Alternatives that sidestep the whole problem **Partitioning** (local or remote) gives each worker step its **own reader instance** scoped to a partition (e.g. a primary-key range), so there is no shared mutable state and no need for synchronization — plus better restartability per partition. **Remote chunking** distributes processing/writing to workers while a master reads. Reach for these when the shared-reader synchronization becomes the bottleneck or restart correctness matters.
- If SynchronizedItemStreamReader serializes reads, where does the speedup actually come from?From the processor and writer running concurrently across threads. Only read() holds the lock; processing and writing each chunk happen in parallel. So it helps I/O-bound processing/writing, not read-bound steps.
- Why is partitioning often preferable to a synchronized multithreaded reader?Partitioning gives each worker its own reader over a disjoint data range, eliminating shared mutable state and the read bottleneck, and provides per-partition restartability — at the cost of more configuration.
- Does marking the reader @StepScope make it thread-safe?No. Step scope creates one instance per step execution, still shared by all threads within that step. You still need SynchronizedItemStreamReader or partitioning.
saying these in an interview costs you the question
- Thinking each thread gets its own reader instance
- Believing @StepScope solves thread safety
- Assuming paging readers are automatically safe (their counters still race)
- Expecting a multithreaded step to speed up a read-bound step despite the synchronized read()