When would you choose partitioning over a multi-threaded step, and how do you decide the partitioning strategy?
answer
- multi-threaded = 1 StepExecution, thread-safe reader, weak restart
- partitioning = N StepExecutions, per-slice restart, remote-capable
- key: stable, high-cardinality, even
- balance slices; watch skew
- budget partitions vs threads vs DB pool
basics
~20 sChoose partitioning when work splits cleanly into independent slices and you want per-slice restart or to scale across machines. A multi-threaded step parallelizes one StepExecution's chunks but needs a thread-safe reader and gives up clean restart. Partition on a stable, evenly-distributed key.
solid answer
~50 sBoth add parallelism, but differently. A **multi-threaded step** (`taskExecutor` on one step) runs multiple chunks of a *single* StepExecution concurrently; it's simple but the reader must be thread-safe, ordering is lost, and reliable restart is compromised because one shared ExecutionContext can't cleanly track concurrent progress. **Partitioning** gives each slice its own worker StepExecution: independent transactions, per-partition restart, and the option to run workers remotely across nodes — at the cost of a Partitioner and a divisible data model. I pick partitioning when the data splits into independent slices (ID ranges, files, tenants, date buckets), I need robust restart, or I must scale beyond one JVM. Strategy-wise: partition on a **stable, high-cardinality, evenly-distributed key** to avoid skew; balance slice sizes (compute MIN/MAX or counts first); keep partition count decoupled from thread-pool size; size the thread pool against the DB connection pool; and make writes idempotent so restarts and crashed partitions don't double-apply.
code
java · 28 lines// Multi-threaded step: one StepExecution, chunks run concurrently.
// Reader MUST be thread-safe; restart guarantees are weak.
@Bean
public Step multiThreadedStep(JobRepository repo, PlatformTransactionManager tx,
ItemReader<Record> threadSafeReader,
ItemWriter<Record> writer, TaskExecutor exec) {
return new StepBuilder("mtStep", repo)
.<Record, Record>chunk(100, tx)
.reader(threadSafeReader) // e.g. SynchronizedItemStreamReader wrapper
.writer(writer)
.taskExecutor(exec) // parallelize chunks of THIS step
.build();
}
// Modulo partitioning: even distribution regardless of ID gaps.
public class ModuloPartitioner implements Partitioner {
@Override public Map<String, ExecutionContext> partition(int gridSize) {
Map<String, ExecutionContext> parts = new HashMap<>();
for (int k = 0; k < gridSize; k++) {
ExecutionContext ec = new ExecutionContext();
ec.putInt("mod", k);
ec.putInt("gridSize", gridSize);
parts.put("partition-" + k, ec); // deterministic name
}
return parts;
}
}
// Worker reader: WHERE MOD(id, :gridSize) = :mod (each slice ~ equal size)go deeper
Know that both add parallelism but partitioning splits into independent slices.
Contrast shared-StepExecution multi-threading vs per-partition StepExecutions and restart implications.
Reason about key choice, skew, and resource budgeting (threads vs partitions vs connections).
Own the end-to-end design: divisibility, balance strategy, idempotency, local-vs-remote evolution, and observability.
## Two different parallelism models ### Multi-threaded step Attach a `TaskExecutor` directly to a single step (`.taskExecutor(...)`). Spring Batch then processes **multiple chunks of the same StepExecution** concurrently. - **Pros**: minimal config; no Partitioner; good quick win for a CPU/IO-bound step with a thread-safe reader. - **Cons**: the `ItemReader` **must be thread-safe** (or wrapped in `SynchronizedItemStreamReader`), which often serializes reads and caps gains; **item ordering is not preserved**; and **restart is unreliable** because a single StepExecution/ExecutionContext can't accurately record which items across concurrent chunks were processed. Saving state is effectively disabled/dangerous. ### Partitioning A manager step splits work into N slices, each a **separate worker StepExecution** with its own reader/writer, transactions, and ExecutionContext. - **Pros**: clean **per-partition restart**; each slice is independent so no shared-reader contention; can run **remotely** across nodes by swapping the PartitionHandler; natural fit for file-per-worker or range-per-worker. - **Cons**: requires the data to be **divisible** and a `Partitioner` to express the division; more moving parts; susceptible to **data skew** if slices are unbalanced. ### Rule of thumb - Data splits cleanly + need robust restart or multi-node scaling → **partitioning**. - Single indivisible stream, simple speed-up, restart not critical → **multi-threaded step**. - (For a different bottleneck — expensive processing/writing but reads are the natural single stream — *remote chunking* keeps one reader and distributes processing; that's a sibling topic.) ## Designing the partition strategy 1. **Pick a good partition key**: stable (not mutated by the job), high cardinality, and **evenly distributed**. Auto-increment IDs are convenient but leave gaps after deletes → uneven ranges. Consider **modulo/hash partitioning** (`WHERE MOD(id, N) = k`) for even distribution regardless of gaps, at the cost of full scans per partition. 2. **Balance slice sizes**: query `MIN/MAX` or row counts up front and size ranges so partitions have similar work; skew means the slowest partition sets wall-clock time. 3. **Choose partition count vs concurrency independently**: many small partitions improve balance and restart granularity, but each adds metadata overhead and (when concurrent) a DB connection. gridSize (partition count) and pool threads are separate knobs. 4. **Budget resources**: max concurrent partitions ≤ DB connection pool headroom; a bounded `ThreadPoolTaskExecutor` prevents connection starvation/deadlock. 5. **Idempotency**: because a crashed partition reruns (possibly after partial writes), make writes idempotent (upserts, natural keys, dedupe) so restart is safe. 6. **Avoid partitioning on mutated columns**: if the job updates the column you partition by, boundaries shift on restart — partition on an immutable key. 7. **Observability**: per-partition StepExecutions expose per-slice counts; use them to detect skew and locate the failing slice. ## Local vs remote decision Start **local** (`TaskExecutorPartitionHandler`) — simplest, no infrastructure. Move to **remote** (`MessageChannelPartitionHandler` over a broker) only when a single JVM's CPU/memory is the ceiling and you need to spread workers across nodes; the Partitioner and worker step stay the same, so it's an incremental change. ## Common principal-level pitfalls - Treating partitioning as a drop-in for any slow step even when data can't be cleanly divided. - Ignoring skew: 'N partitions' but one holds 80% of rows → almost no speed-up. - Over-partitioning: thousands of tiny partitions swamp the metadata store and scheduler. - Concurrency exceeding the connection pool → deadlocks under load. - Non-idempotent writers turning a restart into duplicate data.
- Why is restart less reliable with a multi-threaded step than with partitioning?A multi-threaded step has a single StepExecution and one ExecutionContext, but multiple chunks commit concurrently in nondeterministic order. There's no clean per-thread record of exactly which items were processed, so resuming from a consistent point isn't guaranteed — Spring Batch effectively can't safely save/restore reader state. Partitioning avoids this by giving each slice its own StepExecution and commit point.
- Your range partitions are wildly uneven because of gaps from deleted rows. What do you do?Switch from contiguous ID ranges to modulo/hash partitioning (WHERE MOD(id, N) = k) so each slice gets roughly 1/N of surviving rows regardless of gaps, or precompute count-based boundaries (e.g. via percentiles/NTILE) instead of naive min/max division. Trade-off: modulo forces a full scan per partition.
saying these in an interview costs you the question
- Claiming multi-threaded step and partitioning are interchangeable
- Ignoring that a multi-threaded step needs a thread-safe reader and loses clean restart
- Partitioning on contiguous ID ranges with heavy gaps (skew)
- Setting concurrency above the DB connection pool
- Non-idempotent writers making restart double-apply data