skip to content

As an architect, how do you choose between a split, a multi-threaded step, and partitioning for scaling a batch job?

level: principalimportance: should knowfreq 40%

answer

  1. split = different work; partitioning = same work, sliced data
  2. multi-threaded step = chunk threads, poor restart, thread-safe reader
  3. partitioning = Partitioner + PartitionHandler, local or remote, clean restart
  4. compose: split of partitioned steps
  5. bounded pool sized vs DB connections

basics

~20 s

Use a split to run different, independent steps in parallel. Use a multi-threaded step to run chunks of one step across threads. Use partitioning to run the same step over many data slices, optionally across machines. Split parallelizes work; partitioning parallelizes data.

solid answer

~50 s

These solve different scaling problems. A split (FlowBuilder.split) gives flow-level parallelism: distinct, independent steps run concurrently on a TaskExecutor and join afterwards — ideal when a job has several unrelated stages (load three files, refresh two caches). A multi-threaded step runs the chunks of a single step across a thread pool; it's simple but hard to restart safely and needs thread-safe, cursor-free readers. Partitioning splits one step's dataset into partitions, each processed by an identical worker step, coordinated by a Partitioner and PartitionHandler — it restarts cleanly per partition and can scale out to remote workers (with remote partitioning). So: split = different work in parallel; multi-threaded step = same step, chunk-level threads, in-JVM; partitioning = same step, data-sliced, local or remote, restartable. In practice you often combine them — a split whose branches are each partitioned steps.

code

java · 20 lines
java
// Composition: a split whose branches are each a partitioned step.
Step partitionedLoad = new StepBuilder("partitionedLoad", jobRepository)
        .partitioner("workerLoad", fileRangePartitioner)
        .step(workerLoadStep)
        .gridSize(8)
        .taskExecutor(partitionExecutor)   // data-level parallelism
        .build();

Flow loadFlow = new FlowBuilder<SimpleFlow>("loadFlow")
        .start(partitionedLoad).build();
Flow reindexFlow = new FlowBuilder<SimpleFlow>("reindexFlow")
        .start(reindexStep).build();

Flow split = new FlowBuilder<SimpleFlow>("split")
        .split(stageExecutor)              // flow-level parallelism
        .add(loadFlow, reindexFlow)        // two independent stages in parallel
        .build();

Job job = new JobBuilder("job", jobRepository)
        .start(split).end().build();

go deeper

for a junior

Knows split runs steps in parallel; may not distinguish partitioning.

for a middle

Can define each mechanism and pick the obvious one per scenario.

for a senior

Weighs restartability, thread-safety, and resource limits when choosing among them.

for a principal

Owns the org's scaling strategy, composes split with partitioning, and reasons about metadata contention, connection-pool sizing, idempotency, and remote scale-out.

## The three in-scope-vs-sibling mechanisms Spring Batch offers several scaling strategies. The one this leaf owns is the **split**; the others are contrasted here to make the choice clear. ### 1. Split (parallel steps) — *flow-level parallelism* - **API:** `FlowBuilder.split(taskExecutor).add(flowA, flowB)`. - **Parallelizes:** different, independent **steps/flows**. - **Granularity:** whole steps. - **Restart:** clean — each branch step has its own `StepExecution`. - **Scope:** single JVM (branches run on the local `TaskExecutor`). - **Best when:** the job naturally decomposes into unrelated stages you want overlapped in time. ### 2. Multi-threaded step — *chunk-level parallelism within one step* - **API:** `StepBuilder...tasklet/chunk(...).taskExecutor(taskExecutor)`. - **Parallelizes:** the **chunks** of a **single** step across threads. - **Granularity:** chunks. - **Restart:** problematic — chunk ordering/state across threads makes reliable restart hard; typically you accept re-processing or disable restart. Reader must be **thread-safe** (avoid stateful cursor readers, or synchronize with `SynchronizedItemStreamReader`). - **Scope:** single JVM. - **Best when:** one step is CPU/IO-bound over a stream and you want quick in-process speedup and can tolerate the restart caveats. ### 3. Partitioning — *data-level parallelism, one step* - **API:** `StepBuilder.partitioner(stepName, partitioner).step(workerStep).gridSize(n).taskExecutor(...)` with a `Partitioner` (creates partition `ExecutionContext`s) and a `PartitionHandler`. - **Parallelizes:** the **same** step over disjoint **data partitions** (e.g. id ranges, files). - **Granularity:** data slices, each a full `StepExecution`. - **Restart:** clean — each partition is an independent `StepExecution`, only failed partitions re-run. - **Scope:** **local** (`TaskExecutorPartitionHandler`) or **remote** (`MessageChannelPartitionHandler` over messaging) — scales across machines. - **Best when:** a single step must process a huge dataset and you need horizontal scale and clean restart. ### 4. (Related) Remote chunking — sends items over messaging to remote workers; heavier, master reads and workers process. Mentioned for completeness; different from partitioning in that the master does the reading. ## Decision guide | Question | Choose | |---|---| | Are these *different* independent stages? | **Split** | | Is it *one* step, big stream, single JVM, restart not critical? | **Multi-threaded step** | | Is it *one* step, huge data, need clean restart and/or multi-node? | **Partitioning** | | Need the master to read and fan work out to remote nodes? | **Remote chunking** | ## Composition These aren't mutually exclusive. A common high-throughput pattern is a **split whose branches are each partitioned steps** — different stages in parallel, each stage internally data-parallel. You can also nest flows arbitrarily. ## Architectural considerations - **Correctness first:** split and partitioning give clean restart; multi-threaded step trades restartability for simplicity. - **Resource bounds:** always use a **bounded** `ThreadPoolTaskExecutor`; unbounded `SimpleAsyncTaskExecutor` in a split with many branches (or in a multi-threaded step with high concurrency) can exhaust threads and DB connections. Size the pool against the DB connection pool. - **JobRepository contention:** high fan-out (many partitions or branches) hammers the batch metadata tables; watch for lock contention. - **Idempotency:** any parallel branch/partition should be independently recoverable, since there's no cross-thread transaction. - **Observability:** name threads (`setThreadNamePrefix`) and tag metrics per branch/partition to diagnose stragglers, since the slowest unit dictates wall-clock. ## Bottom line **Split = parallelize different work (flows).** Partitioning/multi-threaded step = parallelize the same work over data. Reach for a split when the job's *structure* is parallel; reach for partitioning when the *data volume* is the bottleneck.

  • Why is a split cleanly restartable while a naive multi-threaded step often is not?
    Each split branch is a distinct step with its own StepExecution, so restart resumes only the failed branch. A multi-threaded step spreads one step's chunks across threads with no deterministic ordering of committed chunks, so Spring Batch can't reliably reconstruct where to resume — you typically accept reprocessing or disable restart.
  • When would you combine a split with partitioning?
    When a job has several independent heavy stages (favoring a split) and at least one of those stages processes a huge dataset (favoring partitioning). You make the split's branches partitioned steps: stages run in parallel, and each stage is internally data-parallel.

saying these in an interview costs you the question

  • Saying a split parallelizes a single step's data (that's partitioning)
  • Recommending a multi-threaded step where reliable restart is required
  • Ignoring bounded thread pools and DB connection sizing at high fan-out
  • Treating split and partitioning as interchangeable rather than solving different problems

context