What is a parallel step (split) in Spring Batch, and what problem does it solve?
answer
- independent flows, concurrent
- FlowBuilder.split(taskExecutor).add(...)
- barrier/join before continuing
- step-level, not data-level parallelism
- any branch fails -> split fails
basics
~10 sA split runs several independent steps/flows at the same time within one job instead of one after another, using a thread pool. The job waits for all of them to finish before moving on.
solid answer
~40 sA split is Spring Batch's way of running multiple independent flows concurrently inside a single job. Normally steps execute sequentially; with a split you group flows and hand them a TaskExecutor, so each flow runs on its own thread in parallel. The classic use case is when several steps have no data dependency on each other — e.g. loading three unrelated files, or refreshing two caches — and running them serially wastes wall-clock time. The framework automatically joins (waits for all branches) before the job continues, and aggregates their statuses so the split fails if any branch fails. You build it with FlowBuilder.split(taskExecutor).add(flow1, flow2). The key precondition: the parallel flows must be truly independent, since ordering between them is not guaranteed.
code
java · 22 lines// Two independent imports run in parallel, then the job continues
Flow loadCustomers = new FlowBuilder<SimpleFlow>("loadCustomers")
.start(customerStep)
.build();
Flow loadProducts = new FlowBuilder<SimpleFlow>("loadProducts")
.start(productStep)
.build();
Flow parallel = new FlowBuilder<SimpleFlow>("parallelLoads")
.split(new SimpleAsyncTaskExecutor("batch-"))
.add(loadCustomers, loadProducts) // run concurrently
.build();
@Bean
Job importJob(JobRepository repo, Step reportStep) {
return new JobBuilder("importJob", repo)
.start(parallel) // both loads run in parallel...
.next(reportStep) // ...then this runs after BOTH finish (join)
.end()
.build();
}go deeper
Know that a split runs independent steps in parallel within one job and waits for all to finish.
Should recall the FlowBuilder.split(taskExecutor).add(...) API and the join/barrier behavior.
Should distinguish split (step-level parallelism) from partitioning (data-level) and know the independence requirement and failure aggregation.
Frames split within the broader scaling toolbox and evaluates when flow-level parallelism is the right tool vs. partitioning or async processing.
## What it is In Spring Batch, a **Job** is made of **Step**s that, by default, run **sequentially** — step B starts only after step A finishes. A **split** (also called a *parallel step* or *parallel flow*) lets you run several **flows** at the **same time** within a single job execution. A **Flow** is a reusable grouping of one or more steps. A **split** takes two or more flows and executes them **concurrently**, each on its own thread supplied by a **TaskExecutor** (Spring's abstraction over a thread pool). ## Why it exists Many batch jobs contain steps that do not depend on each other's output. For example: - Import `customers.csv`, `orders.csv`, and `products.csv` — three independent reads. - Rebuild two unrelated report tables. Running them one after another wastes time. A split runs them together, cutting the job's total wall-clock time roughly to the length of the slowest branch (subject to I/O and CPU limits). ## How you declare it ```java Flow splitFlow = new FlowBuilder<SimpleFlow>("splitFlow") .split(new SimpleAsyncTaskExecutor()) .add(flowA, flowB, flowC) .build(); ``` `split(TaskExecutor)` returns a **SplitBuilder**; `.add(Flow...)` registers the branches; `.build()` produces the combined flow, which you then `start(...)` in a `JobBuilder`. ## Join semantics The split acts as a **barrier**: the job does not proceed past the split until **every** branch has terminated. The framework then **aggregates** the branch statuses — if any branch ends in `FAILED`, the split (and the job) is `FAILED`; the job is `COMPLETED` only if all branches complete. ## The one hard rule The parallel flows must be **independent**. There is no ordering guarantee among them, they may touch the database concurrently, and they run on different threads. If flow B needs data flow A produced, do **not** put them in the same split — chain them sequentially instead. ## When NOT to use it - When steps have data dependencies (use normal sequential flow). - When you want to parallelize *within* a single step over a large dataset — that's **partitioning** or a **multi-threaded step**, not a split. A split is about **step-level / flow-level** parallelism (different work), whereas partitioning is about **data-level** parallelism (same work, different slices).
- What must be true about the flows you put in a split?They must be independent — no data or ordering dependency between them, since they run concurrently on separate threads with no guaranteed order.