skip to content

Parallel Steps (Split)

A split runs independent flows concurrently within a job and joins before continuing. The simplest form of parallelism, and the right answer when two unrelated steps have no ordering constraint.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

explore

questions

5

How do you configure a split with FlowBuilder and a TaskExecutor, and what does each call do?

level: middleimportance: must knowfreq 50%

answer

  1. wrap each unit as a Flow first
  2. split() -> SplitBuilder, add() -> branches, build() -> SimpleFlow
  3. TaskExecutor = the thread pool
  4. start(split).next(after).end().build()
  5. SyncTaskExecutor = no parallelism

basics

~10 s

Build each branch as a Flow, then call new FlowBuilder(name).split(taskExecutor).add(flowA, flowB).build(). split() supplies the thread pool, add() registers the concurrent branches, and the resulting flow becomes a step in the job.

solid answer

~40 s

You first wrap each parallel unit of work in a Flow (via FlowBuilder.start(step).build()). Then you create a container flow: new FlowBuilder<SimpleFlow>("split").split(taskExecutor) returns a SplitBuilder; .add(flowA, flowB, ...) registers the branches that will run concurrently; .build() yields a SimpleFlow. That flow is passed to JobBuilder.start(splitFlow) and closed with .end().build() (a FlowJobBuilder). The TaskExecutor argument is essential — it's the thread pool each branch runs on. Passing a SyncTaskExecutor would make branches run sequentially, defeating the purpose, so you use SimpleAsyncTaskExecutor for demos or, in production, a bounded ThreadPoolTaskExecutor. The split joins automatically: whatever you chain with .next() after the split runs only once all branches have completed. Each branch is a full step with its own StepExecution and ExecutionContext.

code

java · 29 lines
java
@Bean
TaskExecutor batchTaskExecutor() {
    ThreadPoolTaskExecutor exec = new ThreadPoolTaskExecutor();
    exec.setCorePoolSize(4);
    exec.setMaxPoolSize(4);
    exec.setThreadNamePrefix("batch-split-");
    exec.initialize();
    return exec;
}

@Bean
Job parallelJob(JobRepository jobRepository,
                Step stepA, Step stepB, Step finalStep,
                TaskExecutor batchTaskExecutor) {

    Flow flowA = new FlowBuilder<SimpleFlow>("flowA").start(stepA).build();
    Flow flowB = new FlowBuilder<SimpleFlow>("flowB").start(stepB).build();

    Flow split = new FlowBuilder<SimpleFlow>("split")
            .split(batchTaskExecutor)
            .add(flowA, flowB)
            .build();

    return new JobBuilder("parallelJob", jobRepository)
            .start(split)
            .next(finalStep)   // runs only after flowA AND flowB complete
            .end()
            .build();
}

go deeper

for a junior

Can name FlowBuilder.split().add().build() at a high level.

for a middle

Knows the exact builder chain, the role of the TaskExecutor, and how to embed the split in a JobBuilder.

for a senior

Also reasons about executor choice (bounded pool), thread-safety of branches, and join placement of subsequent steps.

for a principal

Sets org-wide conventions for executor sizing/naming and integrates split configuration with observability and resource limits.

## The building blocks **`FlowBuilder<Q>`** — fluent builder for a `Flow`. `<Q>` is the concrete flow type, almost always `SimpleFlow`. **`Flow`** — a named grouping of steps/transitions you can embed in a job or another flow. **`TaskExecutor`** — Spring's thread-pool abstraction (`org.springframework.core.task.TaskExecutor`). The split runs each branch by submitting it to this executor. ## Step-by-step 1. **Wrap each parallel unit in a Flow:** ```java Flow flowA = new FlowBuilder<SimpleFlow>("flowA").start(stepA).build(); Flow flowB = new FlowBuilder<SimpleFlow>("flowB").start(stepB).build(); ``` Each branch can itself be multiple sequential steps — e.g. `.start(step1).next(step2).build()`. 2. **Create the split container:** ```java Flow split = new FlowBuilder<SimpleFlow>("mySplit") .split(taskExecutor) // -> SplitBuilder<SimpleFlow> .add(flowA, flowB) // register branches -> FlowBuilder .build(); // -> SimpleFlow ``` - **`.split(TaskExecutor)`** switches the builder into split mode and returns a `SplitBuilder`. The executor it receives is where the branches will run. - **`.add(Flow...)`** registers the flows that run **concurrently**. You may pass two or more; passing one is legal but pointless. - **`.build()`** compiles it into a `SimpleFlow` whose internal state is a `SplitState`. 3. **Embed the split in the job:** ```java Job job = new JobBuilder("myJob", jobRepository) .start(split) // JobFlowBuilder .next(afterStep) // runs after the join .end() // -> FlowJobBuilder .build(); ``` ## What actually happens at runtime - The `SplitState` submits each branch to the `TaskExecutor` and collects a `Future` per branch. - It **blocks** until all futures complete (the **join** / barrier). - It **aggregates** the per-branch `FlowExecutionStatus` — the worst status wins, so any `FAILED` branch makes the split `FAILED`. ## Common configuration mistakes - **Forgetting the executor / passing `SyncTaskExecutor`** — branches then run sequentially; no parallelism. - **`SimpleAsyncTaskExecutor` in production** — it spawns a **new thread per task, unbounded** (unless you set a concurrency limit), which can exhaust threads. Prefer a `ThreadPoolTaskExecutor`. - **Sharing mutable state between branches** — races, because they run on different threads. - **Expecting `.next()` inside a branch to influence another branch** — branches are isolated. ## Kotlin note The same API works in Kotlin; just supply the generic explicitly: `FlowBuilder<SimpleFlow>("mySplit")`.

  • What happens if you pass a SyncTaskExecutor to split()?
    The branches execute one after another on the calling thread — there's no parallelism, so the split behaves like sequential steps. You must pass an asynchronous executor (e.g. ThreadPoolTaskExecutor) for real concurrency.
  • Can a single branch of a split contain more than one step?
    Yes. Each branch is a Flow, so you can chain steps with .start(s1).next(s2).build(). Those steps run sequentially within the branch, while the branches themselves run in parallel.

saying these in an interview costs you the question

  • Thinking split() alone parallelizes without needing a TaskExecutor
  • Believing SimpleAsyncTaskExecutor is safe for production (it's unbounded by default)
  • Assuming .next() after the split runs before all branches finish

context

open as a page

What is a parallel step (split) in Spring Batch, and what problem does it solve?

level: juniorimportance: should knowfreq 45%

basics

~10 s

A split runs several independent steps/flows at the same time within one job instead of one after another, using a thread pool. The job waits for all of them to finish before moving on.

open as a page

What thread-safety and executor concerns arise from running steps in a split, and how do you mitigate them?

level: middleimportance: should knowfreq 30%

basics

~10 s

Branches run on different threads, so any shared mutable state can race. Use a bounded ThreadPoolTaskExecutor (not unbounded SimpleAsyncTaskExecutor), keep readers/writers per-branch or thread-safe, and size the pool against your DB connection pool.

open as a page

In a split, what happens when one branch fails, and how does the join/status aggregation work? Is the job restartable?

level: seniorimportance: should knowfreq 38%

basics

~20 s

The split waits for every branch to finish (join), then aggregates statuses: if any branch failed, the whole split is FAILED, so the job fails. On restart, completed branches/steps are skipped and only the failed ones re-run.

open as a page

As an architect, how do you choose between a split, a multi-threaded step, and partitioning for scaling a batch job?

level: principalimportance: should knowfreq 40%

basics

~20 s

Use a split to run different, independent steps in parallel. Use a multi-threaded step to run chunks of one step across threads. Use partitioning to run the same step over many data slices, optionally across machines. Split parallelizes work; partitioning parallelizes data.

open as a page