skip to content

How do you configure a split with FlowBuilder and a TaskExecutor, and what does each call do?

level: middleimportance: must knowfreq 50%

answer

  1. wrap each unit as a Flow first
  2. split() -> SplitBuilder, add() -> branches, build() -> SimpleFlow
  3. TaskExecutor = the thread pool
  4. start(split).next(after).end().build()
  5. SyncTaskExecutor = no parallelism

basics

~10 s

Build each branch as a Flow, then call new FlowBuilder(name).split(taskExecutor).add(flowA, flowB).build(). split() supplies the thread pool, add() registers the concurrent branches, and the resulting flow becomes a step in the job.

solid answer

~40 s

You first wrap each parallel unit of work in a Flow (via FlowBuilder.start(step).build()). Then you create a container flow: new FlowBuilder<SimpleFlow>("split").split(taskExecutor) returns a SplitBuilder; .add(flowA, flowB, ...) registers the branches that will run concurrently; .build() yields a SimpleFlow. That flow is passed to JobBuilder.start(splitFlow) and closed with .end().build() (a FlowJobBuilder). The TaskExecutor argument is essential — it's the thread pool each branch runs on. Passing a SyncTaskExecutor would make branches run sequentially, defeating the purpose, so you use SimpleAsyncTaskExecutor for demos or, in production, a bounded ThreadPoolTaskExecutor. The split joins automatically: whatever you chain with .next() after the split runs only once all branches have completed. Each branch is a full step with its own StepExecution and ExecutionContext.

code

java · 29 lines
java
@Bean
TaskExecutor batchTaskExecutor() {
    ThreadPoolTaskExecutor exec = new ThreadPoolTaskExecutor();
    exec.setCorePoolSize(4);
    exec.setMaxPoolSize(4);
    exec.setThreadNamePrefix("batch-split-");
    exec.initialize();
    return exec;
}

@Bean
Job parallelJob(JobRepository jobRepository,
                Step stepA, Step stepB, Step finalStep,
                TaskExecutor batchTaskExecutor) {

    Flow flowA = new FlowBuilder<SimpleFlow>("flowA").start(stepA).build();
    Flow flowB = new FlowBuilder<SimpleFlow>("flowB").start(stepB).build();

    Flow split = new FlowBuilder<SimpleFlow>("split")
            .split(batchTaskExecutor)
            .add(flowA, flowB)
            .build();

    return new JobBuilder("parallelJob", jobRepository)
            .start(split)
            .next(finalStep)   // runs only after flowA AND flowB complete
            .end()
            .build();
}

go deeper

for a junior

Can name FlowBuilder.split().add().build() at a high level.

for a middle

Knows the exact builder chain, the role of the TaskExecutor, and how to embed the split in a JobBuilder.

for a senior

Also reasons about executor choice (bounded pool), thread-safety of branches, and join placement of subsequent steps.

for a principal

Sets org-wide conventions for executor sizing/naming and integrates split configuration with observability and resource limits.

## The building blocks **`FlowBuilder<Q>`** — fluent builder for a `Flow`. `<Q>` is the concrete flow type, almost always `SimpleFlow`. **`Flow`** — a named grouping of steps/transitions you can embed in a job or another flow. **`TaskExecutor`** — Spring's thread-pool abstraction (`org.springframework.core.task.TaskExecutor`). The split runs each branch by submitting it to this executor. ## Step-by-step 1. **Wrap each parallel unit in a Flow:** ```java Flow flowA = new FlowBuilder<SimpleFlow>("flowA").start(stepA).build(); Flow flowB = new FlowBuilder<SimpleFlow>("flowB").start(stepB).build(); ``` Each branch can itself be multiple sequential steps — e.g. `.start(step1).next(step2).build()`. 2. **Create the split container:** ```java Flow split = new FlowBuilder<SimpleFlow>("mySplit") .split(taskExecutor) // -> SplitBuilder<SimpleFlow> .add(flowA, flowB) // register branches -> FlowBuilder .build(); // -> SimpleFlow ``` - **`.split(TaskExecutor)`** switches the builder into split mode and returns a `SplitBuilder`. The executor it receives is where the branches will run. - **`.add(Flow...)`** registers the flows that run **concurrently**. You may pass two or more; passing one is legal but pointless. - **`.build()`** compiles it into a `SimpleFlow` whose internal state is a `SplitState`. 3. **Embed the split in the job:** ```java Job job = new JobBuilder("myJob", jobRepository) .start(split) // JobFlowBuilder .next(afterStep) // runs after the join .end() // -> FlowJobBuilder .build(); ``` ## What actually happens at runtime - The `SplitState` submits each branch to the `TaskExecutor` and collects a `Future` per branch. - It **blocks** until all futures complete (the **join** / barrier). - It **aggregates** the per-branch `FlowExecutionStatus` — the worst status wins, so any `FAILED` branch makes the split `FAILED`. ## Common configuration mistakes - **Forgetting the executor / passing `SyncTaskExecutor`** — branches then run sequentially; no parallelism. - **`SimpleAsyncTaskExecutor` in production** — it spawns a **new thread per task, unbounded** (unless you set a concurrency limit), which can exhaust threads. Prefer a `ThreadPoolTaskExecutor`. - **Sharing mutable state between branches** — races, because they run on different threads. - **Expecting `.next()` inside a branch to influence another branch** — branches are isolated. ## Kotlin note The same API works in Kotlin; just supply the generic explicitly: `FlowBuilder<SimpleFlow>("mySplit")`.

  • What happens if you pass a SyncTaskExecutor to split()?
    The branches execute one after another on the calling thread — there's no parallelism, so the split behaves like sequential steps. You must pass an asynchronous executor (e.g. ThreadPoolTaskExecutor) for real concurrency.
  • Can a single branch of a split contain more than one step?
    Yes. Each branch is a Flow, so you can chain steps with .start(s1).next(s2).build(). Those steps run sequentially within the branch, while the branches themselves run in parallel.

saying these in an interview costs you the question

  • Thinking split() alone parallelizes without needing a TaskExecutor
  • Believing SimpleAsyncTaskExecutor is safe for production (it's unbounded by default)
  • Assuming .next() after the split runs before all branches finish

context