In a split, what happens when one branch fails, and how does the join/status aggregation work? Is the job restartable?
answer
- join waits for all, even after a failure
- worst-status-wins aggregation (Max aggregator)
- any FAILED -> split FAILED -> post-split step skipped
- restart re-runs only failed branch
- no cross-branch rollback; make branches idempotent
basics
~20 sThe split waits for every branch to finish (join), then aggregates statuses: if any branch failed, the whole split is FAILED, so the job fails. On restart, completed branches/steps are skipped and only the failed ones re-run.
solid answer
~50 sA split is a barrier: even if one branch fails, the framework still waits for the others to terminate rather than cancelling them. It then aggregates each branch's FlowExecutionStatus and the worst status wins — one FAILED branch makes the split FAILED, and the step you chained with .next() after the split does not run. Because Spring Batch persists a StepExecution per branch step in the JobRepository, restarting the failed job re-executes the split: branches whose steps already completed are skipped (their COMPLETED StepExecution is reused), and only the failed branch re-runs. That restart behavior depends on steps being restartable and idempotent. A subtle gotcha: since branches commit to the same database concurrently, a failure in one branch does not roll back work another branch already committed — you must design each branch to be independently recoverable.
code
java · 20 lines// If branchB's step fails, branchA still runs to completion (join),
// the split reports FAILED, and reportStep is NOT executed.
Flow branchA = new FlowBuilder<SimpleFlow>("branchA").start(stepA).build();
Flow branchB = new FlowBuilder<SimpleFlow>("branchB").start(stepB).build();
Flow split = new FlowBuilder<SimpleFlow>("split")
.split(taskExecutor)
.add(branchA, branchB)
.build();
Job job = new JobBuilder("job", jobRepository)
.start(split)
.next(reportStep) // skipped if any branch FAILED
.end()
.build();
// On restart with the same job parameters:
// - stepA (COMPLETED) is skipped
// - stepB (FAILED) re-runs
// - reportStep runs only if both now COMPLETEDgo deeper
Knows a failed branch fails the job.
Understands join-then-aggregate and that the post-split step is skipped on failure.
Explains worst-status aggregation, per-branch StepExecution restart, and the absence of cross-branch rollback/idempotency implications.
Designs branches for independent recoverability, reasons about JobRepository contention and hang/timeout risks at scale, and sets restart/idempotency policy.
## The join is unconditional When a split runs, the internal `SplitState` submits each branch to the `TaskExecutor` and then **waits for all of them** to complete — this is the **join** (barrier). Crucially, if branch A throws while branch B is still running, the framework does **not** proactively cancel B; it lets in-flight branches finish, then reports the combined result. This avoids leaving half-killed threads and partially-torn-down resources. ## Status aggregation Each branch produces a `FlowExecutionStatus` (`COMPLETED`, `FAILED`, `STOPPED`, etc.). The split aggregates them with a **worst-status-wins** rule via a `FlowExecutionAggregator` (default `MaxValueFlowExecutionStatusAggregator`): - All branches `COMPLETED` -> split `COMPLETED`. - Any branch `FAILED` -> split `FAILED`. - Any branch `STOPPED` (and none failed) -> split `STOPPED`. A `FAILED` split means the job is `FAILED`, and any step chained after the split (`.next(...)`) is **not** executed. ## Persistence & restart Every step in every branch has its own **`StepExecution`** persisted in the **`JobRepository`**. Because the branches share a single `JobExecution` but have distinct `StepExecution`s, Spring Batch tracks each one independently. On **restart** of the failed `JobInstance` (same job parameters): - Steps that reached `COMPLETED` are **not re-run** (standard restart semantics; assuming default `allowStartIfComplete=false`). - The **failed** branch's step(s) re-execute from where they can restart (e.g. chunk-restart for a fault-tolerant chunk step). - The split re-joins and, if all now complete, the job proceeds to the post-split step. ## Concurrency gotchas 1. **No cross-branch transaction/rollback.** Branches commit independently to the DB on separate threads. Branch A failing does **not** undo branch B's committed chunks. Design each branch to be independently recoverable / idempotent. 2. **JobRepository contention.** All branch `StepExecution` updates hit the same batch metadata tables concurrently. This is supported, but a very high split fan-out can create lock contention on the metadata schema. 3. **Shared mutable state = races.** Branches run on different threads; never share non-thread-safe objects (e.g. a mutable field on a shared bean) between them. 4. **Executor exhaustion.** If the pool is smaller than the number of branches, some branches queue and effective parallelism drops; if it's a `SimpleAsyncTaskExecutor` without a limit, you can spawn too many threads. 5. **Timeouts / hangs.** A branch that blocks forever blocks the join forever — the whole job hangs. Give branch work sensible timeouts. ## When this matters in design Use a split only when branches are **independent and each individually restartable**. If branch B depends on branch A's committed output, sequence them (`.next()`), because the split gives no ordering and no cross-branch atomicity.
- If branch A commits 10k rows and then branch B fails, are A's rows rolled back on job failure?No. Branches commit independently on their own threads; there is no cross-branch transaction. A's committed work remains. You must make each branch idempotent/recoverable so a restart doesn't double-apply or corrupt data.
- Does a failing branch immediately cancel the other running branches?No. The split lets in-flight branches finish (the join is unconditional), then aggregates statuses. This avoids abruptly killing threads mid-work; the failure is reported after all branches terminate.
saying these in an interview costs you the question
- Claiming a branch failure rolls back the whole split atomically
- Thinking a failing branch instantly aborts sibling branches
- Assuming the post-split step still runs after a branch fails
- Believing restart re-runs already-completed branches