skip to content

In a split, what happens when one branch fails, and how does the join/status aggregation work? Is the job restartable?

level: seniorimportance: should knowfreq 38%

answer

  1. join waits for all, even after a failure
  2. worst-status-wins aggregation (Max aggregator)
  3. any FAILED -> split FAILED -> post-split step skipped
  4. restart re-runs only failed branch
  5. no cross-branch rollback; make branches idempotent

basics

~20 s

The split waits for every branch to finish (join), then aggregates statuses: if any branch failed, the whole split is FAILED, so the job fails. On restart, completed branches/steps are skipped and only the failed ones re-run.

solid answer

~50 s

A split is a barrier: even if one branch fails, the framework still waits for the others to terminate rather than cancelling them. It then aggregates each branch's FlowExecutionStatus and the worst status wins — one FAILED branch makes the split FAILED, and the step you chained with .next() after the split does not run. Because Spring Batch persists a StepExecution per branch step in the JobRepository, restarting the failed job re-executes the split: branches whose steps already completed are skipped (their COMPLETED StepExecution is reused), and only the failed branch re-runs. That restart behavior depends on steps being restartable and idempotent. A subtle gotcha: since branches commit to the same database concurrently, a failure in one branch does not roll back work another branch already committed — you must design each branch to be independently recoverable.

code

java · 20 lines
java
// If branchB's step fails, branchA still runs to completion (join),
// the split reports FAILED, and reportStep is NOT executed.
Flow branchA = new FlowBuilder<SimpleFlow>("branchA").start(stepA).build();
Flow branchB = new FlowBuilder<SimpleFlow>("branchB").start(stepB).build();

Flow split = new FlowBuilder<SimpleFlow>("split")
        .split(taskExecutor)
        .add(branchA, branchB)
        .build();

Job job = new JobBuilder("job", jobRepository)
        .start(split)
        .next(reportStep)   // skipped if any branch FAILED
        .end()
        .build();

// On restart with the same job parameters:
//   - stepA (COMPLETED) is skipped
//   - stepB (FAILED) re-runs
//   - reportStep runs only if both now COMPLETED

go deeper

for a junior

Knows a failed branch fails the job.

for a middle

Understands join-then-aggregate and that the post-split step is skipped on failure.

for a senior

Explains worst-status aggregation, per-branch StepExecution restart, and the absence of cross-branch rollback/idempotency implications.

for a principal

Designs branches for independent recoverability, reasons about JobRepository contention and hang/timeout risks at scale, and sets restart/idempotency policy.

## The join is unconditional When a split runs, the internal `SplitState` submits each branch to the `TaskExecutor` and then **waits for all of them** to complete — this is the **join** (barrier). Crucially, if branch A throws while branch B is still running, the framework does **not** proactively cancel B; it lets in-flight branches finish, then reports the combined result. This avoids leaving half-killed threads and partially-torn-down resources. ## Status aggregation Each branch produces a `FlowExecutionStatus` (`COMPLETED`, `FAILED`, `STOPPED`, etc.). The split aggregates them with a **worst-status-wins** rule via a `FlowExecutionAggregator` (default `MaxValueFlowExecutionStatusAggregator`): - All branches `COMPLETED` -> split `COMPLETED`. - Any branch `FAILED` -> split `FAILED`. - Any branch `STOPPED` (and none failed) -> split `STOPPED`. A `FAILED` split means the job is `FAILED`, and any step chained after the split (`.next(...)`) is **not** executed. ## Persistence & restart Every step in every branch has its own **`StepExecution`** persisted in the **`JobRepository`**. Because the branches share a single `JobExecution` but have distinct `StepExecution`s, Spring Batch tracks each one independently. On **restart** of the failed `JobInstance` (same job parameters): - Steps that reached `COMPLETED` are **not re-run** (standard restart semantics; assuming default `allowStartIfComplete=false`). - The **failed** branch's step(s) re-execute from where they can restart (e.g. chunk-restart for a fault-tolerant chunk step). - The split re-joins and, if all now complete, the job proceeds to the post-split step. ## Concurrency gotchas 1. **No cross-branch transaction/rollback.** Branches commit independently to the DB on separate threads. Branch A failing does **not** undo branch B's committed chunks. Design each branch to be independently recoverable / idempotent. 2. **JobRepository contention.** All branch `StepExecution` updates hit the same batch metadata tables concurrently. This is supported, but a very high split fan-out can create lock contention on the metadata schema. 3. **Shared mutable state = races.** Branches run on different threads; never share non-thread-safe objects (e.g. a mutable field on a shared bean) between them. 4. **Executor exhaustion.** If the pool is smaller than the number of branches, some branches queue and effective parallelism drops; if it's a `SimpleAsyncTaskExecutor` without a limit, you can spawn too many threads. 5. **Timeouts / hangs.** A branch that blocks forever blocks the join forever — the whole job hangs. Give branch work sensible timeouts. ## When this matters in design Use a split only when branches are **independent and each individually restartable**. If branch B depends on branch A's committed output, sequence them (`.next()`), because the split gives no ordering and no cross-branch atomicity.

  • If branch A commits 10k rows and then branch B fails, are A's rows rolled back on job failure?
    No. Branches commit independently on their own threads; there is no cross-branch transaction. A's committed work remains. You must make each branch idempotent/recoverable so a restart doesn't double-apply or corrupt data.
  • Does a failing branch immediately cancel the other running branches?
    No. The split lets in-flight branches finish (the join is unconditional), then aggregates statuses. This avoids abruptly killing threads mid-work; the failure is reported after all branches terminate.

saying these in an interview costs you the question

  • Claiming a branch failure rolls back the whole split atomically
  • Thinking a failing branch instantly aborts sibling branches
  • Assuming the post-split step still runs after a branch fails
  • Believing restart re-runs already-completed branches

context