How do restart, StepExecution metadata, and thread-safety work for a partitioned step?
answer
- 1 manager + N worker StepExecutions
- each partition = own ExecutionContext + txns
- restart reruns only FAILED partitions
- stable partition names are mandatory
- concurrent => @StepScope reader/writer
basics
~20 sEach partition is a separate worker StepExecution with its own metadata and transactions. On restart, completed partitions are skipped and only failed/unfinished ones rerun. Because partitions run concurrently, reader/writer beans must be @StepScope or otherwise thread-safe.
solid answer
~40 sA partitioned step creates one manager StepExecution plus N worker StepExecutions — each worker has its **own** ExecutionContext, transaction boundaries, and read/write/commit counts in the batch metadata tables. That granularity gives clean restart: if the job is restarted, the Partitioner regenerates the same named partitions, Spring Batch matches them to prior worker StepExecutions, marks the COMPLETED ones as skippable, and reruns only the FAILED/unfinished partitions from their last commit point. For that to work the Partitioner must produce **stable, deterministic partition names** across runs. Concurrency raises thread-safety concerns: the workers run in parallel, so any reader/processor/writer holding mutable state must be `@StepScope` (a fresh instance per partition) rather than a shared singleton. The manager aggregates worker exit statuses; one failed partition fails the manager step, but successful partitions' work is preserved and not redone.
code
java · 43 lines// Deterministic partition names => restart can match completed workers.
public class StableRangePartitioner implements Partitioner {
private final long min, max;
StableRangePartitioner(long min, long max) { this.min = min; this.max = max; }
@Override public Map<String, ExecutionContext> partition(int gridSize) {
long size = (max - min) / gridSize + 1;
Map<String, ExecutionContext> parts = new HashMap<>();
long start = min; int i = 0;
while (start <= max) {
long end = Math.min(start + size - 1, max);
ExecutionContext ec = new ExecutionContext();
ec.putLong("minId", start);
ec.putLong("maxId", end);
// Name derived deterministically from the slice, NOT from time/UUID:
parts.put("partition-" + start + "-" + end, ec);
start += size; i++;
}
return parts;
}
}
// Concurrency safety: @StepScope so each partition gets its own restartable reader.
@Bean
@StepScope
public JdbcPagingItemReader<Record> reader(
@Value("#{stepExecutionContext['minId']}") Long minId,
@Value("#{stepExecutionContext['maxId']}") Long maxId,
DataSource ds) {
// JdbcPagingItemReader saves its page position to the ExecutionContext,
// so a FAILED partition resumes from the last commit on restart.
return new JdbcPagingItemReaderBuilder<Record>()
.name("pagingReader") // name enables state saving for restart
.dataSource(ds)
.pageSize(200)
.selectClause("SELECT id, payload")
.fromClause("FROM records")
.whereClause("WHERE id BETWEEN :minId AND :maxId")
.parameterValues(Map.of("minId", minId, "maxId", maxId))
.sortKeys(Map.of("id", Order.ASCENDING))
.rowMapper(new RecordRowMapper())
.build();
}go deeper
Know that each partition is its own StepExecution and only failed ones rerun on restart.
Explain deterministic naming and @StepScope for concurrent readers/writers.
Detail metadata rows, aggregation, per-partition commit-point resume, and idempotency needs.
Design for crash consistency: stable boundaries, idempotent writes, and safe restart under data mutation and skew.
## Metadata: many StepExecutions, not one In a normal step, a Step run = one **StepExecution** row. In a partitioned step you get: - **1 manager StepExecution** (the `PartitionStep`), which does no item I/O. - **N worker StepExecutions**, one per partition, each a first-class row in `BATCH_STEP_EXECUTION` with its own `read_count`, `write_count`, `commit_count`, status, and a linked **`BATCH_STEP_EXECUTION_CONTEXT`** holding that partition's ExecutionContext (its slice params plus any restart data the reader saves). Worker step names are typically the base worker step name plus the partition name (e.g. `workerStep:partition3`), so they're individually identifiable. ## Restart semantics — the big payoff Because each partition is its own StepExecution: 1. On restart of a FAILED job, the manager step re-runs. 2. The **`Partitioner` runs again** and must produce the **same partition names** as the first run. 3. Spring Batch matches new partition names to the previous worker StepExecutions: - Partitions that were **COMPLETED** are recognized and **not re-executed**. - Partitions that **FAILED** (or never ran) are re-executed. A restartable reader (one that saved its position in the ExecutionContext) resumes from the **last commit point**, not from the start of the slice. This is why **deterministic partition naming** is essential: if names change between runs, Spring Batch can't match old executions and may reprocess completed slices. ## Thread-safety under concurrency Workers run concurrently on the PartitionHandler's TaskExecutor. Implications: - **Readers/writers with mutable cursor/state must be `@StepScope`** so each partition gets its own instance — sharing a single `JdbcCursorItemReader` across threads corrupts its cursor. - **Stateless components** (e.g. a pure `ItemProcessor`) can be singletons safely. - Downstream systems the writers touch must tolerate concurrent writes (e.g. no single-connection assumptions). ## Aggregation of results When all workers finish, a **`StepExecutionAggregator`** (default `DefaultStepExecutionAggregator`) rolls up worker counts and statuses onto the manager StepExecution: if any worker is FAILED, the manager is FAILED; the read/write/commit counts are summed. This aggregated view is what the job sees. ## Transactions and skip/retry Each partition has independent transactions and its own skip/retry limits — a skip in one partition doesn't consume another's skip budget. Chunk commit and rollback are per worker StepExecution. ## Gotchas & edge cases - **Non-deterministic Partitioner** (e.g. re-querying MIN/MAX that changed because completed partitions altered the data): can shift boundaries on restart and reprocess or miss rows. Prefer fixing boundaries or making processing idempotent. - **In-flight partition on crash**: a partition that was RUNNING when the JVM died may be left in an inconsistent RUNNING/UNKNOWN state; Spring Batch marks it for rerun on restart, but you should ensure writes are idempotent so partial-then-rerun doesn't double-apply. - **Data mutated by processing** breaks range assumptions if you partition by a column you also update; partition on a stable key. - **Aggregated counts** on the manager are sums; per-partition detail lives on the worker rows — check those when diagnosing skew or a specific failure.
- Why must partition names be deterministic across runs?On restart the Partitioner runs again; Spring Batch matches the new partition names to previously persisted worker StepExecutions. If a name is identical to a COMPLETED partition, that partition is skipped. Random/time-based names won't match, so completed slices would be reprocessed (and possibly double-applied).
- If one partition out of ten fails, what reruns on restart?Only that failed partition's worker StepExecution reruns — resuming from its last commit point if the reader is restartable. The nine COMPLETED partitions are recognized and skipped, so their work isn't redone.
saying these in an interview costs you the question
- Thinking a partitioned step is a single StepExecution
- Using UUID/timestamp partition names, breaking restart matching
- Sharing one stateful reader across concurrent partitions
- Assuming restart reruns all partitions, not just failed ones
- Partitioning on a column the processing mutates