skip to content

Partitioning

A partitioner splits the data into named partitions, each with its own context, and a handler runs a worker step per partition locally or remotely. This is the standard answer for scaling a batch job, because each worker reads its own slice.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

explore

questions

5

What is partitioning in Spring Batch, and why would you use it?

level: juniorimportance: must knowfreq 70%

answer

  1. master splits, workers process
  2. Partitioner -> Map<name, ExecutionContext>
  3. one StepExecution per partition
  4. PartitionHandler runs the workers
  5. restart reruns only failed partitions

basics

~20 s

Partitioning splits a step's work into several named partitions that run in parallel — each processes a slice of the data (e.g. an ID range). It scales a long-running step by using multiple threads or machines.

solid answer

~40 s

Partitioning is Spring Batch's main horizontal-scaling pattern for a single step. A master (manager) step uses a Partitioner to divide the input into N named partitions, each described by its own ExecutionContext (e.g. minId/maxId, or a filename). A PartitionHandler then runs one worker StepExecution per partition, in parallel. The worker step is an ordinary chunk-oriented step; each instance just reads its assigned slice from its ExecutionContext. Because every partition is a separate StepExecution with its own metadata and transaction, partitioning also gives you fine-grained restart — only failed partitions rerun. Use it when one step's data set is large and cleanly divisible (row ranges, files, tenants) and you want to cut wall-clock time by processing slices concurrently, locally across threads or remotely across nodes.

code

java · 24 lines
java
// The manager (master) step: it splits work via a Partitioner
// and delegates execution to a PartitionHandler.
@Bean
public Step managerStep(JobRepository jobRepository,
                        Partitioner rangePartitioner,
                        PartitionHandler partitionHandler) {
    return new StepBuilder("managerStep", jobRepository)
            .partitioner("workerStep", rangePartitioner) // names the worker step
            .partitionHandler(partitionHandler)
            .build();
}

// The worker step is an ordinary chunk step reused by every partition.
@Bean
public Step workerStep(JobRepository jobRepository,
                       PlatformTransactionManager tx,
                       ItemReader<Record> reader,
                       ItemWriter<Record> writer) {
    return new StepBuilder("workerStep", jobRepository)
            .<Record, Record>chunk(100, tx)
            .reader(reader)   // reader is @StepScope, reads its slice
            .writer(writer)
            .build();
}

go deeper

for a junior

Know the one-liner: split a step into parallel slices, each processing part of the data.

for a middle

Explain the Partitioner + PartitionHandler split of responsibility and the per-partition ExecutionContext.

for a senior

Add master/worker StepExecution semantics, @StepScope reader injection, and per-partition restart.

for a principal

Weigh partitioning vs multi-threaded step vs remote chunking, and design slice strategy for balance and idempotency.

## The problem partitioning solves A Spring Batch **Step** normally reads-processes-writes items one chunk at a time on a single thread. If the data set is huge (millions of rows, thousands of files), that step becomes the bottleneck. **Partitioning** scales that one step out: it divides the step's total work into independent slices called **partitions**, and runs each slice as its own **worker StepExecution**, in parallel. ## Master / worker (manager / worker) model Partitioning introduces two roles: - **Master step** (also called *manager* step): a special step whose job is to *split* the work and *distribute* it. It does not process items itself. In config it is built via `StepBuilder.partitioner(workerStepName, partitioner)`, producing a `PartitionStep`. - **Worker step**: an ordinary chunk-oriented step (reader/processor/writer). Multiple instances of it run concurrently, one per partition. Each worker instance is a full **StepExecution** with its own **ExecutionContext**, its own transactions, and its own row in the batch metadata tables. ## The two collaborators 1. **`Partitioner`** — an interface with `Map<String, ExecutionContext> partition(int gridSize)`. It decides *how* to split. It returns a map of **partition name → ExecutionContext**. Each `ExecutionContext` carries the parameters that tell that worker which slice to handle (for example `minId`/`maxId`, or a resource path). `gridSize` is a *hint* for how many partitions to create; the Partitioner may honor it or ignore it. Spring ships `MultiResourcePartitioner` to partition across a set of files, but you usually write your own for ID/date ranges. 2. **`PartitionHandler`** — decides *where/how* the workers run. `TaskExecutorPartitionHandler` (the default when using the step builder) runs workers **locally** on a `TaskExecutor` (thread pool). `MessageChannelPartitionHandler` sends partitions to **remote** workers over messaging. ## End-to-end flow 1. Master step starts. 2. A `StepExecutionSplitter` (default `SimpleStepExecutionSplitter`) calls the `Partitioner`, gets the map of named ExecutionContexts, and creates one worker `StepExecution` per entry, seeding each with its ExecutionContext. 3. The `PartitionHandler` executes those worker StepExecutions (concurrently). 4. Each worker runs the *same* worker step definition but reads its slice-specific values from the **step ExecutionContext** — typically injected with `@Value("#{stepExecutionContext['minId']}")` on a `@StepScope` reader. 5. When all workers finish, the master **aggregates** their exit statuses. If any partition FAILED, the master step is FAILED. ## Reading the partition slice Because the worker reader must be parameterized per partition, it must be `@StepScope` (late-bound). The values placed by the Partitioner into each partition's ExecutionContext become available via SpEL as `stepExecutionContext['key']`. ## Restart semantics Each partition is a distinct StepExecution, so on restart Spring Batch reruns **only the partitions that did not complete** — successful partitions are skipped. This is a major operational benefit over a single monolithic step. ## When to use it - The step's data is large and **cleanly divisible** into independent slices with no cross-slice ordering dependency. - You can express each slice as a small set of parameters (range bounds, filename, tenant id). - You want to shorten wall-clock time via concurrency (local threads) or spread load across nodes (remote). ## Gotchas - The Partitioner runs **once, up front**; it must be able to compute the slice boundaries before processing (e.g. query MIN/MAX ids first). It can't dynamically rebalance mid-run. - Slices should be roughly balanced, or the slowest partition dominates wall-clock time (data skew). - Workers run concurrently — the reader/processor/writer components must be thread-safe or `@StepScope` so each partition gets its own instance. ## Not to be confused with Partitioning is **not** the same as a multi-threaded step (`taskExecutor` on a single step, which parallelizes chunks of one StepExecution and needs a thread-safe/restart-unfriendly reader). Partitioning gives each slice its own StepExecution and clean restart.

  • What object carries each partition's parameters to its worker?
    The partition's own ExecutionContext, returned by the Partitioner in the Map<String, ExecutionContext>. The worker reader reads those values via SpEL, e.g. @Value("#{stepExecutionContext['minId']}").
  • Does the master step process any items itself?
    No. The master (manager) step only splits the work into partitions and coordinates the workers, then aggregates their statuses. All item reading/processing/writing happens in the worker StepExecutions.

saying these in an interview costs you the question

  • Thinking the master step reads/writes the data itself
  • Confusing partitioning with a single multi-threaded step
  • Believing all partitions share one ExecutionContext
  • Assuming partitions can rebalance dynamically at runtime

context

open as a page

How does the Partitioner interface work, and how does a worker step read its assigned slice?

level: middleimportance: must knowfreq 62%

basics

~10 s

Partitioner has one method: partition(int gridSize) returning Map<String, ExecutionContext>. Each entry is a named partition whose ExecutionContext holds that slice's parameters (e.g. minId/maxId). The worker's @StepScope reader reads them via SpEL like #{stepExecutionContext['minId']}.

open as a page

What is the role of PartitionHandler and TaskExecutorPartitionHandler in partitioning?

level: seniorimportance: should knowfreq 50%

basics

~10 s

The PartitionHandler executes the worker StepExecutions produced from the partitions and collects their results. TaskExecutorPartitionHandler is the local implementation: it runs each worker on a TaskExecutor (thread pool) and waits for all to finish.

open as a page

How do restart, StepExecution metadata, and thread-safety work for a partitioned step?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Each partition is a separate worker StepExecution with its own metadata and transactions. On restart, completed partitions are skipped and only failed/unfinished ones rerun. Because partitions run concurrently, reader/writer beans must be @StepScope or otherwise thread-safe.

open as a page

When would you choose partitioning over a multi-threaded step, and how do you decide the partitioning strategy?

level: principalimportance: should knowfreq 34%

basics

~20 s

Choose partitioning when work splits cleanly into independent slices and you want per-slice restart or to scale across machines. A multi-threaded step parallelizes one StepExecution's chunks but needs a thread-safe reader and gives up clean restart. Partition on a stable, evenly-distributed key.

open as a page