skip to content

What is Remote Chunking in Spring Batch, and which part of the chunk-oriented step runs where?

level: juniorimportance: must knowfreq 45%

answer

  1. Master reads, workers process+write
  2. Reading NOT scaled — single point
  3. Data itself crosses the wire
  4. @EnableBatchIntegration
  5. Use when process/write is the bottleneck

basics

~20 s

A scaling pattern where one master step reads items and sends chunks over messaging (like a queue) to remote worker processes. The workers do the processing and writing, then reply. Reading stays central; processing/writing scales out.

solid answer

~40 s

Remote Chunking splits a single chunk-oriented step across JVMs using messaging middleware (via Spring Integration). The master (manager) runs the ItemReader, groups items into chunks, and sends each chunk as a message to a request channel backed by a broker (JMS, RabbitMQ, Kafka). Remote workers listen on that channel, run the ItemProcessor and ItemWriter on their chunk, and send a response back on a reply channel. The master aggregates responses to track chunk success/failure and commit progress. Only processing and writing are distributed — reading is still single-threaded on the master. That makes it a good fit when the read is cheap but processing or writing is the bottleneck (e.g. CPU-heavy transforms, slow external writes). It differs from remote partitioning, which distributes the whole read-process-write.

code

java · 22 lines
java
// spring-batch-integration on the classpath
@Configuration
@EnableBatchIntegration
public class ManagerConfig {

    @Autowired
    private RemoteChunkingManagerStepBuilderFactory managerStepBuilderFactory;

    // Master reads items and ships chunks to the 'requests' channel,
    // aggregating worker acks arriving on the 'replies' channel.
    @Bean
    public TaskletStep managerStep(ItemReader<Order> reader,
                                   MessageChannel requests,
                                   PollableChannel replies) {
        return this.managerStepBuilderFactory.get("managerStep")
                .<Order, Order>chunk(100)   // chunk size sent per message
                .reader(reader)             // ONLY the reader runs here
                .outputChannel(requests)    // outbound ChunkRequest
                .inputChannel(replies)      // inbound ChunkResponse
                .build();
    }
}

go deeper

for a junior

Should grasp the core split: master reads, workers process+write, messages carry the chunks.

for a middle

Should note reading isn't scaled and that data itself is sent over the broker.

for a senior

Should contrast with partitioning and reason about when the read vs process/write is the bottleneck.

for a principal

Should discuss serialization cost, middleware durability implications, and system-level throughput trade-offs.

**Remote Chunking** is one of Spring Batch's four scaling strategies (multi-threaded step, parallel steps, remote partitioning, remote chunking). It scales a *single* chunk-oriented step across multiple JVMs/machines by moving the **process** and **write** phases off the master and onto remote workers, communicating over **messaging middleware** through **Spring Integration** channels. **The chunk-oriented model (baseline):** A normal Spring Batch step loops: read N items (the `chunk` size), process each, then write the whole chunk in one transaction. `ItemReader`, `ItemProcessor`, `ItemWriter` all run in the same thread/JVM. **What Remote Chunking changes:** - **Master (a.k.a. manager):** runs the `ItemReader` single-threaded. As it accumulates a chunk, instead of processing/writing locally it serializes the chunk into a `ChunkRequest` message and puts it on an outbound (request) channel. The message travels through a broker — ActiveMQ/Artemis (JMS), RabbitMQ (AMQP), or Kafka — to workers. - **Worker:** a separate application (often many instances) listening on the request channel. Each worker deserializes the chunk, runs the `ItemProcessor` and `ItemWriter` on it, and sends a `ChunkResponse` back on a reply channel indicating success or failure. - **Master again:** consumes the responses, updates the `StepExecution` (write counts, status), and knows when all outstanding chunks are accounted for so the step can complete. **Key characteristics:** 1. **Reading is NOT scaled.** The reader is a single point on the master. If reading is your bottleneck, remote chunking won't help — use partitioning instead. 2. **Actual data crosses the wire.** The items themselves are serialized into the messages. This has two consequences: items must be serializable, and the middleware must be **durable/reliable** (guaranteed delivery), because a lost chunk message = lost/unprocessed data with no easy recovery. This is stricter than remote partitioning, which only ships small metadata (partition ranges). 3. **It's a form of remote worker load-balancing.** Multiple workers compete for messages on the same queue; the broker distributes chunks, giving natural back-pressure and horizontal scaling of process/write throughput. **When to use it:** processing or writing dominates the runtime (heavy CPU transforms, enrichment via slow remote calls, slow target systems) while the source read is comparatively cheap and inherently sequential. **Enabling it:** annotate config with `@EnableBatchIntegration` (from `spring-batch-integration`), which supplies `RemoteChunkingManagerStepBuilderFactory` (build the master step) and `RemoteChunkingWorkerBuilder` (build the worker integration flow). Under the hood the master uses a `ChunkMessageChannelItemWriter` as its writer, and each worker uses a `ChunkProcessorChunkHandler` service activator to run process+write.

  • Which phases of the step run on the master versus the workers?
    Read runs only on the master (single-threaded). Process and write run on the remote workers. The master also aggregates the workers' responses to update the StepExecution and decide completion.
  • Does remote chunking help if reading from the database is the slow part?
    No. The reader stays single-threaded on the master, so it becomes the bottleneck. For a slow read you want remote partitioning, where each worker reads its own partition.

saying these in an interview costs you the question

  • Saying the workers read the data (only the master reads)
  • Claiming remote chunking scales the reader
  • Confusing it with multi-threaded step (single JVM) — remote chunking spans JVMs
  • Thinking only metadata is sent (the actual items are serialized and sent)

context