What is Remote Chunking in Spring Batch, and which part of the chunk-oriented step runs where?
answer
- Master reads, workers process+write
- Reading NOT scaled — single point
- Data itself crosses the wire
- @EnableBatchIntegration
- Use when process/write is the bottleneck
basics
~20 sA scaling pattern where one master step reads items and sends chunks over messaging (like a queue) to remote worker processes. The workers do the processing and writing, then reply. Reading stays central; processing/writing scales out.
solid answer
~40 sRemote Chunking splits a single chunk-oriented step across JVMs using messaging middleware (via Spring Integration). The master (manager) runs the ItemReader, groups items into chunks, and sends each chunk as a message to a request channel backed by a broker (JMS, RabbitMQ, Kafka). Remote workers listen on that channel, run the ItemProcessor and ItemWriter on their chunk, and send a response back on a reply channel. The master aggregates responses to track chunk success/failure and commit progress. Only processing and writing are distributed — reading is still single-threaded on the master. That makes it a good fit when the read is cheap but processing or writing is the bottleneck (e.g. CPU-heavy transforms, slow external writes). It differs from remote partitioning, which distributes the whole read-process-write.
code
java · 22 lines// spring-batch-integration on the classpath
@Configuration
@EnableBatchIntegration
public class ManagerConfig {
@Autowired
private RemoteChunkingManagerStepBuilderFactory managerStepBuilderFactory;
// Master reads items and ships chunks to the 'requests' channel,
// aggregating worker acks arriving on the 'replies' channel.
@Bean
public TaskletStep managerStep(ItemReader<Order> reader,
MessageChannel requests,
PollableChannel replies) {
return this.managerStepBuilderFactory.get("managerStep")
.<Order, Order>chunk(100) // chunk size sent per message
.reader(reader) // ONLY the reader runs here
.outputChannel(requests) // outbound ChunkRequest
.inputChannel(replies) // inbound ChunkResponse
.build();
}
}go deeper
Should grasp the core split: master reads, workers process+write, messages carry the chunks.
Should note reading isn't scaled and that data itself is sent over the broker.
Should contrast with partitioning and reason about when the read vs process/write is the bottleneck.
Should discuss serialization cost, middleware durability implications, and system-level throughput trade-offs.
**Remote Chunking** is one of Spring Batch's four scaling strategies (multi-threaded step, parallel steps, remote partitioning, remote chunking). It scales a *single* chunk-oriented step across multiple JVMs/machines by moving the **process** and **write** phases off the master and onto remote workers, communicating over **messaging middleware** through **Spring Integration** channels. **The chunk-oriented model (baseline):** A normal Spring Batch step loops: read N items (the `chunk` size), process each, then write the whole chunk in one transaction. `ItemReader`, `ItemProcessor`, `ItemWriter` all run in the same thread/JVM. **What Remote Chunking changes:** - **Master (a.k.a. manager):** runs the `ItemReader` single-threaded. As it accumulates a chunk, instead of processing/writing locally it serializes the chunk into a `ChunkRequest` message and puts it on an outbound (request) channel. The message travels through a broker — ActiveMQ/Artemis (JMS), RabbitMQ (AMQP), or Kafka — to workers. - **Worker:** a separate application (often many instances) listening on the request channel. Each worker deserializes the chunk, runs the `ItemProcessor` and `ItemWriter` on it, and sends a `ChunkResponse` back on a reply channel indicating success or failure. - **Master again:** consumes the responses, updates the `StepExecution` (write counts, status), and knows when all outstanding chunks are accounted for so the step can complete. **Key characteristics:** 1. **Reading is NOT scaled.** The reader is a single point on the master. If reading is your bottleneck, remote chunking won't help — use partitioning instead. 2. **Actual data crosses the wire.** The items themselves are serialized into the messages. This has two consequences: items must be serializable, and the middleware must be **durable/reliable** (guaranteed delivery), because a lost chunk message = lost/unprocessed data with no easy recovery. This is stricter than remote partitioning, which only ships small metadata (partition ranges). 3. **It's a form of remote worker load-balancing.** Multiple workers compete for messages on the same queue; the broker distributes chunks, giving natural back-pressure and horizontal scaling of process/write throughput. **When to use it:** processing or writing dominates the runtime (heavy CPU transforms, enrichment via slow remote calls, slow target systems) while the source read is comparatively cheap and inherently sequential. **Enabling it:** annotate config with `@EnableBatchIntegration` (from `spring-batch-integration`), which supplies `RemoteChunkingManagerStepBuilderFactory` (build the master step) and `RemoteChunkingWorkerBuilder` (build the worker integration flow). Under the hood the master uses a `ChunkMessageChannelItemWriter` as its writer, and each worker uses a `ChunkProcessorChunkHandler` service activator to run process+write.
- Which phases of the step run on the master versus the workers?Read runs only on the master (single-threaded). Process and write run on the remote workers. The master also aggregates the workers' responses to update the StepExecution and decide completion.
- Does remote chunking help if reading from the database is the slow part?No. The reader stays single-threaded on the master, so it becomes the bottleneck. For a slow read you want remote partitioning, where each worker reads its own partition.
saying these in an interview costs you the question
- Saying the workers read the data (only the master reads)
- Claiming remote chunking scales the reader
- Confusing it with multi-threaded step (single JVM) — remote chunking spans JVMs
- Thinking only metadata is sent (the actual items are serialized and sent)