skip to content

Why must the messaging middleware in Remote Chunking provide guaranteed delivery, and what goes wrong otherwise?

level: seniorimportance: should knowfreq 30%

answer

  1. Message delivery = data delivery
  2. Lost request → silent data loss
  3. Lost response → hang / reprocess on restart
  4. At-least-once → need idempotent writes
  5. Persistent broker + manual acks + transactional consume

basics

~20 s

Because the actual items travel in the messages. If a request or response message is lost, those items are never processed/written or the master never learns a chunk finished — causing silent data loss or a hung/failed step. So you need a durable broker with acknowledgements.

solid answer

~50 s

In remote chunking the data itself is serialized into `ChunkRequest`/`ChunkResponse` messages, so message delivery *is* data delivery. If a request is dropped, those items are silently skipped — no exception, just missing output. If a response is dropped, the master's `ChunkMessageChannelItemWriter` waits for acks that never come, so the step can hang or fail even though the work was done, and restart may reprocess. Therefore you need durable, acknowledged, at-least-once messaging (persistent JMS/ActiveMQ, RabbitMQ with acks, Kafka with proper commit semantics) and ideally transactional message consumption on the worker so process+write and the ack succeed or fail together. This also forces you to think about idempotent writes, because at-least-once delivery can redeliver a chunk after a worker crash. Contrast with remote partitioning, which only ships metadata and lets durable *sources* be re-read, so it tolerates weaker guarantees.

go deeper

for a junior

May just know 'you need a reliable queue'.

for a middle

Should explain that items are in the messages, so loss = data loss.

for a senior

Should enumerate lost-request vs lost-response vs duplicate failure modes and the need for idempotent writes.

for a principal

Should design end-to-end: transactional consume, ack-after-commit, idempotency keys, back-pressure, and contrast with partitioning's recoverability.

This question probes whether a candidate understands the **data-integrity implications** of putting real payloads on a message bus. **Why it matters uniquely for chunking:** In remote chunking the message payload contains the **actual business items**. That's fundamentally different from remote partitioning (metadata only). So the reliability of the transport directly determines the correctness of the batch job. **Failure modes without guaranteed delivery:** 1. **Lost `ChunkRequest` (master→worker):** The chunk of items simply never reaches any worker. No processor/writer runs on them. There's no error on the master unless it's tracking outstanding counts — and even then it might just appear as missing acks. Result: **silent data loss** (records that should have been written aren't). 2. **Lost `ChunkResponse` (worker→master):** The worker actually processed and wrote the chunk, but the master never receives the acknowledgement. `ChunkMessageChannelItemWriter` counts sent-vs-received; missing responses mean it waits (potential **hang**) or eventually fails the step. On restart the master may **reprocess** already-written chunks — hence the need for **idempotent writes**. 3. **Duplicate delivery (at-least-once):** A worker crash after writing but before ack can cause the broker to redeliver the chunk to another worker → double write unless writes are idempotent or transactional. 4. **Reordering:** The components use sequence counters, but a transport that silently drops/duplicates undermines correlation. **What 'guaranteed delivery' means in practice:** - **Persistent/durable messages** so a broker restart doesn't lose in-flight chunks. - **Consumer acknowledgements** (client-ack / manual ack) so a message is only removed after the worker has committed process+write — not on mere receipt. - **Transactional or exactly-once-ish semantics** where possible: bind the worker's message ack to the same transaction as the write, so either both commit or the message is redelivered. - Concretely: ActiveMQ/Artemis with persistent JMS queues, RabbitMQ with publisher confirms + manual consumer acks, or Kafka with careful offset-commit-after-write. **Design consequences to mention:** - **Idempotent writers** (upserts, natural keys) to survive redelivery/restart. - **Serializable items** — payloads must serialize cleanly and not be enormous (large chunks = large messages = broker pressure). - **Back-pressure / throttling** so the master doesn't flood the queue faster than workers drain it. **Why partitioning is more forgiving:** it sends only `StepExecutionRequest` metadata; the worker re-reads from the durable *source of truth* (DB/file). A lost or redelivered partition request is cheaply recoverable — the data itself was never at risk on the wire. **Bottom line:** Remote chunking trades reliability requirements for the ability to centralize a cheap read; you *must* pair it with a durable, acknowledged broker and idempotent writes, or you risk silent correctness bugs that are very hard to detect.

  • At-least-once delivery can redeliver a chunk. How do you keep the output correct?
    Make the ItemWriter idempotent — upserts keyed on a natural/business key, or dedup on a unique constraint — so reprocessing the same chunk after a crash or restart doesn't double-write. Ideally bind the worker's message ack to the write transaction so ack and write commit together.
  • Why is remote partitioning less sensitive to message loss than remote chunking?
    Partitioning sends only metadata (partition ranges); the worker re-reads the actual data from the durable source. A lost or duplicated partition request is cheaply recoverable. In chunking the data lives only in the message, so loss means the items are gone.

saying these in an interview costs you the question

  • Saying delivery reliability doesn't matter because Spring Batch retries automatically
  • Ignoring idempotency under at-least-once delivery
  • Assuming a lost chunk raises an obvious error (it can be silent)
  • Using a non-durable/in-memory transport for production chunking

context