skip to content

A consumer pulls a message off a queue and crashes before sending an acknowledgement back to the broker. What happens to that message, and what delivery guarantee does this behavior typically produce?

level: middleimportance: must knowfreq 75%

answer

  1. ack after success, not on receipt
  2. unacked message gets redelivered
  3. at-least-once = no loss, possible duplicates
  4. idempotency key absorbs duplicates
  5. visibility timeout / unacked delivery / offset commit

basics

~20 s

The broker assumes the message wasn't handled and gives it to another consumer after a timeout, so the message isn't lost. But this means the same message might get processed twice - that's called at-least-once delivery.

solid answer

~40 s

Most queue brokers only remove a message once the consumer explicitly acknowledges (acks) successful processing. If the consumer crashes, disconnects, or simply never acks within a visibility/lock timeout, the broker treats the message as un-processed and redelivers it - either to the same consumer on reconnect or to another available consumer. This makes the message durable across consumer failures, but it means the same message can be delivered and processed more than once, for example if the consumer actually finished the work but crashed before the ack made it back to the broker. This is at-least-once delivery: no message loss, but duplicates are possible. Achieving exactly-once effects on top of this requires making the consumer's processing idempotent - keyed by a message ID or business key so a duplicate delivery is a safe no-op.

go deeper

for a junior

Should know that a crash before ack causes redelivery, not message loss, and that this can create duplicates.

for a middle

Should be able to name the concrete mechanism (visibility timeout / unacked delivery / offset commit) and explain why acking too early silently breaks the guarantee.

for a senior

Should design idempotent consumers using message/business keys and reason about visibility-timeout tuning trade-offs relative to actual processing time.

for a principal

Should set org-wide conventions for idempotency-key propagation across service boundaries and evaluate when transactional/exactly-once-within-a-system mechanisms (e.g., Kafka transactions) are worth their added complexity versus at-least-once plus idempotency.

## What an acknowledgement decides Acknowledgement is the mechanism a message broker uses to know when it's safe to consider a message done and remove or advance past it — and it's the single biggest lever determining what delivery guarantee your system actually gets. ## The ack cycle **Mechanism, step by step.** 1. When a consumer receives a message, most brokers don't delete it immediately. Instead they mark it as **in-flight** or invisible and start a timer — RabbitMQ calls this an unacked delivery tied to the consumer's channel, SQS calls it a visibility timeout, Kafka (which works slightly differently, via consumer offsets rather than per-message acks) tracks a committed offset per partition. 2. The consumer processes the message and, only after finishing successfully, sends an **acknowledgement** (RabbitMQ: `basic.ack`; SQS: `DeleteMessage` call; Kafka: commit the offset past that record). 3. Only at that point does the broker consider the message consumed and either delete it (queue model) or advance the read cursor (log model, where the message physically stays but is no longer re-read by that consumer group on restart). 4. If the consumer never sends that ack — because it crashed, the process was killed, the network dropped, or it explicitly rejected the message — the broker's timer expires and it treats the message as failed: it becomes visible/available again and gets redelivered, either back to the same consumer or to a different one in a competing-consumers setup. ## Why the broker waits **Why this design exists:** it's the only way to survive consumer failure without losing work. - If the broker deleted a message the instant it handed it out (sometimes called **auto-ack** or fire-and-forget), any consumer crash between receipt and completion would silently lose that message forever — unacceptable for anything business-critical like payments or order processing. - Requiring an explicit ack after successful completion means the broker, not the consumer, is the source of truth for whether the work got done, and a crashed consumer simply results in the work being retried by someone else. ## Duplicates are the cost **Trade-off — duplicates are the cost.** The mechanism that prevents message loss inherently creates the possibility of duplicate delivery. Consider the exact crash timing: a consumer finishes the actual business work (charges a credit card, writes a database row) but crashes in the tiny window before the ack packet reaches the broker. From the broker's point of view, no ack arrived, so it redelivers — and now the business work runs a second time. This is why virtually every at-least-once system is described as giving at-least-once delivery, not exactly-once, and why this needs to be paired with **idempotent processing**: the consumer's handler should be written so that applying the same message twice, matched by a message ID, idempotency key, or natural business key, has the same effect as applying it once — for example, setting the account balance to X rather than adding $10 to the balance, or checking whether a given message ID has already been processed before doing the side-effecting work. ## Failure modes in production - **Acks before the work finishes.** The most common bug in production is a consumer that acks the message before finishing the work (to be safe against long processing times, or because of a badly placed try/catch), which silently converts the system to at-most-once — if the process then crashes mid-work, the message is gone and the work never completes, with no error raised anywhere. - **Forgetting a redelivery limit.** The opposite failure — forgetting to configure a sane redelivery/retry limit — lets a poison message (one that always fails to process, due to a malformed payload for example) get redelivered forever, endlessly cycling through consumers, consuming CPU and log volume without ever making progress. - **Visibility-timeout mis-tuning.** A third failure mode: setting it too short for the actual processing time means the broker gives up and redelivers a message that's still being legitimately processed, causing two consumers to work on it concurrently and possibly double-write; setting it far too long means a genuinely crashed consumer's message sits invisible and unprocessed for an unnecessarily long time before anyone else picks it up. ## An idempotency key absorbing a redelivery A concrete example: an e-commerce order-processing worker consumes `ChargeCard` messages from an SQS queue. The worker successfully calls the payment gateway and gets a success response, but the EC2 instance is terminated by a spot-instance reclaim before the `DeleteMessage` call completes. SQS's visibility timeout expires, the message reappears, and a second worker instance picks it up and calls the payment gateway again. Because the team stamped every `ChargeCard` message with a unique idempotency key passed straight through to the payment gateway's API, the gateway recognizes the duplicate key and returns the original charge's result instead of charging the customer twice — the at-least-once redelivery is absorbed safely by that idempotency design rather than causing a double charge.

  • How is Kafka's offset-commit model different from RabbitMQ/SQS per-message acking?
    Kafka doesn't delete messages on ack; it just advances a per-partition, per-consumer-group offset marking read up to here. Messages remain in the log for the whole retention window regardless of consumption, so a consumer group can rewind and replay past messages, which per-message-delete queue systems generally can't do once a message is acked and removed.
  • What's the difference between at-least-once and exactly-once delivery, and is true exactly-once actually achievable?
    At-least-once guarantees no loss but allows duplicates; exactly-once means each message has effect exactly one time. True end-to-end exactly-once across independent systems is generally unachievable without coordination, so in practice exactly-once is really at-least-once delivery plus idempotent processing, or transactional/atomic commit protocols (like Kafka's transactional producer/consumer) that give exactly-once within that specific system's boundary.
  • If a consumer explicitly rejects/nacks a message instead of crashing, what's the difference in broker behavior?
    An explicit nack/reject typically triggers immediate redelivery (or dead-lettering, depending on configuration) rather than waiting for a timeout, so the broker can react faster than in the silent-crash case where it has to wait out the full visibility/lock timeout before assuming failure.

It's like a manager who only crosses a task off the whiteboard once the employee reports back done. If the employee gets hit by a bus mid-task and never reports back, the manager assumes it's not done and reassigns it to someone else - even if the original employee actually finished it right before collapsing.

saying these in an interview costs you the question

  • Thinks a message is removed from the queue the moment it's delivered to a consumer
  • Believes at-least-once delivery guarantees no duplicates
  • Acks a message before processing completes to be safe
  • Doesn't know what an idempotency key is or why it's needed here
  • Assumes exactly-once is trivially achievable with no extra design work

context