skip to content

Why do Competing Consumers setups typically guarantee at-least-once message delivery rather than exactly-once, and what must consumer code do to stay correct under that guarantee?

level: middleimportance: must knowfreq 75%

answer

  1. ack-then-delete, ambiguous crash = redeliver
  2. at-least-once is the default, not a bug
  3. idempotent ops: SET not ADD
  4. dedupe by message ID or idempotency key
  5. exactly-once needs a shared transaction boundary

basics

~20 s

The system can't perfectly promise a message is handled exactly one time — crashes and retries mean it might get handled twice. So the code that processes a message has to be written so doing it twice causes no harm, like setting a value instead of adding to it.

solid answer

~40 s

At-least-once delivery is a structural consequence of how these systems recover from failure: a message is only removed from the queue after an explicit acknowledgment, and any ambiguity between 'processed but ack lost' and 'not processed' is resolved by redelivering. That means the same message can legitimately be delivered and processed more than once. Exactly-once would require the broker and every downstream side effect to participate in a single atomic transaction, which most queue/consumer stacks don't support end-to-end. So consumer code must be idempotent: processing the same message twice must produce the same result as processing it once. Techniques include using the message ID (or a business key) to dedupe via a database unique constraint or upsert, making operations naturally idempotent (SET rather than ADD), or maintaining a short-lived processed-ID cache for near-duplicate suppression.

go deeper

for a junior

Should know that a message might get processed more than once, and that this is expected behavior, not a bug to 'fix' by tuning timeouts.

for a middle

Should be able to explain why the broker can't tell 'processed but ack lost' apart from 'not processed,' and know at least one concrete idempotency technique (dedup key, SET vs ADD).

for a senior

Should design consumer operations to be idempotent by default and know the trade-offs of the common deduplication mechanisms (unique constraints, TTL caches, FIFO dedup windows).

for a principal

Should reason about where exactly-once is actually achievable (single-system transactional boundaries like Kafka) versus where it's a myth once side effects leave that boundary, and set organizational conventions (idempotency keys as a standard contract) rather than solving it ad hoc per consumer.

## The guarantee and where it comes from Delivery semantics describe the guarantee a messaging system makes about how many times a given message will be delivered to a consumer, and Competing Consumers systems almost universally land on **at-least-once** rather than exactly-once. The mechanism is the acknowledgment protocol described by the visibility-timeout pattern: a consumer receives a message, does its work, and then explicitly tells the broker "done, delete this." The broker cannot distinguish between two situations that look identical from its side: 1. The consumer crashed before finishing the work, so the message genuinely needs to be retried. 2. The consumer finished the work perfectly but crashed (or the network dropped the response) before the acknowledgment reached the broker. In case 2, the work already happened, but the broker has no way to know that — from its point of view, it just didn't hear back in time. The only safe default is to treat both cases the same way and redeliver, because silently treating case 1 as case 2 (i.e., never retrying) would lose messages, which is almost always the worse failure. The cost of that choice is that case 2 becomes a duplicate: the same message really was processed successfully once already, and it gets processed again. ## Why exactly-once rarely survives **Exactly-once delivery** — meaning each message is guaranteed to produce its effect exactly one time, no more, no less — is a much stronger guarantee that requires the broker and the consumer's side effects to be coordinated as a single atomic unit, typically via a shared transactional store or a two-phase-commit-like protocol. A few specialized systems approximate this in narrow contexts (Kafka's transactional producer/consumer APIs offer "exactly-once semantics" when both the read and the write stay inside the Kafka ecosystem), but the moment a consumer's side effect reaches outside that boundary — writing to an external database, calling a third-party API, sending an email — there is no longer a shared transaction to make atomic, and the guarantee reduces back to at-least-once from the application's perspective. This is why the practical, portable rule for Competing Consumers is: assume at-least-once, and design accordingly, rather than chase a stronger guarantee that rarely survives contact with real infrastructure. ## Idempotency, the constraint the consumer absorbs The trade-off, then, isn't really a choice — it's a constraint the consumer must absorb. This is where **idempotency** comes in: a consumer operation is idempotent if running it N times has the same observable effect as running it once. - **Make the operation naturally idempotent.** The cleanest way to get this is to design the operation itself to be naturally idempotent — an "update the account balance to $500" message is idempotent by construction (it's a SET), whereas "add $50 to the account balance" is not (running it twice adds $100). - **Deduplication keyed on something stable.** When the operation can't be made naturally idempotent — e.g., "charge the customer's card" really is an action, not a state assignment — the standard technique is deduplication keyed on something stable: the message's unique ID, or a business-level idempotency key the producer attaches (many payment APIs formalize this as an explicit idempotency key the caller supplies). - **Record the keys already processed.** The consumer records processed keys, typically via a database unique constraint or a fast key-value store with a TTL, and checks that record before doing the side effect — if the key's already been seen, it short-circuits and returns the prior result instead of repeating the action. ## What the failure mode looks like The failure mode when this discipline is skipped shows up concretely and often expensively. - A **payment-processing consumer** that isn't idempotent, running behind a Competing Consumers queue, will occasionally double-charge a customer: the charge succeeds, the acknowledgment is lost to a network blip or a mid-flight crash, the message is redelivered to (possibly) a different instance, and it charges the card again because nothing recorded that the first charge already happened. - An **inventory-decrement consumer** that isn't idempotent will oversell stock under load, because a redelivered "decrement by 1" message decrements again. These bugs are notoriously hard to catch in testing because they only manifest under the exact conditions Competing Consumers is built to handle — crashes, redeliveries, timing races — which are rare and load-dependent, so they surface in production, at scale, and often correlate with deploys (which kill consumer instances mid-message) or transient network issues, not with any obvious code change. ## Where it shows up A concrete, well-known real-world pattern: AWS's own SQS documentation explicitly states standard queues provide at-least-once delivery (with occasional out-of-order delivery too) and recommends designing consumers to be idempotent; SQS FIFO queues add deduplication (via a short dedup window keyed on message content or an explicit deduplication ID) specifically to narrow — not eliminate — the practical rate of duplicates, while still describing the overall contract as effectively-once rather than a hard guarantee, precisely because the same crash-before-ack ambiguity exists at the FIFO layer too.

  • Could you just make the broker wait longer for an acknowledgment to avoid duplicates?
    No — waiting longer only delays when the ambiguity gets resolved, it doesn't remove it. The broker still can't distinguish 'consumer is slow but alive' from 'consumer died and the ack will never come,' so at some point it has to guess by timing out, and that guess can still be wrong in either direction. Longer waits mostly trade duplicate risk for slower crash recovery.
  • What's a practical way to implement deduplication without a dedicated dedup service?
    A common pattern is a unique constraint on the message/idempotency ID in whatever database the consumer already writes to: the consumer attempts an insert or upsert keyed on that ID as part of the same transaction as the business write, and if the insert violates the unique constraint, it knows this message was already processed and skips (or returns the prior result). This piggybacks on infrastructure you already have instead of adding a new moving part.
  • Does at-least-once delivery mean messages can also arrive out of order?
    Often yes, especially on non-FIFO queues, because redelivery and multiple consumers mean a message that was retried can land after messages that were originally behind it. At-least-once and ordering are separate guarantees — a queue can offer one without the other, and systems that need both have to explicitly engineer for it, usually at some throughput cost.

Like a food-delivery app that resends your order to the kitchen if it doesn't get a 'received' confirmation in time — usually because the confirmation itself got lost, not because the order didn't arrive. The kitchen has to recognize 'oh, order #482 again' and not cook two meals, rather than blindly making a duplicate every time the ticket reprints.

saying these in an interview costs you the question

  • Claims the broker guarantees exactly-once delivery by default
  • Writes consumer logic that increments/appends without any dedup safeguard
  • Doesn't recognize 'ack lost after successful processing' as the real source of duplicates
  • Thinks retries only happen when processing genuinely failed
  • No mention of idempotency keys or unique-constraint-based deduplication

context