A producer publishes 'order-placed' events and retries publishing on timeout, so the broker may assign a different message ID to what is logically the same business event. Why is deduplicating on the broker-assigned message ID insufficient here, and what should the dedup key be instead?
answer
- broker message ID = delivery attempt, not business event
- producer retry after timeout -> new message ID
- idempotency key generated once, reused across retries
- producer-consumer contract, not consumer-only fix
- Stripe client libraries reuse the key across retries
basics
~20 sIf the sender retries and the queue gives the retry a brand-new message ID, checking 'have I seen this exact message ID' won't catch it as a duplicate. Instead, dedupe on something the business itself considers the same thing every time, like the order's own ID, so retries collapse together no matter what ID the broker assigns.
solid answer
~50 sBroker-assigned message IDs identify a delivery attempt, not a business event; if the producer times out and retries the publish, the broker may mint a fresh message ID for that retry, even though semantically it's the same order-placed event being sent again. Deduplicating on that broker ID only catches redelivery of the exact same physical message, the broker resending because it wasn't acked; it does nothing for producer-side retries, which look like two entirely distinct, never-before-seen messages to the consumer. The fix is to dedupe on a business-level idempotency key that the producer generates once and keeps stable across its own retries, typically a UUID it creates before the first publish attempt and embeds in the payload, or a natural key like order_id plus operation type if the operation is inherently unique per order, so every retry of the same logical event carries the same key regardless of what transport-level message ID the broker assigns.
go deeper
Should sense that the same thing happening twice can arrive under different technical IDs, even if they can't yet name the fix precisely.
Should understand that dedup keys need to represent the business event, not the transport message, and give a basic example of a business key.
Should design the producer-side discipline, generate once, reuse across retries, and identify this as a producer-consumer contract, not a consumer-only fix.
Should reason about multi-producer scenarios, deterministic key derivation across services, and how this interacts with the transactional outbox pattern where the outbox row's own primary key often serves as the natural idempotency key.
## What a broker message ID actually identifies A message ID assigned by a broker — Kafka's offset, SQS's `MessageId`, or a UUID a client library generates per publish call — identifies **one delivery attempt of one physical message on the wire**. It says nothing about the business intent behind that message. 1. When a producer calls `publish()` and the network times out before it gets an acknowledgment from the broker, the producer typically cannot tell whether the publish actually succeeded server-side or was lost in flight. 2. The standard, safe response is to retry the publish. 3. If the retry succeeds, the broker now holds two physical messages: the original, which may or may not have actually landed, and the retry, each with its own broker-assigned message ID, both representing the same logical order-placed event from the producer's point of view. ## Why deduplicating on that ID is blind A consumer that deduplicates on the broker's message ID will treat these as two entirely separate, never-before-seen messages; its dedup check for 'have I processed message ID X' will pass for both, because X is different for each. That defeats the entire purpose of idempotent consumption: - The consumer **correctly protects against** the broker redelivering the exact same physical message, for example because its own ack was lost or a consumer-group rebalance caused reassignment before an offset commit. - But it **is blind to producer-side retries**, which are actually the more common source of true duplicates in many systems, since they originate from ordinary network flakiness on the publish path rather than from consumer crashes. ## The fix: key on business-event identity The fix is to dedupe on a key chosen for **business-event identity, not delivery-attempt identity**, commonly called an idempotency key or idempotency token. The producer: 1. generates this key once, before its first publish attempt — a UUID generated client-side, or a natural key derived from the business operation such as `order_id` combined with an operation type like `order-placed` if that combination is guaranteed unique, 2. embeds it in the message payload or a header, 3. and reuses the identical key on every retry of that same logical publish. The consumer's dedup check then keys off that field instead of the broker's message ID, so both the original and the retried physical message, despite having different broker IDs, collapse to the same dedup-table entry, and the second one is correctly recognized and skipped. ## Why this is a contract, not a consumer-local fix The trade-off is that this now requires **producer-side discipline that the consumer cannot enforce on its own**: if a producer's client library doesn't consistently reuse the same idempotency key across retries, for instance if the key is generated inside the retry loop instead of before it, the dedup key changes every attempt just like the broker's message ID does, and the whole scheme silently fails to protect anything. This pushes real design responsibility onto the producer side, and it means idempotent-consumer design isn't a purely consumer-local concern; it's a contract between producer and consumer about what identifies the same event, which needs to be documented and tested on both sides, not something a consumer team can unilaterally guarantee by adding a dedup table. ## Where it silently fails **Failure modes seen in production:** - **A common one is a producer library that retries transparently at the transport layer** without exposing or reusing an idempotency key at the message level, so every retry is invisible to the application code and looks like an entirely fresh publish with a fresh ID; teams discover this only when duplicate business events show up despite a seemingly correct consumer-side dedup table. - **Another is a natural-key choice that isn't actually unique** across the operations that will use it, for example using bare `order_id` as the dedup key when an order can legitimately be updated multiple times, each producing its own `order-updated` event, which would wrongly treat every legitimate subsequent update as a duplicate of the first and silently drop it. ## The producer's half of the pattern **A concrete real-world illustration** is the same Stripe `Idempotency-Key` pattern used elsewhere, but seen from the producer's obligation rather than the server's: Stripe's own client libraries generate the idempotency key once per logical operation and explicitly reuse it across the library's internal retry attempts, which is precisely the discipline this question is testing for. Without that client-side reuse, Stripe's server-side idempotency-key deduplication would be completely ineffective against the very retries it exists to protect against.
- Where should the idempotency key be generated, inside the retry loop or before it?Before it, and exactly once per logical operation. If the key generation happens inside the retry loop, every retry attempt would compute a fresh key, which reproduces the exact bug this pattern is meant to solve; the consumer would once again see what look like distinct events.
- What if two different producers can legitimately publish the same business event, such as two services both reacting to the same upstream trigger?Then the idempotency key needs to be derived from something both producers can independently compute identically, such as a deterministic hash of the triggering event's own ID plus the operation type, rather than a randomly generated UUID unique to one producer's process; otherwise the two producers' same event would get two different keys and dedup would fail to catch the overlap.
- Is a natural business key like order_id always a safe idempotency key on its own?Only if it's actually unique to the specific operation being deduplicated; order_id alone is unsafe for an event type that legitimately fires multiple times per order, like order-updated, because it would make every real, distinct update look like a duplicate of the first. It needs to be combined with something that makes it unique per logical occurrence, such as an operation type plus a sequence number or a version.
It's like a customer resubmitting a paper form because the first submission's mailbox receipt got lost, and the post office stamping each mailing with a new tracking number. If the office only checks tracking numbers for duplicates, it'll process both submissions as separate requests, unless the form itself has the customer's own reference number on it that stays the same across both mailings.
saying these in an interview costs you the question
- Assumes broker message ID and business-event identity are always the same thing
- Places the fix entirely in the consumer with no mention of producer-side key generation/reuse
- Suggests generating the idempotency key freshly on each retry attempt
- Uses a non-unique natural key like a bare order_id for an event type that can legitimately repeat
- Doesn't recognize producer-side retries as a source of duplicates distinct from broker redelivery