When should you key your dedup store on a business idempotency key versus (topic, partition, offset)?
answer
- business key = logical identity, survives republish
- (t,p,o) = physical position, only guards redelivery
- republished event => new offset => looks new
- header idempotency-key UUID
- no producer id => fall back to (t,p,o)
basics
~20 sUse a business key (like orderId) when the same logical event can reappear at a different offset or on another topic — it survives re-publishing. Use (topic,partition,offset) only when there's no natural business id and duplicates come purely from Kafka redelivery.
solid answer
~50 sThe dedup key defines what 'the same message' means. A **business idempotency key** — a stable id the producer embeds (orderId, paymentRequestId, a UUID header) — identifies the same *logical* event no matter how many times or where it's published. It dedups across producer retries, topic re-publishing, fan-out, and even multiple source topics. The cost: the producer must reliably generate and attach it. **(topic, partition, offset)** is always available from the ConsumerRecord and needs no producer cooperation, but it identifies a *physical* position — the exact same logical event re-published gets a new offset and looks new. So (t,p,o) only protects against Kafka's own redelivery of the *same* record (rebalances, restarts), not against upstream duplicates. Practical rule: prefer a business key when the domain has a natural unique id and you care about logical duplicates; fall back to (t,p,o) when you only need to guard against consumer redelivery and no business id exists.
go deeper
Know there are two key choices: a business id from the message, or the message's offset coordinates.
Explain that a business key survives re-publishing while (t,p,o) only guards Kafka redelivery, and pick accordingly.
Reason about producer cooperation, headers vs payload, repartitioning, and composite keys for observability.
Define the dedup-key contract across producing and consuming teams and standardize header conventions org-wide.
## What a dedup key actually asserts The key you store and check answers one question: **"have I already handled *this*?"** — and the definition of *this* is entirely determined by the key. Two duplicates are only detected if they share the same key. ## Business idempotency key A **business idempotency key** is a stable identifier for a *logical* event, produced upstream and carried in the message — commonly in the value (`orderId`, `paymentRequestId`) or in a Kafka **record header** (e.g. an `idempotency-key` header holding a UUID). **Detects duplicates from:** - Producer retries (the producer resent the same logical event; even with the **idempotent producer** enabled, an application-level resend after an ack timeout can re-emit). - Re-publishing / replay (an event reprocessed and re-emitted onto the topic). - Fan-out and multiple source topics (the same logical id arriving via different physical paths). **Requires:** the upstream system to generate a stable, collision-free id and attach it. If the producer doesn't, you can't use this key. ## (topic, partition, offset) Every `ConsumerRecord` exposes `topic()`, `partition()`, and `offset()`. The triple is globally unique for a physical record placement and needs zero producer cooperation. **Detects duplicates from:** - Consumer redelivery only — the *same physical record* read again after a rebalance, restart, or offset-commit failure. **Does NOT detect:** - The same logical event published twice. Each publish lands at a *different* offset, so the triples differ and both pass the dedup check. ## Choosing | Situation | Prefer | |---|---| | Natural unique domain id exists, care about logical dupes / replays / cross-topic | Business key | | No domain id; only guarding against consumer redelivery of the same record | (t,p,o) | | Producer can't be trusted to attach a key | (t,p,o) | | Need dedup across compaction / re-publishing pipelines | Business key | ## Edge cases - **Partition reassignment doesn't change (t,p,o)** for a given record — partition and offset are stable for that record, so the triple remains a valid redelivery guard. - **Repartitioning the topic** (changing partition count) can move *future* logical events to different partitions; the triple still works for redelivery but is meaningless across a topic rewrite — another reason a business key is more robust. - A **composite** approach is valid: store the business key as the dedup key but also record (t,p,o) for observability/debugging.
- Where in a Kafka record would you carry a business idempotency key, and why a header?In a record header (e.g. 'idempotency-key'). Headers keep dedup metadata out of the business payload, are cheap to read without deserializing the value, and survive format/schema changes to the value.
- If the idempotent producer is enabled, do you still need consumer-side dedup?Often yes. The idempotent producer only suppresses duplicate writes from a single producer session's retries; it doesn't cover consumer redelivery, application-level resends after timeouts, or cross-topic/replay duplicates.
saying these in an interview costs you the question
- Claiming (topic,partition,offset) dedups re-published logical events — it doesn't, they get new offsets.
- Assuming the idempotent producer removes the need for consumer dedup.
- Using a non-unique field (like a timestamp) as the key.
- Putting the key only in the value when a header is cleaner and cheaper to read.