A teammate says 'we enabled enable.idempotence on the producer, so our consumers are exactly-once now.' What's wrong with that statement?
answer
- producer idempotence = dedup producer RETRIES (PID + seq)
- scope: write-to-broker, one partition, one session
- does nothing for consumer reprocessing
- transactions build on top of it (not the same)
- consumer exactly-once = EOS or idempotent consumer
basics
~20 sProducer idempotence only stops the producer's own retries from writing duplicate records into a partition. It does nothing for consumer-side processing. A consumer can still reprocess a record after a crash, so consumer-side idempotence is a completely separate concern.
solid answer
~40 s`enable.idempotence=true` (default since Kafka 3.0) tags each produced record with a producer id and per-partition sequence number so the broker drops duplicates caused by *producer retries* (e.g. a network hiccup where the ack was lost and the producer resends). It guarantees no duplicate and in-order writes within a partition for that producer session. It has zero effect on the consumer side: at-least-once consumer delivery still happens because a consumer can crash after processing but before committing its offset and then reprocess. End-to-end exactly-once for a consumer's side effects requires either Kafka EOS (only for Kafka-internal sinks) or an idempotent consumer (for external sinks). The two idempotences solve different problems: producer idempotence dedups the write into Kafka; consumer idempotence dedups the processing of records out of Kafka.
go deeper
Know producer idempotence only stops duplicate writes from producer retries, not consumer reprocessing.
Explain the PID + sequence number mechanism and that consumer-side exactly-once is a separate concern.
Distinguish idempotent producer from transactions, note acks=all/default-since-3.0, and place each in the pipeline.
Articulate the full duplicate taxonomy (producer retry, consumer reprocess) and which mechanism owns each, guiding team conventions.
**Two different 'idempotences'.** The word appears on both ends of Kafka and people conflate them. **Producer idempotence (`enable.idempotence=true`).** When a producer sends a record and the broker writes it but the *acknowledgement* is lost (network blip, timeout), the producer retries — and without protection that retry appends a *second copy* of the record. Idempotent producer prevents this: the broker assigns the producer a **producer id (PID)** and the producer attaches a monotonically increasing **sequence number** per partition. The broker remembers the last sequence per (PID, partition) and rejects/ignores a record whose sequence it has already seen, so a retry doesn't duplicate. It also preserves write ordering within a partition (which is why `max.in.flight.requests.per.connection` can stay up to 5 safely). Since Kafka 3.0 it's on by default and requires `acks=all`. **Scope: producer-to-broker writes, within one partition, for one producer session.** **What it does NOT do.** It says nothing about consumers. After the record is durably (and uniquely) in the partition, a consumer still reads it under at-least-once semantics: process, then commit offset, with the well-known window where a crash before commit causes reprocessing. Producer idempotence cannot reach across to the consumer's database write, HTTP call, or offset commit. **So the teammate's claim is wrong because:** 1. It addresses the wrong end of the pipeline — write-side dedup, not processing-side dedup. 2. Consumers can and will reprocess records after crashes/rebalances regardless of producer settings. 3. Exactly-once *processing* needs a different mechanism: EOS transactions (Kafka-internal) or idempotent consumer logic (external sinks). **Related confusions to head off.** The idempotent producer is also not the same as transactions — transactions (a `transactional.id` plus `sendOffsetsToTransaction`) build *on top of* the idempotent producer to add atomic multi-partition writes and offset inclusion. And idempotent producer does not dedup logically-duplicate records you deliberately send twice from application code; it only dedups a single send that got retried by the client. **Correct framing.** Enabling producer idempotence is good hygiene (free, default) and removes one source of duplicates — producer retries. But 'effectively exactly-once' as seen by downstream systems is owned by the consumer: make its side effects idempotent, or use EOS where the sink is Kafka itself.
- What mechanism lets the idempotent producer detect a retried duplicate?Each producer gets a producer id (PID), and every record carries a monotonic per-partition sequence number. The broker tracks the last accepted sequence per (PID, partition) and discards any record with an already-seen sequence — that's a retried duplicate.
- Does enabling producer idempotence dedup a record your application code sends twice on purpose?No. It only dedups a single send the client retried internally (same PID/sequence). Two distinct application-level sends are different records with different sequences and both are kept.
saying these in an interview costs you the question
- Saying producer idempotence gives end-to-end exactly-once — it only dedups producer retries.
- Confusing producer idempotence with Kafka transactions/EOS.
- Thinking it dedups deliberate application-level duplicate sends.
- Assuming it affects consumer offset/reprocessing behavior at all.
- Forgetting it's the foundation transactions are built on, not a competitor.