skip to content

Idempotent Producer

How a producer ID plus per-partition sequence numbers let the broker drop duplicate retried batches. Interviewers ask it as the foundation of exactly-once, and to test the difference between no duplicates and transactions.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is Kafka's idempotent producer, and how do you turn it on?

level: juniorimportance: must knowfreq 78%

answer

  1. enable.idempotence=true
  2. default since Kafka 3.0
  3. no duplicates on retry
  4. per-partition, per-session
  5. auto: acks=all, in-flight<=5

basics

~10 s

An idempotent producer guarantees that producer retries don't create duplicate messages on a partition. You enable it by setting enable.idempotence=true (the default since Kafka 3.0).

solid answer

~40 s

Kafka's idempotent producer ensures that a message produced to a partition is written exactly once even when the producer retries after a transient failure (e.g., a network blip where the ack was lost). Normally a retry could append the same record twice; idempotence prevents that by tagging each batch with a producer ID and a per-partition sequence number that the broker uses to detect and drop duplicates. You enable it with enable.idempotence=true. Since Kafka 3.0 it is the default, and the producer auto-configures the required acks=all, retries>0, and max.in.flight.requests.per.connection<=5. The guarantee is per-partition and for the lifetime of a single producer session; it is not cross-session and not multi-partition atomic (that needs transactions).

go deeper

for a junior

Know the config name and that it stops retry duplicates.

for a middle

Know it's the 3.0 default and auto-sets acks=all and bounded in-flight.

for a senior

Articulate the per-partition / per-session scope and the boundary with transactions.

for a principal

Frame it in the delivery-semantics taxonomy and reason about when idempotence suffices vs. when EOS transactions are required.

## The problem A Kafka producer sends a batch of records to a broker and waits for an acknowledgment (ack). If the broker writes the batch but the ack is lost on the way back (network timeout, broker restart), the producer doesn't know it succeeded. Its only safe option is to **retry**. Without protection, the retry appends the same records **again**, producing **duplicates**. This is the classic at-least-once delivery problem. ## What idempotence means here *Idempotent* means "doing it more than once has the same effect as doing it once." An **idempotent producer** guarantees that no matter how many times the producer retries a given send, the record lands on the partition **exactly once**. ## How you enable it Set the producer config: ``` enable.idempotence=true ``` Since **Kafka 3.0** this is the **default**, so you usually don't set it explicitly. When enabled, the producer enforces a set of compatible configs automatically (and will throw a `ConfigException` if you set conflicting values): - `acks=all` — every in-sync replica must acknowledge. - `retries > 0` (defaults to Integer.MAX_VALUE) — so it actually retries. - `max.in.flight.requests.per.connection <= 5` — bounded so ordering and dedup can be preserved. ## What it does NOT give you - **It is per-partition.** It dedups within one topic-partition, not across partitions atomically. - **It is per-producer-session.** The guarantee covers retries within the lifetime of one producer instance; it does not survive a producer crash/restart (a new session gets a new producer ID). Cross-session, cross-partition atomicity requires **transactions** (`transactional.id` + `initTransactions()`). - It does not deduplicate at the application level — if your app code calls `send()` twice with the same payload, those are two distinct records. ## Mechanism in one line The broker assigns the producer a **producer ID (PID)** and the producer attaches a **monotonically increasing sequence number per partition** to each batch; the broker remembers the last sequence it saw and drops any batch it has already written.

  • Is idempotence on by default?
    Yes, since Kafka 3.0 enable.idempotence defaults to true (and acks defaults to all).
  • Does idempotence stop your app from sending the same message twice?
    No. It only deduplicates the producer's own retries of a single send. Two separate send() calls are two distinct records.

saying these in an interview costs you the question

  • Claiming idempotence gives exactly-once across multiple partitions or across producer restarts — that needs transactions.
  • Saying it deduplicates application-level duplicate sends.
  • Saying you must set acks and retries manually — the producer enforces compatible values automatically.

context

open as a page

How does the broker actually detect and drop duplicate batches from an idempotent producer?

level: middleimportance: must knowfreq 70%

basics

~20 s

Each producer gets a producer ID (PID). Each batch carries a per-partition sequence number that increases by one. The broker remembers the last sequence written per (PID, partition) and discards any batch it has already seen.

open as a page

What does idempotence guarantee versus what it does NOT, and when do you need transactions instead?

level: seniorimportance: must knowfreq 64%

basics

~10 s

Idempotence guarantees exactly-once writes to a single partition within one producer session. It does NOT give cross-partition atomicity, cross-session deduplication, or atomic consume-process-produce. Those require transactions (transactional.id, initTransactions, begin/commit).

open as a page

Which producer configs does enable.idempotence require, and what happens if you set a conflicting value?

level: middleimportance: should knowfreq 58%

basics

~10 s

Idempotence requires acks=all, retries>0, and max.in.flight.requests.per.connection<=5. If you explicitly set a conflicting value (e.g., acks=1), the producer fails fast with a ConfigException.

open as a page

Why does idempotence require acks=all, and how do durability settings (min.insync.replicas) interact with the exactly-once-per-partition guarantee?

level: seniorimportance: should knowfreq 42%

basics

~20 s

acks=all means a write is acknowledged only after all in-sync replicas store it, so an acked record can't be silently lost on leader failover. Idempotence dedups retries, but only acks=all makes the underlying write durable enough for the guarantee to mean exactly-once.

open as a page