skip to content

How do the producerId, producerEpoch, and baseSequence fields in a record batch enable idempotent (exactly-once-into-the-log) production, and how does the broker use them to detect duplicates?

level: seniorimportance: must knowfreq 40%

answer

  1. PID + epoch + baseSequence in header
  2. per (PID, epoch, partition) increasing seq
  3. retry with seen seq → ack but drop (dup)
  4. gap → OUT_OF_ORDER_SEQUENCE_NUMBER
  5. epoch = fencing zombies; max.in.flight 5 safe

basics

~20 s

Each batch header carries a producerId (PID), producerEpoch, and baseSequence number. The broker tracks the last sequence it accepted per (PID, partition). If a retried batch repeats sequences, the broker drops it as a duplicate, so retries don't create duplicate records.

solid answer

~50 s

The RecordBatch v2 header holds a **producerId** (PID, assigned by the broker via InitProducerId), a **producerEpoch** (bumps to fence old/zombie producers), and a **baseSequence** (the sequence number of the first record; per-record sequences run baseSequence..baseSequence+lastOffsetDelta). With **enable.idempotence=true**, each (PID, epoch, partition) has a strictly increasing sequence. The broker keeps the **last accepted sequence** per (PID, partition) in memory (and snapshots it). On append it checks: if the incoming baseSequence equals last+1, accept; if it's a sequence already seen (a retry), it **acks success but discards** the duplicate (returns DUPLICATE_SEQUENCE_NUMBER internally); if there's a gap, it rejects with **OUT_OF_ORDER_SEQUENCE_NUMBER**. This is what makes producer retries safe — a TCP-level resend of an already-written batch is deduplicated. It guarantees no duplicates and no reordering **per partition, for the life of a PID/epoch session**; it is not end-to-end exactly-once (that needs transactions). max.in.flight.requests.per.connection up to 5 is still safe under idempotence.

go deeper

for a junior

Know there's a producer id and sequence number that let Kafka ignore duplicate retries.

for a middle

Explain enable.idempotence and that the broker tracks last sequence per producer/partition.

for a senior

Detail PID/epoch/baseSequence semantics, duplicate vs out-of-order handling, and the per-partition/per-session scope.

for a principal

Distinguish idempotence from transactions/EOS, reason about PID expiry, fencing, in-flight ordering, and failure modes across broker restarts.

## The duplicate problem Producers retry on transient failures (e.g. a network blip after the broker wrote the batch but before the ack arrived). Without protection, the retry writes the batch **again**, creating duplicates. Kafka's **idempotent producer** (`enable.idempotence=true`, default since 3.0) solves this using three batch-header fields. ## The three fields 1. **producerId (PID)** — a unique long the broker assigns when the producer calls **InitProducerId**. It identifies the producer session. 2. **producerEpoch** — a short that the broker bumps to **fence** stale producers. If a new instance with the same transactional id (or a recovered session) gets a higher epoch, batches from the old epoch are rejected as **zombies** (fencing). 3. **baseSequence** — the sequence number of the **first** record in the batch. Sequences are **per (PID, epoch, partition)** and strictly increasing; record i in the batch has sequence `baseSequence + i`. ## How the broker deduplicates For each (PID, partition) the broker tracks the **last 5 accepted batches' sequence ranges** (it keeps recent state, not just one number, to validate up to 5 in-flight requests). On append it compares the incoming **baseSequence** to what it last accepted: - **baseSequence == lastAcceptedSeq + 1** → in order → **accept**. - **baseSequence <= lastAcceptedSeq** (already written) → it's a **retry/duplicate** → the broker **returns success to the producer but does not write the records again** (internally `DUPLICATE_SEQUENCE_NUMBER`). The producer is happy; the log is clean. - **baseSequence > lastAcceptedSeq + 1** (a gap) → **OUT_OF_ORDER_SEQUENCE_NUMBER** → the producer must resend the missing batches; this preserves no-reordering. ## Why max.in.flight = 5 is still safe Before idempotence, `max.in.flight.requests.per.connection > 1` with retries could reorder batches. Because the broker now validates sequence contiguity (rejecting gaps), the idempotent producer can keep **up to 5** in-flight requests without losing ordering — the broker enforces order via sequences and the producer re-queues on OUT_OF_ORDER. ## Guarantees and limits - **Scope:** no duplicates and no reordering **per partition**, for the lifetime of a (PID, epoch) session. PID state can be lost (broker restart beyond snapshot retention, `producer.id.expiration.ms`, or topic `retention`), after which dedup history resets. - **Not end-to-end EOS:** idempotence prevents *log* duplicates from retries; it does **not** make multi-partition writes atomic or make consume-process-produce exactly-once. That requires **transactions** (transactional.id, `__transaction_state`, read_committed), which reuse the same PID/epoch fields plus control batches. - **Fencing:** epoch bumps stop zombie producers (e.g. a hung instance that resumes after a replacement took over). ## Where it lives on disk PID, epoch, and baseSequence are all in the **uncompressed batch header**, so the broker validates them without decompressing the records — the same property that enables zero-copy and cheap append.

  • What does the broker return if a batch arrives with a sequence gap, and why?
    OUT_OF_ORDER_SEQUENCE_NUMBER. A gap means an earlier batch is missing, so accepting this one would reorder/lose data; the producer must resend the missing sequences. This is how no-reordering is preserved even with up to 5 in-flight requests.
  • Does idempotent production give exactly-once end-to-end? What's missing?
    No. It only removes duplicate writes from producer retries, per partition, per PID session. End-to-end exactly-once needs transactions (transactional.id, atomic multi-partition commits, read_committed consumers).
  • What is the producerEpoch for?
    Fencing. The broker bumps the epoch so that a stale/zombie producer (same identity, older epoch) is rejected, preventing it from writing after a newer instance took over.

saying these in an interview costs you the question

  • Claiming idempotence gives end-to-end exactly-once (it only dedups log writes per partition).
  • Saying duplicates cause an error to the client — a duplicate retry is acked as success and silently dropped.
  • Thinking max.in.flight must be 1 with idempotence (up to 5 is safe).
  • Ignoring that PID/dedup state can expire (producer.id.expiration.ms), resetting history.

context