When does an idempotent producer throw OutOfOrderSequenceException, and what does it imply about delivery guarantees?
answer
- PID + per-partition sequence numbers
- gap in sequence -> reject (no silent hole)
- earlier batch lost / dropped via delivery.timeout
- broker producer-state expiry / unclean election
- fatal: recreate producer (new PID) or abort txn
basics
~20 sIt means the broker received a producer's batch with a sequence number that doesn't follow the last one it accepted, so a gap exists — usually because an earlier batch was permanently lost or its state expired. It signals the producer can no longer guarantee ordered, gap-free delivery for that session.
solid answer
~50 sWith enable.idempotence=true, each batch per partition carries a Producer ID (PID) and a monotonic sequence number. The broker only accepts sequence n+1 after n. OutOfOrderSequenceException is raised when the broker sees a sequence that skips ahead of (or behind) what it expects — meaning an earlier batch in the sequence was lost and never durably written, so accepting the new one would leave a gap and silently drop data. Common causes: the broker's idempotent state for that PID was evicted (producer idle longer than the broker retains state, or after unclean leader election / log truncation), or an earlier batch exhausted delivery.timeout.ms and was dropped while later batches kept their sequence numbers. It is a fatal exception for that producer instance: the ordering/dedup guarantee is broken, so the producer should be recreated (new PID) or, with transactions, the transaction aborted. It is the client's signal that exactly-once/ordered semantics could not be upheld.
go deeper
Recognize the name and that it relates to idempotent producers and ordering, not just duplicates.
Explain it as a sequence-number gap the broker rejects, caused by a lost earlier batch.
Identify causes (dropped batch via delivery.timeout, broker state expiry, unclean election) and that it's fatal, needing producer recreation.
Frame it as the system's fail-loud guardrail for the ordering/EOS contract, tie it to durability config (acks=all, ISR), and design app-level replay/recovery around it.
## Setup Idempotence (`enable.idempotence=true`) gives each producer a **Producer ID (PID)** and tags every record batch sent to a partition with a **monotonically increasing sequence number** starting at 0. The broker remembers, per `(PID, partition)`, the last sequence it durably accepted. It will only accept the **next** expected sequence; a duplicate (same sequence) is silently dropped (dedup), and a **gap** is rejected. ## When OutOfOrderSequenceException fires The broker throws/returns `OutOfOrderSequenceException` to the client when the incoming batch's sequence number does **not** equal `lastAcceptedSequence + 1` — i.e. the producer skipped ahead. Accepting it would mean some earlier batch was never written, leaving an undetected hole in the per-partition stream. The broker refuses rather than silently drop. Typical root causes: - **An earlier batch was permanently lost.** It exhausted `delivery.timeout.ms` (or `retries`) and was failed/abandoned, but later batches retained higher sequence numbers and reached the broker, so a gap appears. - **Broker-side idempotent state expired or was lost.** The broker retains producer state for a bounded time/size (`transactional.id.expiration.ms` / producer-state retention). If a producer goes idle long enough, or after **unclean leader election** or log truncation, the broker's notion of the last sequence can be reset or inconsistent, so a resumed producer's sequence looks out of order. - **Replication anomalies** under failure can leave the new leader without the latest accepted sequence. ## What it implies This is treated as a **fatal** exception for the idempotent producer instance: the invariant that guarantees no-duplicate, no-reorder, no-gap within the session has been violated and **cannot be silently recovered**. Practically: - A **non-transactional** idempotent producer must be **closed and recreated** to get a fresh PID and reset sequences (accepting that records around the gap may have been lost — the app must decide whether to resend from a durable source). - A **transactional** producer surfaces it as an abortable/fatal error; the in-flight transaction is aborted and, depending on the error, the producer may need re-initialization (`initTransactions` / new instance). ## Why the design is conservative Kafka deliberately fails loudly here. Silently accepting an out-of-order sequence would convert a **detectable** data-loss/ordering event into **silent corruption** of the partition's record stream. The exception is the system telling you: 'I can no longer prove your at-least-once-with-ordering / exactly-once contract — handle it.' ## Operational guidance - Treat it as a signal of an underlying durability incident (unclean elections, undersized `delivery.timeout.ms` causing dropped early batches, or producers idling past state retention), not just a transient blip. - Ensure `acks=all` and adequate `min.insync.replicas` to minimize loss-driven gaps. - For long-idle producers, be aware state can expire; design the app to recreate the producer and replay from the source of truth on this error.
- Is OutOfOrderSequenceException retriable by the producer like a NetworkException?No. It is fatal for that producer session — the ordering/dedup invariant is broken. The producer must be recreated (new PID) or, in a transaction, the transaction aborted; you cannot simply resend the same batch.
- Why does Kafka reject the out-of-order batch instead of just appending it?Appending would leave an undetected gap where an earlier batch should be, turning a detectable loss/ordering event into silent data corruption. Failing loudly lets the application react (e.g. replay from source).
saying these in an interview costs you the question
- Treating OutOfOrderSequenceException as a transient/retriable error to be auto-resent.
- Saying it just means duplicates — it actually signals a gap / lost earlier batch.
- Believing idempotence survives indefinite producer idleness and broker state expiry.
- Assuming you can recover by resending the same batch with the same sequence.