What are the real guarantees and limitations of Kafka exactly-once semantics, and how do fencing, EOSMode, and non-Kafka side effects factor in?
answer
- EOS = idempotent producer + transactions + read_committed
- 'effectively once', not 'runs once'
- scope = within Kafka only; DB not covered
- fencing via transactional.id + epoch
- EOSMode.V2 = 1 producer/group; V1 removed in 3.x
basics
~20 sKafka EOS guarantees exactly-once only for read-process-write within Kafka: idempotent producers plus transactions make produce+offset-commit atomic. It does not cover external systems like databases or HTTP calls, so those must be idempotent. Fencing and read_committed complete the picture.
solid answer
~50 sKafka's exactly-once is scoped to Kafka: an idempotent producer removes retry duplicates within a partition, and transactions make multi-partition writes plus consumer-offset commits atomic. Combined with downstream read_committed consumers, a read-process-write pipeline is 'effectively once' — processing may re-run on abort, but observable output and offset progression happen once. Its limits: (1) it does not span external side effects — DB rows, cache writes, outbound HTTP are not rolled back and can repeat, so make them idempotent or use the outbox pattern; ChainedKafkaTransactionManager is deprecated. (2) Correctness depends on stable transactional.id fencing; a misconfigured or shared id, or a producer stuck open, breaks it and stalls read_committed consumers at the LSO. (3) EOSMode.V2 (default, broker ≥ 2.5) uses one producer per consumer group with sendOffsetsToTransaction(groupMetadata) for fencing, replacing V1's producer-per-partition which is removed in Spring Kafka 3.x. (4) It only holds if every downstream reader is read_committed.
code
java · 13 lines// Making non-Kafka side effects safe under 'effectively once' reprocessing:
@KafkaListener(topics = "payments-in", groupId = "payments")
public void onPayment(PaymentEvent e) {
// Kafka tx covers this send + the offset commit...
kafkaTemplate.send("payments-out", process(e));
// ...but NOT this DB write. On abort/reprocess it could repeat,
// so make it idempotent (upsert keyed by the event id / dedup guard):
ledger.upsertByEventId(e.id(), e.amount());
}
// EOSMode is configured on container properties (V2 is the default):
// factory.getContainerProperties().setEosMode(ContainerProperties.EOSMode.V2);go deeper
Know EOS is 'effectively once' within Kafka and doesn't cover external systems.
List the three mechanisms and the read_committed dependency.
Explain fencing, EOSMode.V2 vs V1, and idempotency for side effects.
Design cross-system consistency (outbox, TM nesting), reason about LSO/timeout tuning, id stability under autoscaling, and when to skip EOS entirely.
## What EOS actually is Kafka's 'exactly-once semantics' is a composition of three mechanisms, not a single magic flag: 1. **Idempotent producer** (`enable.idempotence=true`, implied by transactions): the broker tracks a producer id (PID) + per-partition sequence number and drops duplicate retries, so a network retry doesn't append the same record twice **to one partition**. 2. **Transactions**: atomic writes across **multiple partitions/topics** plus **consumer offset commits** — all commit or all abort. 3. **Consumer `read_committed`**: readers never see aborted or in-flight transactional records. Together these give **exactly-once for read-process-write *within Kafka***. Crucially it is **'effectively once'**: on abort/crash the input record is re-consumed and your code re-executes, but because the prior output was aborted (invisible to `read_committed`) and the offset never advanced, the *observable* result is once. ## Fencing (the safety backbone) Each `transactional.id` has a broker-tracked **epoch**. `initTransactions()` bumps the epoch; a stale ('zombie') producer with a lower epoch gets `ProducerFencedException` and cannot commit. This is what allows safe restarts. **Failure modes:** two live instances sharing a `transactional.id` fence each other; an unstable id (e.g., random per boot) loses fencing across restarts; a producer that hangs mid-transaction holds it open and stalls all `read_committed` consumers at the **Last Stable Offset** until the `transaction.timeout.ms` expires and the broker aborts it. ## EOSMode (Spring Kafka) - **`EOSMode.V2`** (default; was `BETA`): requires broker **and** client **≥ 2.5**. Uses a **single producer per consumer group**, and `sendOffsetsToTransaction` is passed the **consumer group metadata** so the broker fences based on group generation. Scales well. - **`EOSMode.V1`** (was `ALPHA`): one producer per **topic/partition/group**, needed because older brokers fenced only by `transactional.id`. It doesn't scale (producer explosion) and is **removed in Spring Kafka 3.0**. Operationally V2 means fewer producers, less memory, and fencing that survives rebalances. ## The big limitation — non-Kafka side effects A Kafka transaction spans **only** Kafka records and offsets. If your listener also writes to a database, calls an external API, or mutates a cache, those are **not** part of the transaction and are **not rolled back** on abort; on reprocessing they can repeat. Options: - **Idempotent side effects** (upserts keyed by event id; dedup tables). - **Transactional outbox**: write to a DB table in a DB transaction, relay to Kafka separately (e.g., Debezium/CDC) — moves the atomicity boundary to the DB. - **Nested transaction managers** (Kafka TM outermost, DB `@Transactional` inside, or reverse) — narrows but does not eliminate the window; `ChainedKafkaTransactionManager` is **deprecated**. There is **no true XA/2PC** between Kafka and a database; you engineer for idempotency instead. ## When to use / not use - **Use** when you need pipeline correctness and can tolerate the throughput/latency cost (extra round trips, `read_committed` LSO latency, transaction overhead). - **Avoid** for pure fire-and-forget high-throughput streams where at-least-once + idempotent consumers is simpler and cheaper. - **Watch**: `transaction.timeout.ms` vs. `max.poll.interval.ms`; keep transactions short; monitor LSO lag; ensure `transactional.id` stability across deployments and autoscaling. ## Common misconceptions - 'EOS means my code runs once' — no, it can re-run; output is once. - 'EOS covers my database write' — no, only Kafka. - 'Idempotent producer alone gives exactly-once' — no, that's per-partition dedup only; you also need transactions + read_committed. - 'read_committed is optional' — without it, aborted output is observable and the guarantee is void.
- Kafka can't do XA with a database — how do you get end-to-end consistency with a DB write?Use the transactional outbox: within a single DB transaction, write both your business rows and an outbox row; a separate relay (CDC like Debezium, or a poller) publishes the outbox to Kafka. Atomicity lives in the DB; the Kafka side becomes at-least-once and consumers dedupe. This avoids needing distributed 2PC across Kafka and the DB.
- What happens if a transactional producer hangs mid-transaction?The transaction stays open, so downstream read_committed consumers stall at the Last Stable Offset for those partitions. The broker eventually aborts it once transaction.timeout.ms elapses, releasing the LSO. This is why you keep transactions short and monitor LSO/consumer lag.
- Does the idempotent producer alone give exactly-once?No. The idempotent producer only removes duplicate retries to a single partition (PID + sequence number). Exactly-once for a pipeline also needs transactions for atomic multi-partition writes and offset commits, plus read_committed consumers downstream.
saying these in an interview costs you the question
- Claiming Kafka EOS covers database writes or supports XA/2PC with a DB.
- Saying exactly-once means the listener executes exactly once.
- Presenting the idempotent producer as sufficient for exactly-once.
- Recommending ChainedKafkaTransactionManager as the current best practice (it's deprecated).
- Ignoring transactional.id stability/fencing or the read_committed requirement.