skip to content

EOS Scope and Limits

Where exactly-once stops: one cluster, Kafka to Kafka, no external side effects. A great senior-level question, because it punctures the claim that Kafka simply gives you exactly-once.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is the scope boundary of Kafka's exactly-once semantics (EOS), and why doesn't it automatically extend to external systems?

level: juniorimportance: must knowfreq 70%

answer

  1. EOS = cluster-local read-process-write
  2. offsets live in __consumer_offsets (a topic)
  3. external sink = not in the transaction
  4. DB/REST/other cluster = outside
  5. atomic = output records + offset commit

basics

~20 s

Kafka EOS only guarantees exactly-once within a single Kafka cluster: a read-process-write where the input topic, output topic, and consumer offsets all live in that same cluster. External databases, APIs, or other clusters are outside the transaction, so it can't cover them.

solid answer

~40 s

Kafka's exactly-once semantics are a cluster-local guarantee. The Kafka transaction atomically commits produced records to output topics AND the consumer-offset commits (sendOffsetsToTransaction) for the input topic — but only because both the data and the offsets are stored as Kafka topics in the same cluster. The transaction coordinator and idempotent producer (with a producer epoch fencing zombies) make that atomic. Anything outside the cluster — a Postgres write, a REST call, an email, an S3 put, or even a second Kafka cluster — is not enrolled in the transaction. So a consume-transform-produce loop is EOS, but consume-then-write-to-DB is not unless you add an external mechanism (idempotent sink, transactional outbox, or 2PC). Misunderstanding this leads people to assume their whole pipeline is exactly-once when only the Kafka hop is.

go deeper

for a junior

Know the headline: EOS only covers the Kafka-to-Kafka hop in one cluster; databases and external calls are not covered.

for a middle

Explain that offsets are stored as a topic, which is why the offset commit can join the transaction, and name the read-process-write pattern.

for a senior

Articulate the coordinator/PID/epoch mechanics and enumerate which pipeline shapes are and aren't EOS, including read_committed requirements downstream.

for a principal

Frame the boundary as an architectural contract: design pipelines so external effects sit behind idempotency or outbox, and teach teams not to over-trust EOS labeling.

## What EOS actually is **Delivery semantics** describe how many times a message can be processed end-to-end: - **At-most-once**: messages may be lost, never duplicated. - **At-least-once**: messages are never lost but may be duplicated (the default for most Kafka setups). - **Exactly-once**: each message is reflected in the result exactly one time — no loss, no duplication. Kafka delivers **exactly-once semantics (EOS)** for a specific pattern: **read-process-write** (also called consume-transform-produce) where the source topic, the destination topic, and the consumer's committed offsets are all inside **one Kafka cluster**. ## Why the boundary is the cluster Kafka achieves EOS with two building blocks: 1. **The idempotent producer** — each producer gets a Producer ID (PID) and a monotonically increasing sequence number per partition, so the broker deduplicates retried writes within a producer session. 2. **Transactions** — a `transactional.id` lets a producer wrap multiple partition writes *and* the consumer offset commit into one atomic unit via `producer.sendOffsetsToTransaction(...)`. A **transaction coordinator** (a broker) writes commit/abort markers to an internal `__transaction_state` topic. The critical insight: **consumer offsets are themselves stored as a Kafka topic** (`__consumer_offsets`) in the same cluster. So "advance my input position" and "write my output records" are *both* writes to topics in the same cluster, and Kafka can commit them atomically. The moment any side effect lives **outside** the cluster — a relational database, a cache, an HTTP endpoint, a file store, an email, or a *different* Kafka cluster — that write is invisible to the transaction coordinator. There is no protocol enrolling it in the commit. So Kafka cannot roll it back on abort or guarantee it happened exactly once. ## Consequences - `consume → transform → produce` (all in-cluster): **EOS achievable** with `processing.guarantee=exactly_once_v2` (Streams) or manual transactions. - `consume → write to external DB`: **not EOS by itself**. You must make the sink **idempotent** (e.g., upsert by a deterministic key) or use a **transactional outbox**. - `produce → MirrorMaker 2 → another cluster`: the second cluster's copy is **at-least-once**; cross-cluster EOS is not provided. ## Edge note Even inside the cluster, EOS requires the *whole* loop to use the same transactional producer and `isolation.level=read_committed` on downstream consumers; otherwise downstream readers see aborted/uncommitted records.

  • Why can the consumer offset commit be part of the Kafka transaction but a database write cannot?
    Consumer offsets are stored in the __consumer_offsets topic inside the same cluster, so committing them is just another Kafka write the transaction coordinator controls. A database write goes to a system the coordinator has no protocol with, so it can't be atomically committed or aborted alongside the Kafka writes.
  • Give an example of a pipeline people wrongly assume is exactly-once.
    Consume from a topic, enrich, then INSERT into Postgres. The Kafka consume is at-least-once, and on a redelivery you re-INSERT, producing duplicates unless the DB write is idempotent (upsert by key) or wrapped in an outbox pattern.

saying these in an interview costs you the question

  • Claiming Kafka EOS makes the entire end-to-end pipeline (including DBs and APIs) exactly-once
  • Saying EOS works across two Kafka clusters out of the box
  • Confusing the idempotent producer (single-session dedupe) with full transactions
  • Forgetting that downstream consumers need isolation.level=read_committed to actually see only committed data

context

open as a page

When a Kafka consumer or sink connector writes to an external system, why is at-least-once the practical floor, and how do you make the sink effectively exactly-once?

level: middleimportance: must knowfreq 65%

basics

~20 s

Because the Kafka offset commit and the external write happen in two separate systems, a crash between them causes redelivery and a repeat write — so it's at-least-once. You reach effective exactly-once by making the sink idempotent: upsert by a deterministic key or dedupe on a stored offset/ID so reprocessing produces no duplicate effect.

open as a page

Explain the transactional outbox pattern and why teams choose it over distributed 2PC to integrate a database with Kafka.

level: seniorimportance: must knowfreq 60%

basics

~20 s

The outbox pattern writes the business row and an 'outbox' event row in one local database transaction, then a separate relay (often Debezium CDC) reads the outbox and publishes to Kafka. It avoids two-phase commit (2PC) across the DB and Kafka by relying only on the local DB transaction plus at-least-once publishing with idempotency.

open as a page

Why does cross-cluster replication with MirrorMaker 2 not provide exactly-once semantics, and what are the consequences for consumers on the target cluster?

level: seniorimportance: should knowfreq 45%

basics

~20 s

MirrorMaker 2 is a Kafka Connect consume-from-source, produce-to-target pipeline across two clusters. Because the source consume and target produce span different clusters with no shared transaction, it's at-least-once: failures cause duplicate records on the target. Offsets and producer state don't carry over, so target consumers must tolerate duplicates and remapped offsets.

open as a page

As a platform architect, how would you reason about end-to-end exactly-once across a pipeline that ingests from an external source, processes in Kafka, and lands in an external sink?

level: principalimportance: should knowfreq 35%

basics

~20 s

Treat the pipeline as three hops: source→Kafka, Kafka→Kafka, Kafka→sink. Only the Kafka-internal hop can be true EOS. The source and sink boundaries are at-least-once unless you add idempotency, dedupe keys, or store-local atomicity. End-to-end exactly-once is engineered at the edges, not granted by Kafka.

open as a page