skip to content

A team enabled exactly_once_v2 in a Streams app but a downstream service still sees duplicate/phantom records. What's the likely cause and fix?

level: middleimportance: must knowfreq 50%

answer

  1. EOS controls the writer, not arbitrary readers
  2. default isolation.level = read_uncommitted
  3. fix: downstream read_committed
  4. read_committed reads only up to LSO
  5. Streams' own consumers already read_committed

basics

~10 s

The downstream consumer is almost certainly reading with isolation.level=read_uncommitted (the default), so it sees records from aborted transactions and uncommitted writes. Fix: set isolation.level=read_committed on the downstream consumer so it only reads committed output.

solid answer

~40 s

EOS only controls how the Streams app writes — atomically, with abort markers for failed transactions. It does NOT change how an arbitrary downstream consumer reads. A plain consumer defaults to isolation.level=read_uncommitted, meaning it reads every record on the partition, including ones from transactions that were later aborted (phantom/duplicate-looking records) and records before their commit marker. The fix is to set the downstream consumer's isolation.level=read_committed, which makes it skip aborted records and only read up to the Last Stable Offset (committed data). This is exactly why Kafka Streams sets read_committed on its own internal consumers automatically — but a separate microservice or Connect sink consuming the output topic is your responsibility. Note read_committed adds some latency (must wait for commit markers / LSO).

go deeper

for a junior

Know the downstream consumer must be set to read_committed to avoid seeing aborted records.

for a middle

Explain why aborted-transaction records physically exist and that isolation.level default is read_uncommitted; diagnose the duplicate symptom.

for a senior

Discuss LSO-driven latency, Connect overrides, and that end-to-end exactly-once also needs an idempotent sink.

for a principal

Define org conventions: enforce read_committed on all consumers of EOS topics and idempotent/transactional terminal sinks.

## The mental model gap People assume turning on EOS in the producing app makes the output topic ‘exactly once’ for everyone. It doesn't. EOS guarantees the **writer** behaves atomically: committed transactions land, aborted ones are marked aborted. But a topic partition physically contains the data from aborted transactions until log compaction/retention — they're just flagged. Whether a **reader** sees them depends entirely on the reader's `isolation.level`. ## isolation.level - **read_uncommitted (default)**: the consumer returns ALL records up to the high watermark, including records from open or aborted transactions. So it sees ‘phantom’ records that were never truly committed — looking like duplicates or spurious data. - **read_committed**: the consumer only returns records from **committed** transactions and only up to the **Last Stable Offset (LSO)** — the offset before the earliest still-open transaction. Aborted records are filtered out using the transaction markers and the aborted-transaction index. This is what you need downstream. ## Why the symptom looks like duplicates When a Streams task aborts a transaction (crash, rebalance, fenced producer) and then reprocesses, both the aborted attempt's records AND the successful retry's records are physically on the partition. A read_uncommitted consumer reads both → apparent duplicates. read_committed drops the aborted set. ## The fix Set on every downstream consumer of the EOS output topic: ``` isolation.level=read_committed ``` For Kafka Connect sinks, set `consumer.override.isolation.level=read_committed` (or the worker-level consumer config). Streams’ own internal consumers are already read_committed under EOS — only **external** consumers need this. ## Tradeoffs / caveats - **Latency**: read_committed can only advance to the LSO, so a long open upstream transaction delays visibility and can make consumer lag look high. - **Not retroactive**: changing isolation.level doesn't rewrite already-consumed data; you may need to reset/reprocess. - **Still need EOS upstream**: read_committed alone doesn't help if the producer wasn't transactional — there'd be no markers to filter on. - **End-to-end**: true exactly-once delivery to an external system also needs that system to be idempotent or to participate transactionally; Kafka can't make a non-transactional sink exactly-once by itself. ## Quick diagnosis checklist 1. Is the producing app actually on exactly_once_v2 and committing (not silently falling back)? 2. Is the downstream consumer read_committed? (most common miss) 3. Is the downstream sink idempotent for the final delivery?

  • Why doesn't Kafka Streams' EOS automatically fix the downstream consumer?
    EOS configures only the Streams app's own producer/consumers. An external microservice or Connect sink is a separate consumer with its own isolation.level, defaulting to read_uncommitted; you must set read_committed there.
  • What is the latency cost of read_committed?
    It can only read up to the Last Stable Offset, so an open upstream transaction delays visibility and can inflate reported consumer lag until that transaction commits.

saying these in an interview costs you the question

  • Assuming enabling EOS on the producer makes the topic exactly-once for all readers.
  • Not knowing read_uncommitted is the consumer default.
  • Thinking read_committed alone (without transactional upstream) gives EOS.
  • Forgetting Connect sinks need consumer.override.isolation.level=read_committed.

context