skip to content

What delivery guarantee does Debezium provide, how do you achieve effectively-exactly-once end to end, and what is the role of exactly-once snapshot/streaming and Kafka Connect EOS support?

level: principalimportance: should knowfreq 38%

answer

  1. default = at-least-once (periodic offset commit)
  2. duplicates from crash before offset commit + snapshot/stream overlap
  3. idempotent upsert/delete by primary key = effectively exactly-once
  4. KIP-618 Connect 3.3 exactly.once.source.support + transaction.boundary
  5. incremental snapshot watermarking dedupes mid-chunk
  6. end-to-end needs read_committed + transactional consumers

basics

~20 s

Debezium delivers at-least-once: after a crash it may re-emit some events because offsets are committed periodically, not per-event. Effective exactly-once is achieved downstream by idempotency — each event carries the primary key, so consumers/sinks upsert by key and re-applied duplicates are harmless. Kafka Connect 3.3+ added exactly-once source support that can make source-side delivery exactly-once.

solid answer

~50 s

By default Debezium is **at-least-once**: it commits its source position (binlog/LSN offset) to the Connect offsets topic periodically, so a crash between emitting events and committing the offset causes those events to be **re-emitted** on restart. Duplicates are also possible across the snapshot→streaming boundary. The standard way to get **effectively exactly-once** is **idempotent consumption**: every event carries the source **primary key** as the Kafka message key, so sinks perform **upserts/deletes keyed by PK** — re-applying a duplicate yields the same final state. From **Kafka Connect 3.3 (KIP-618)** there is **exactly-once support for source connectors** (`exactly.once.source.support` on the worker + connector-defined transaction boundaries), which wraps record production and offset commits in Kafka transactions so source-side delivery becomes exactly-once. End-to-end EOS additionally requires consumers reading with `read_committed` and processing transactionally (e.g. Kafka Streams EOS). So the pragmatic design is idempotent-by-key plus, where supported, Connect EOS.

go deeper

for a junior

Know Debezium can deliver duplicates (at-least-once) and that consumers should handle them.

for a middle

Explain why duplicates occur (periodic offset commits) and that PK-keyed upserts make consumption idempotent.

for a senior

Design idempotent sinks, reason about snapshot/stream overlap, and know read_committed for transactional consumers.

for a principal

Decide between idempotent-by-key vs Connect EOS (KIP-618), define end-to-end exactly-once across snapshot, streaming, and consumers, and own the cost/latency trade-offs.

**Start from the guarantee.** Out of the box, **Debezium provides at-least-once delivery**. Here's the mechanism and why duplicates arise: - Debezium reads the log and produces records to Kafka. It records its progress (binlog file+position / Postgres LSN) as **connector offsets**, committed to the Connect `offset.storage.topic` **periodically** (`offset.flush.interval.ms`), not after every single record. - If the worker crashes after some records were published but **before** the corresponding offset was committed, on restart Debezium resumes from the *last committed* offset and **re-publishes** the records produced after it. Hence duplicates. - The **snapshot→streaming handoff** can also duplicate: a row read during the snapshot (`op:'r'`) may also appear as a streamed `u`/`d` if it changed during/after the snapshot window. **Ordering is preserved per key.** Within a table's topic-partition, events for a given primary key are in commit order (Debezium keys by PK so all changes to one row land on one partition). This ordering is what makes idempotent replay safe. **Effectively-exactly-once via idempotency (the classic answer).** Because every event carries the **primary key** as the Kafka **message key** and the full `after`/`before` state, downstream sinks can be made idempotent: - Treat `c`/`u`/`r` as **UPSERT by PK**, `d`/tombstone as **DELETE by PK**. - Re-applying the same event (or an older duplicate) converges to the same final row state — duplicates are harmless. - Many sink connectors (e.g. JDBC sink with `insert.mode=upsert`, Elasticsearch sink keyed by document id) implement exactly this. This turns at-least-once transport into **effectively-exactly-once outcome** without distributed transactions. It is the most common, most robust design. **Kafka Connect exactly-once source support (KIP-618, Connect 3.3+).** Apache Kafka added native **exactly-once semantics for source connectors**: - Enable on the worker with `exactly.once.source.support = enabled` (rolling upgrade via `preparing` first). - The connector defines **transaction boundaries** (`transaction.boundary` = `poll` / `interval` / `connector`); Connect then produces source records **and** commits source offsets within a **single Kafka transaction**, so either both happen or neither. - This eliminates source-side duplicates from the offset-commit race. Debezium supports this for connectors that opt in; it requires an idempotent/transactional producer under the hood (`enable.idempotence`, transactional id managed by Connect). - It increases latency/throughput cost and requires careful operational handling, so teams often still prefer idempotent sinks. **'Exactly-once snapshot' concerns.** The phrase usually refers to (a) not double-counting snapshot vs streaming and (b) consistent snapshot boundaries. Mechanisms: the pinned log position at snapshot start (so streaming reconciles overlap), **incremental snapshots with watermarking** (open/close window markers in the change stream let Debezium dedupe rows that change mid-chunk during an online snapshot), and idempotent downstream application. True end-to-end exactly-once for the snapshot still ultimately leans on idempotency or Connect EOS. **End-to-end EOS requires the whole chain.** Even with source EOS, a consumer can still double-process. Full exactly-once needs: - Consumers using **`isolation.level=read_committed`** so they only see committed transactional records. - Transactional/idempotent processing downstream (e.g. **Kafka Streams** with `processing.guarantee=exactly_once_v2`, or sinks that commit offsets and writes atomically / idempotently). **Practical recommendation.** Design for **at-least-once + idempotent, key-based application** as the default; layer on **Connect exactly-once source support** where the connector and version support it and you need to eliminate source duplicates; ensure consumers read committed and process idempotently/transactionally. Don't claim 'Debezium is exactly-once' unqualified — it is at-least-once by default, made effectively exactly-once by these techniques.

  • Why is Debezium at-least-once rather than exactly-once by default?
    It commits source offsets (binlog pos/LSN) to the offsets topic periodically, not atomically with each produced record. A crash after producing records but before committing the offset causes those records to be re-emitted on restart.
  • How do you turn at-least-once transport into an exactly-once outcome without distributed transactions?
    Idempotent, key-based application: every event carries the primary key, so sinks upsert on insert/update and delete on tombstone, keyed by PK. Re-applying a duplicate converges to the same final state, so duplicates are harmless.
  • What does Kafka Connect exactly-once source support (KIP-618) actually guarantee, and what's still needed end to end?
    It commits produced records and source offsets in a single Kafka transaction, removing source-side duplicates (Connect 3.3+, exactly.once.source.support + transaction.boundary). End-to-end still needs consumers using read_committed and processing transactionally/idempotently.

saying these in an interview costs you the question

  • Claiming Debezium is exactly-once out of the box (it is at-least-once by default).
  • Saying ordering across the whole topic is guaranteed (it's per-key/per-partition).
  • Believing idempotent sinks require distributed XA transactions (they just upsert by PK).
  • Thinking Connect source EOS alone gives end-to-end exactly-once without read_committed consumers.
  • Confusing 'no duplicates' with 'no reprocessing' — duplicates can be reprocessed but converge if idempotent.

context