skip to content

Change Data Capture with Debezium

Log-based change data capture with Debezium: initial snapshots, the before/after envelope, delete tombstones, and schema changes. A very common interview topic, since CDC is the standard way to get a database into Kafka.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

6

What is Change Data Capture (CDC) with Debezium, and why is log-based CDC preferred over query-based polling?

level: juniorimportance: must knowfreq 78%

answer

  1. reads the transaction log, not the table
  2. MySQL binlog / Postgres WAL
  3. one event per insert/update/delete
  4. polling misses DELETEs
  5. Kafka Connect source connector

basics

~20 s

Debezium is a Kafka Connect source connector that reads a database's transaction log (MySQL binlog, Postgres WAL) and streams every row insert, update, and delete to Kafka as events. Log-based CDC catches all changes, including deletes, with low impact on the database.

solid answer

~40 s

Change Data Capture means capturing row-level changes (inserts/updates/deletes) from a database and emitting them as a stream. Debezium is a set of Kafka Connect source connectors that do this by reading the database's commit log directly: MySQL's binary log (binlog), Postgres's Write-Ahead Log (WAL), etc. It runs inside Connect workers, deserializes log entries, and produces one Kafka message per change. Log-based CDC is preferred over query-based polling (SELECT ... WHERE updated_at > last_run) because the log records every committed change in order — so deletes and intermediate updates are never missed, there is no need for an updated_at column, ordering is preserved, and the database is barely impacted since reading the log is cheap compared to repeated full-table scans.

go deeper

for a junior

Know that Debezium streams row changes from the DB log into Kafka and that it catches deletes which polling misses.

for a middle

Explain the binlog/WAL mechanism and the concrete advantages over polling (deletes, ordering, no updated_at column, low load).

for a senior

Discuss the required DB config (ROW binlog format, wal_level=logical, replication slots) and operational trade-offs of log-based capture.

for a principal

Reason about where CDC fits in an event-driven architecture, downstream contract stability, and when polling or outbox patterns are preferable.

**Change Data Capture (CDC)** is the practice of detecting and capturing changes made to data in a database so other systems can react to them. Instead of periodically asking 'what changed?', CDC delivers a continuous stream of change events. **Debezium** is an open-source CDC platform built on top of **Kafka Connect** (the pluggable integration framework that ships with Apache Kafka). Debezium provides *source connectors* — one per database engine (MySQL, PostgreSQL, MongoDB, SQL Server, Oracle, Db2, Cassandra). A source connector pulls data *into* Kafka. Each Debezium connector runs as one or more *tasks* inside Kafka Connect worker JVMs. **How log-based CDC works:** Relational databases maintain a durable, ordered transaction log to guarantee crash recovery and replication: - **MySQL** writes a **binary log (binlog)** in `ROW` format — every committed row change is recorded. - **PostgreSQL** writes a **Write-Ahead Log (WAL)**; Debezium consumes it via *logical replication* (a logical decoding plugin like `pgoutput` or `wal2json`). Debezium connects as a replication client, reads these log entries as the database commits them, decodes each one into a structured change event, and publishes it to a Kafka topic (one topic per table by default). Because the log is the same mechanism the database uses for its own replicas, Debezium sees an exact, ordered record of every committed change. **Why log-based beats query-based polling:** - **Query-based polling** runs something like `SELECT * FROM t WHERE updated_at > :last`. It requires an `updated_at` column, **misses DELETEs entirely** (a deleted row no longer matches any query), can miss intermediate states (two updates between polls collapse to one), adds query load to the primary, and can lag. - **Log-based CDC** captures *every* committed change including deletes, preserves commit order, needs no extra columns, and imposes minimal load because reading the log is what replicas already do. **Trade-off:** log-based CDC requires database privileges and config (binlog enabled in ROW format with `binlog_row_image=FULL`; in Postgres `wal_level=logical` plus a replication slot and publication). It is operationally heavier to set up but far more correct and efficient at runtime.

  • Why can't query-based polling capture deletes?
    A polling query (WHERE updated_at > last) can only return rows that still exist. A deleted row no longer matches any SELECT, so its disappearance is invisible. The transaction log, by contrast, contains an explicit delete record.
  • Does Debezium run inside the database or separately?
    Separately — it runs as a connector/task inside a Kafka Connect worker process, connecting to the database as a replication client over the network. It is not a database plugin (though Postgres needs a logical-decoding output plugin server-side).

saying these in an interview costs you the question

  • Saying Debezium polls the table with SELECT statements (it reads the commit log).
  • Claiming Debezium runs inside the database server itself.
  • Saying query-based CDC can capture deletes just as well.
  • Confusing Debezium (a source connector) with a sink that writes to a database.

context

open as a page

Describe the structure of a Debezium change event envelope (op/before/after/source) and explain how deletes and tombstones work.

level: middleimportance: must knowfreq 72%

basics

~20 s

Each Debezium event has a value with op (c/u/d/r), before (old row), after (new row), and source metadata. A delete sends an event with op='d', before populated, after null — followed by a separate tombstone message (same key, null value) so log-compacted topics can drop the key.

open as a page

How does Debezium's initial snapshot work, and how does it hand off to streaming the transaction log without missing or duplicating data?

level: seniorimportance: must knowfreq 60%

basics

~20 s

On first start, Debezium reads the current contents of the captured tables (the snapshot) and emits each row as an op='r' event, recording the log position (binlog offset / LSN) at snapshot start. It then begins streaming the transaction log from that recorded position, so every change after the snapshot is captured exactly once.

open as a page

What is a Debezium schema change topic (and schema history), and how does Debezium handle DDL/schema evolution in the source database?

level: middleimportance: should knowfreq 42%

basics

~20 s

Debezium tracks the source table structure so it can correctly decode log entries that only carry column positions/values. MySQL persists captured DDL to an internal schema history topic, and can also publish DDL to a separate schema change topic for consumers. When columns are added/changed, the event schema evolves accordingly.

open as a page

Explain how the Postgres Debezium connector consumes the WAL: logical replication, output plugins (pgoutput/wal2json), replication slots, and the operational risks of slots.

level: seniorimportance: should knowfreq 50%

basics

~20 s

The Postgres connector uses logical replication: Postgres must run with wal_level=logical. A logical decoding output plugin (pgoutput, built in, or wal2json) turns WAL records into change events, delivered through a replication slot that tracks the last LSN the connector confirmed. The big risk: if the connector is down, the slot holds WAL on disk and can fill the disk.

open as a page

What delivery guarantee does Debezium provide, how do you achieve effectively-exactly-once end to end, and what is the role of exactly-once snapshot/streaming and Kafka Connect EOS support?

level: principalimportance: should knowfreq 38%

basics

~20 s

Debezium delivers at-least-once: after a crash it may re-emit some events because offsets are committed periodically, not per-event. Effective exactly-once is achieved downstream by idempotency — each event carries the primary key, so consumers/sinks upsert by key and re-applied duplicates are harmless. Kafka Connect 3.3+ added exactly-once source support that can make source-side delivery exactly-once.

open as a page