skip to content

Debezium

The de-facto open-source CDC platform: a family of source connectors that turn database logs into change events, usually run on Kafka Connect but also embeddable or deployable standalone. If a job description says CDC, this is the tool most likely on the other side of the interview.

on this pageshow

explore

questions

18

In Debezium, how do the snapshot.mode values initial and initial_only differ?

level: juniorimportance: must knowfreq 72%

answer

  1. a start-up decision, not a runtime one
  2. one copies then follows the log
  3. the other copies, then stops
  4. offsets present means no snapshot at all

basics

~20 s

Both read the full contents of the captured tables on first start. With initial the connector then switches to streaming the transaction log and follows it forever; with initial_only it shuts down once the snapshot finishes and never streams.

solid answer

~50 s

`snapshot.mode` decides what a Debezium connector does at start-up, before it begins reading the log. `initial` — the default on the common connectors — takes a consistent snapshot of every captured table, emitting one `r` (read) event per existing row, then hands over to streaming and follows the transaction log indefinitely. `initial_only` performs exactly the same snapshot and then stops the connector; the streaming phase never runs. Crucially, the mode is only consulted when the connector has **no stored offsets**. On any later restart where offsets exist, an `initial` connector skips the snapshot entirely and resumes streaming from the recorded position — which is why restarting a connector does not re-copy your tables. `initial_only` is the one-off bulk-copy mode: use it when you want Debezium's snapshot machinery to seed a target once, without leaving an ongoing CDC feed behind.

code

properties · 5 lines
properties
name=inventory-bulk-copy
connector.class=io.debezium.connector.postgresql.PostgresConnector
topic.prefix=inventory
table.include.list=public.customers,public.orders
snapshot.mode=initial_only

go deeper

for a junior

Be ready to name the mode you would set for a normal always-on pipeline and say what each phase does: read the existing rows, then follow the log. Knowing that snapshot rows carry the read operation is a nice extra.

for a middle

Explain the mechanics: the mode is only consulted when no offsets exist, offsets are keyed by the connector's logical server prefix, and that is why restarts do not re-copy tables. Be able to place always, never and the schema-only modes alongside the two defaults.

for a senior

Show that you plan for the snapshot phase operationally — how long it will run, that the source retains log segments throughout, and what an interrupted snapshot does to a downstream sink that is not idempotent.

for a principal

Own the decision of whether the bulk load belongs in the CDC tool at all. Argue when a one-shot copy plus a log-only connector beats one long snapshot, and what that choice means for cutover windows and source-database risk.

## What snapshot.mode controls A Debezium connector has two possible phases: a **snapshot** phase, in which it reads existing rows out of the source tables with ordinary SELECT queries, and a **streaming** phase, in which it reads the database's transaction log (the MySQL binlog, the Postgres WAL through a replication slot, the SQL Server change tables, and so on). The `snapshot.mode` property decides which of those phases run and when. It is a **start-up decision**, evaluated when the connector task begins — not something that changes behaviour while the connector is running. Snapshot rows are emitted as change events like any other, but with the operation field `op` set to `r` for *read*, distinguishing them from `c` (create), `u` (update) and `d` (delete). The `source` block of those events also carries a `snapshot` marker so a consumer can tell a backfilled row from a live change. ## initial With `snapshot.mode=initial`, the connector, on a first start with no recorded offsets, snapshots every table in its capture set and then transitions to streaming from the log position it recorded at snapshot time. This is the default on most connectors and the mode almost every production pipeline runs: the destination ends up with a complete copy of the table plus every subsequent change. ## initial_only `initial_only` runs the identical snapshot and then **stops**. The connector task completes and does not open the log reader. Nothing further is emitted, no matter how much the source changes afterwards. It exists for one-shot copies: seeding a new warehouse table, exporting a dataset in Debezium's event format for a migration, or producing a reference copy that some other process will keep current. ## Offsets decide, not the config The most common misunderstanding is that `initial` means "snapshot on every start". It does not. Kafka Connect persists a source connector's offsets, and Debezium's source partition is keyed by the connector's `topic.prefix` (the logical name for the server). If offsets exist for that key, the connector skips straight to streaming from the stored position. Consequences that follow directly from this: - Restarting, rebalancing, or bouncing a worker does not re-copy your tables. - Deleting and recreating the connector with the **same name and the same `topic.prefix`** finds the old offsets and still skips the snapshot. - Changing `snapshot.mode` on a connector that already has offsets changes nothing; to force a fresh snapshot you must remove the stored offsets or use a mode that snapshots unconditionally. ## Where the other modes fit Once you understand the two above, the rest of the family reads easily. `always` snapshots on **every** start regardless of offsets — useful in tests, hazardous on large tables. `never` skips the snapshot entirely and goes straight to the log, so pre-existing rows are never emitted. `schema_only` (renamed `no_data` in recent releases) reads the table definitions but not the rows. `when_needed` snapshots whenever the connector decides it must — no offsets, or a stored log position that the server has already purged. `schema_only_recovery` (renamed `recovery`) rebuilds a lost schema history topic without re-reading data. ## Choosing between them Pick `initial` for any pipeline that must stay current — that is nearly all of them. Pick `initial_only` when the destination genuinely wants a point-in-time copy and nothing more, or when you deliberately want to run the copy under supervision and then start a *separate* streaming connector with a mode that skips the snapshot. Bear in mind the operational cost of the snapshot phase in either case: while it runs, the connector is not consuming the log, so the source keeps retaining log segments the connector has not yet read. ## Common mistakes Teams frequently set `initial_only`, see the connector go into a completed state, and file a bug that Debezium "stopped working" — it did exactly what was asked. The mirror-image mistake is expecting `initial` to re-copy a table after someone truncated the destination; it will not, because the offsets are still there. And an interrupted snapshot is worth planning for: on most connectors a classic snapshot that dies part-way begins again from the start on the next attempt (some newer connector versions can resume — check the one you run), so the partially emitted `r` events are re-sent. A sink that upserts on the primary key absorbs that harmlessly; an append-only sink does not.

  • If you delete a Debezium connector and recreate it with the same configuration, does it snapshot again?
    Usually not. Kafka Connect keeps source offsets keyed by connector name and Debezium's source partition, which is derived from `topic.prefix`. Recreate the connector with the same name and the same prefix and it finds those offsets and resumes streaming. Change the prefix (or clear the offsets) and it behaves like a first start, snapshotting again.
  • What happens if an initial snapshot dies halfway through?
    On most connectors the snapshot has not produced a committed streaming offset, so the next start begins the snapshot from the beginning and re-emits the rows already sent. Idempotent, upsert-style sinks absorb the duplicates; append-only sinks accumulate them. Some newer connector versions can resume an interrupted snapshot, so confirm the behaviour for the connector and version you run.
  • How do you force a Debezium connector to re-snapshot a table that it has already streamed for months?
    Do not fight the offsets. Send an ad-hoc incremental snapshot signal for that table: the connector re-reads it chunk by chunk while streaming continues, emitting `r` events interleaved with live changes. Wiping offsets and re-running an `initial` snapshot is the blunt alternative and re-copies everything the connector captures, not just the one table.

initial is photographing the room and then leaving the camera running; initial_only is taking the photograph and packing the camera away.

saying these in an interview costs you the question

  • Thinks initial re-snapshots on every connector restart
  • Believes initial_only keeps streaming once the snapshot completes
  • Assumes editing snapshot.mode alone forces a fresh snapshot
  • Thinks snapshot rows arrive as create events rather than read events
  • Says a failed snapshot always resumes exactly where it stopped

context

open as a page

How do you restrict which tables and columns a Debezium connector emits change events for?

level: middleimportance: must knowfreq 66%

basics

~20 s

Set table.include.list or table.exclude.list to regular expressions matching fully-qualified table names, and column.include.list, column.exclude.list or a column mask for fields. Debezium matches the whole identifier, and the include and exclude forms of one option are mutually exclusive.

open as a page

In Debezium, what does the required topic.prefix connector property control?

level: middleimportance: must knowfreq 62%

basics

~20 s

topic.prefix is the connector's logical source name. Debezium puts it in front of every change topic, names the schema-change, heartbeat and transaction topics with it, and records it in stored state, so it must be unique per connector.

open as a page

What does Debezium's ExtractNewRecordState SMT do to a change event, and how does it treat deletes?

level: middleimportance: must knowfreq 68%

basics

~20 s

Debezium's ExtractNewRecordState SMT flattens the change-event envelope down to the after row image, so a sink sees an ordinary record instead of before/after/op. Deletes arrive with a null after image and are dropped by default.

open as a page

In Debezium, how do you trigger an incremental snapshot of one table without restarting the connector?

level: middleimportance: must knowfreq 58%

basics

~20 s

Send an execute-snapshot signal. Configure signal.data.collection to point at a signalling table with id, type and data columns, then insert a row whose type is execute-snapshot and whose data names the tables to re-read. The connector backfills them in chunks while streaming continues.

open as a page

Why does a Debezium MySQL connector keep its own schema history topic, and what breaks if that topic is lost?

level: seniorimportance: must knowfreq 57%

basics

~20 s

MySQL's binlog carries row values without column names, so the connector stores every captured DDL statement in a private schema history topic to decode rows at any log position. Lose that topic and the connector cannot start until its history is rebuilt.

open as a page

Which Debezium connector class do you configure for a MySQL, PostgreSQL, MongoDB or SQL Server source?

level: juniorimportance: should knowfreq 50%

basics

~10 s

Debezium ships a separate connector class per source: MySqlConnector, PostgresConnector, MongoDbConnector and SqlServerConnector, all under io.debezium.connector. Each reads that engine's own change log and has its own database-side prerequisites.

open as a page

Which outbox table columns does Debezium's EventRouter SMT read, and how does it pick the topic and key?

level: middleimportance: should knowfreq 52%

basics

~20 s

Debezium's EventRouter reads an outbox row's id, aggregate type, aggregate id, event type and payload columns. The aggregate type selects the destination topic, the aggregate id becomes the message key, and the payload becomes the message value.

open as a page

In a Debezium MySQL connector, how does snapshot.mode=never differ from schema_only?

level: middleimportance: should knowfreq 48%

basics

~20 s

Both skip copying existing rows. schema_only first reads the current table definitions into the connector's schema history and starts streaming from the end of the binlog. never captures no schema at all, so the connector must rebuild it by replaying DDL from the binlog.

open as a page

Why would you set heartbeat.interval.ms on a Debezium Postgres connector capturing a low-traffic table?

level: seniorimportance: should knowfreq 52%

basics

~20 s

With no writes to captured tables Debezium emits nothing, commits no new offset, and the PostgreSQL replication slot never advances, so write-ahead log accumulates until the disk fills. Heartbeats emit periodic events so offsets and the slot keep moving.

open as a page

Two Debezium MySQL connectors share the same database.server.id — what happens on the source?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Each Debezium MySQL connector registers as a replication client identified by database.server.id. When two share an id, MySQL drops one of them; both reconnect, kick each other off again, and neither keeps up with the binlog.

open as a page

A captured table gains a nullable column; what do Debezium's Avro consumers and the schema registry see?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Debezium emits a DDL change event and then change events whose value schema carries the new optional field, so a new schema version is registered. Adding a nullable column is compatible; dropping or retyping a column is what stops the pipeline.

open as a page

Your Debezium initial snapshot has run eight hours and the source's replication slot keeps growing — what now?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The connector is snapshotting, so it is not consuming the log yet, and the slot retains every change since the snapshot began until the source disk fills. Triage disk headroom first, then restart with a mode that streams immediately and backfill the tables with incremental snapshots.

open as a page

In Debezium, what do the snapshot-window-open and snapshot-window-close signals do?

level: seniorimportance: should knowfreq 36%

basics

~20 s

They mark the boundaries of one incremental-snapshot chunk in the transaction log. Debezium buffers the chunk's rows, drops from that buffer any key it sees change between the two markers, and emits what remains — so a live change always wins over the older snapshot read.

open as a page

How would you evolve the payload schemas of Debezium outbox events without breaking downstream consumers?

level: principalimportance: should knowfreq 36%

basics

~20 s

Treat the outbox payload as a published contract: additive, optional-only changes by default; an explicit event version for breaking ones, published alongside the old until consumers migrate; and a subject naming strategy that lets several event schemas share one routed topic.

open as a page

When should Debezium's outbox router expand a JSON payload column into a typed structure?

level: middleimportance: nice to knowfreq 24%

basics

~20 s

Expand it when consumers need a real schema for the event body — typically with a registry-backed converter. Leave it as a string when the payload's shape varies between events, because the structure is inferred per record and inconsistent shapes produce unstable schemas.

open as a page

In Debezium, what does snapshot.select.statement.overrides change about a table's snapshot?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

It replaces the default SELECT that reads a table during the snapshot with a statement you supply, so you can filter rows or restrict columns. It affects the snapshot only; streaming afterwards still captures every change to that table.

open as a page

When would you run Debezium Server or the embedded engine instead of Debezium on Kafka Connect?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Debezium Server is a standalone process that streams change events to non-Kafka sinks; the embedded engine is a library running inside your own application. Choose them when there is no Kafka Connect cluster, accepting that you own restarts, scaling and offset storage.

open as a page