skip to content

In Debezium, how do the snapshot.mode values initial and initial_only differ?

level: juniorimportance: must knowfreq 72%

answer

  1. a start-up decision, not a runtime one
  2. one copies then follows the log
  3. the other copies, then stops
  4. offsets present means no snapshot at all

basics

~20 s

Both read the full contents of the captured tables on first start. With initial the connector then switches to streaming the transaction log and follows it forever; with initial_only it shuts down once the snapshot finishes and never streams.

solid answer

~50 s

`snapshot.mode` decides what a Debezium connector does at start-up, before it begins reading the log. `initial` — the default on the common connectors — takes a consistent snapshot of every captured table, emitting one `r` (read) event per existing row, then hands over to streaming and follows the transaction log indefinitely. `initial_only` performs exactly the same snapshot and then stops the connector; the streaming phase never runs. Crucially, the mode is only consulted when the connector has **no stored offsets**. On any later restart where offsets exist, an `initial` connector skips the snapshot entirely and resumes streaming from the recorded position — which is why restarting a connector does not re-copy your tables. `initial_only` is the one-off bulk-copy mode: use it when you want Debezium's snapshot machinery to seed a target once, without leaving an ongoing CDC feed behind.

code

properties · 5 lines
properties
name=inventory-bulk-copy
connector.class=io.debezium.connector.postgresql.PostgresConnector
topic.prefix=inventory
table.include.list=public.customers,public.orders
snapshot.mode=initial_only

go deeper

for a junior

Be ready to name the mode you would set for a normal always-on pipeline and say what each phase does: read the existing rows, then follow the log. Knowing that snapshot rows carry the read operation is a nice extra.

for a middle

Explain the mechanics: the mode is only consulted when no offsets exist, offsets are keyed by the connector's logical server prefix, and that is why restarts do not re-copy tables. Be able to place always, never and the schema-only modes alongside the two defaults.

for a senior

Show that you plan for the snapshot phase operationally — how long it will run, that the source retains log segments throughout, and what an interrupted snapshot does to a downstream sink that is not idempotent.

for a principal

Own the decision of whether the bulk load belongs in the CDC tool at all. Argue when a one-shot copy plus a log-only connector beats one long snapshot, and what that choice means for cutover windows and source-database risk.

## What snapshot.mode controls A Debezium connector has two possible phases: a **snapshot** phase, in which it reads existing rows out of the source tables with ordinary SELECT queries, and a **streaming** phase, in which it reads the database's transaction log (the MySQL binlog, the Postgres WAL through a replication slot, the SQL Server change tables, and so on). The `snapshot.mode` property decides which of those phases run and when. It is a **start-up decision**, evaluated when the connector task begins — not something that changes behaviour while the connector is running. Snapshot rows are emitted as change events like any other, but with the operation field `op` set to `r` for *read*, distinguishing them from `c` (create), `u` (update) and `d` (delete). The `source` block of those events also carries a `snapshot` marker so a consumer can tell a backfilled row from a live change. ## initial With `snapshot.mode=initial`, the connector, on a first start with no recorded offsets, snapshots every table in its capture set and then transitions to streaming from the log position it recorded at snapshot time. This is the default on most connectors and the mode almost every production pipeline runs: the destination ends up with a complete copy of the table plus every subsequent change. ## initial_only `initial_only` runs the identical snapshot and then **stops**. The connector task completes and does not open the log reader. Nothing further is emitted, no matter how much the source changes afterwards. It exists for one-shot copies: seeding a new warehouse table, exporting a dataset in Debezium's event format for a migration, or producing a reference copy that some other process will keep current. ## Offsets decide, not the config The most common misunderstanding is that `initial` means "snapshot on every start". It does not. Kafka Connect persists a source connector's offsets, and Debezium's source partition is keyed by the connector's `topic.prefix` (the logical name for the server). If offsets exist for that key, the connector skips straight to streaming from the stored position. Consequences that follow directly from this: - Restarting, rebalancing, or bouncing a worker does not re-copy your tables. - Deleting and recreating the connector with the **same name and the same `topic.prefix`** finds the old offsets and still skips the snapshot. - Changing `snapshot.mode` on a connector that already has offsets changes nothing; to force a fresh snapshot you must remove the stored offsets or use a mode that snapshots unconditionally. ## Where the other modes fit Once you understand the two above, the rest of the family reads easily. `always` snapshots on **every** start regardless of offsets — useful in tests, hazardous on large tables. `never` skips the snapshot entirely and goes straight to the log, so pre-existing rows are never emitted. `schema_only` (renamed `no_data` in recent releases) reads the table definitions but not the rows. `when_needed` snapshots whenever the connector decides it must — no offsets, or a stored log position that the server has already purged. `schema_only_recovery` (renamed `recovery`) rebuilds a lost schema history topic without re-reading data. ## Choosing between them Pick `initial` for any pipeline that must stay current — that is nearly all of them. Pick `initial_only` when the destination genuinely wants a point-in-time copy and nothing more, or when you deliberately want to run the copy under supervision and then start a *separate* streaming connector with a mode that skips the snapshot. Bear in mind the operational cost of the snapshot phase in either case: while it runs, the connector is not consuming the log, so the source keeps retaining log segments the connector has not yet read. ## Common mistakes Teams frequently set `initial_only`, see the connector go into a completed state, and file a bug that Debezium "stopped working" — it did exactly what was asked. The mirror-image mistake is expecting `initial` to re-copy a table after someone truncated the destination; it will not, because the offsets are still there. And an interrupted snapshot is worth planning for: on most connectors a classic snapshot that dies part-way begins again from the start on the next attempt (some newer connector versions can resume — check the one you run), so the partially emitted `r` events are re-sent. A sink that upserts on the primary key absorbs that harmlessly; an append-only sink does not.

  • If you delete a Debezium connector and recreate it with the same configuration, does it snapshot again?
    Usually not. Kafka Connect keeps source offsets keyed by connector name and Debezium's source partition, which is derived from `topic.prefix`. Recreate the connector with the same name and the same prefix and it finds those offsets and resumes streaming. Change the prefix (or clear the offsets) and it behaves like a first start, snapshotting again.
  • What happens if an initial snapshot dies halfway through?
    On most connectors the snapshot has not produced a committed streaming offset, so the next start begins the snapshot from the beginning and re-emits the rows already sent. Idempotent, upsert-style sinks absorb the duplicates; append-only sinks accumulate them. Some newer connector versions can resume an interrupted snapshot, so confirm the behaviour for the connector and version you run.
  • How do you force a Debezium connector to re-snapshot a table that it has already streamed for months?
    Do not fight the offsets. Send an ad-hoc incremental snapshot signal for that table: the connector re-reads it chunk by chunk while streaming continues, emitting `r` events interleaved with live changes. Wiping offsets and re-running an `initial` snapshot is the blunt alternative and re-copies everything the connector captures, not just the one table.

initial is photographing the room and then leaving the camera running; initial_only is taking the photograph and packing the camera away.

saying these in an interview costs you the question

  • Thinks initial re-snapshots on every connector restart
  • Believes initial_only keeps streaming once the snapshot completes
  • Assumes editing snapshot.mode alone forces a fresh snapshot
  • Thinks snapshot rows arrive as create events rather than read events
  • Says a failed snapshot always resumes exactly where it stopped

context