In Debezium, how do you trigger an incremental snapshot of one table without restarting the connector?
answer
- no restart, no offset surgery
- the connector reads its own instruction
- a small table with id, type and data
- one JSON payload names the tables
- chunked by primary key, alongside streaming
basics
~20 sSend an execute-snapshot signal. Configure signal.data.collection to point at a signalling table with id, type and data columns, then insert a row whose type is execute-snapshot and whose data names the tables to re-read. The connector backfills them in chunks while streaming continues.
solid answer
~40 sDebezium exposes a **signalling channel** for on-demand work. Point `signal.data.collection` at a table in the source database with `id`, `type` and `data` columns, make sure it is not excluded from the capture set, and grant the connector user write access to it. To backfill a table you insert one row: `type` = `execute-snapshot`, `data` = a JSON document listing `data-collections` and, optionally, `"type": "INCREMENTAL"`. The connector picks the signal up out of the log, then reads the named table in chunks ordered by primary key, emitting `r` events **interleaved with ongoing streaming** rather than pausing it. Related signal types let you `pause-snapshot`, `resume-snapshot` and `stop-snapshot` mid-flight. If you prefer not to write to the source, enable the Kafka channel with `signal.enabled.channels` and post the same payload to `signal.kafka.topic`.
code
sql · 7 lines-- signalling table: id VARCHAR(42), type VARCHAR(32), data VARCHAR(2048)
INSERT INTO inventory.dbz_signal (id, type, data)
VALUES (
'backfill-customers-1',
'execute-snapshot',
'{"data-collections": ["inventory.customers"], "type": "INCREMENTAL"}'
);go deeper
Know that a re-read is requested by writing a signal row rather than by restarting anything, and that the connector needs a small signalling table configured up front. Being able to name the execute-snapshot signal is the recall target.
Explain the plumbing: the signal is delivered through the transaction log, the payload names the collections, chunking follows the primary key, and events arrive as reads interleaved with live changes rather than replacing them.
Show operational control — pausing or stopping a backfill that is hurting the source, resuming after a restart from the last chunk, filtering to a subset, and choosing the Kafka channel or a read-only variant when writing to the source is forbidden.
Own the process around it: who is allowed to trigger a re-read of a production table, how the request is audited, and how downstream consumers are told that a wave of read events is a backfill rather than real activity.
## The problem this solves A classic Debezium snapshot happens once, at start-up, for the whole capture set. Real pipelines need re-reads afterwards: a table was added to `table.include.list`, a consumer bug corrupted a destination table, a column was backfilled in place by a DBA (an in-place UPDATE does produce log events, but a data-fix applied by restoring a dump may not), or the connector was started with a mode that skipped the rows. Restarting with a fresh snapshot means re-copying **everything** and stopping the live feed for hours. Incremental snapshots, triggered by a signal, solve exactly this. ## Setting up the signalling channel Create a table in the source database with three columns — `id` (a string, the caller's unique identifier for the request), `type` (the signal name) and `data` (a string holding a JSON document) — and name it in `signal.data.collection` using the qualification the connector expects (`schema.table` on Postgres, `database.table` on MySQL, `database.schema.table` on SQL Server). Two setup mistakes account for most "my signal did nothing" reports: the connector's database user cannot write to the table, and the table has been filtered out of the capture set by a `table.include.list` that does not mention it. The connector reads its own writes to that table out of the log, so it must be captured. ## Sending the signal A backfill request is a single INSERT whose `type` is `execute-snapshot` and whose `data` is a JSON object with a `data-collections` array of fully-qualified table names. The `type` field inside the payload selects the flavour of snapshot; `INCREMENTAL` is the one that runs alongside streaming. Newer releases add a blocking flavour that pauses streaming for the duration, for cases where interleaving is not acceptable. Once accepted, the connector reads the table in chunks ordered by primary key. Each chunk is bounded by window markers so that rows also changing in the log are reconciled rather than duplicated. The emitted events look like snapshot reads — `op` of `r` — with the source metadata flagging that they came from an incremental snapshot, so a consumer can distinguish them from the original snapshot if it needs to. ## Controlling a running snapshot The same channel carries `stop-snapshot` to abandon a backfill outright and `pause-snapshot` / `resume-snapshot` to suspend one that is competing with peak-hour traffic. Because progress is recorded in the connector's offsets, a connector restart in the middle of an incremental snapshot resumes from the last completed chunk rather than starting the table again — a sharp contrast with the classic snapshot phase. ## Requirements and limits Chunking is done by primary key, so the table needs one; where it does not have a usable one you can nominate a `surrogate-key` column in the signal payload. The payload also accepts a filter so you can re-read part of a table rather than all of it — the field is named `additional-conditions` in recent 2.x releases and was a single `additional-condition` string earlier, so check the version you run. Chunk size is governed by `incremental.snapshot.chunk.size`, whose default is deliberately small; raising it reduces the number of window round-trips at the cost of a larger in-memory buffer per chunk. ## The Kafka signal channel Many teams are not allowed to write to a production source database, and some sources are read-only replicas. Setting `signal.enabled.channels` to include `kafka` and configuring `signal.kafka.topic` lets you publish the same payload as a Kafka record instead. The record's key must match the connector's `topic.prefix`, which is how a connector picks out the signals addressed to it from a shared topic. The MySQL connector additionally supports a read-only incremental snapshot mode that avoids writing window markers to the source at all, using log positions instead — that is the combination to reach for when the source must stay untouched. ## What to say in an interview The strong answer is not just "insert a row in the signal table". It is: the signal is delivered *through the log*, which is why the signalling table must be captured; the backfill runs *alongside* streaming rather than replacing it; progress is checkpointed so it survives restarts; and it can be paused or stopped when the source is under load. Those four properties are why incremental snapshots replaced "delete the offsets and re-snapshot everything" as the standard answer to "we need that table again".
- The signal row was inserted but nothing happened — what do you check first?Two things, in order: that the signalling table is actually captured (if `table.include.list` is set and does not name it, the connector never sees the insert, because signals arrive through the log), and that the connector's database user could write the row and read the table. After that, check that the JSON in the data column is well formed and that the table names inside it are qualified the way the connector expects.
- How do you back-fill only part of a table rather than all of it?Include a filter condition in the signal payload alongside the data-collections list. Recent 2.x releases take an `additional-conditions` array pairing each data collection with a filter expression; earlier releases took a single `additional-condition` string. Verify the shape for your version, because a malformed payload is silently ignored rather than rejected loudly.
- What if you are not allowed to write to the source database?Enable the Kafka signal channel via `signal.enabled.channels` and publish the payload to `signal.kafka.topic`, keying the record with the connector's `topic.prefix` so the right connector claims it. On MySQL, the read-only incremental snapshot mode goes further and avoids writing window markers to the source entirely, using log positions to bound each chunk.
saying these in an interview costs you the question
- Says you must delete offsets and re-snapshot everything
- Forgets the signalling table must be in the capture set
- Thinks the incremental snapshot pauses streaming
- Assumes a connector restart abandons the backfill
- Believes any table can be chunked without a primary key