In a Debezium MySQL connector, how does snapshot.mode=never differ from schema_only?
answer
- neither one emits your existing rows
- one records table structure first
- binlog rows carry values, not column names
- the other must replay DDL from the log
- a third, similarly named mode repairs lost history
basics
~20 sBoth skip copying existing rows. schema_only first reads the current table definitions into the connector's schema history and starts streaming from the end of the binlog. never captures no schema at all, so the connector must rebuild it by replaying DDL from the binlog.
solid answer
~50 sNeither mode emits your existing rows — that part is identical, and it is the part teams forget. The difference is schema. A MySQL binlog row event carries column *values* but no column names or types, so the Debezium MySQL connector keeps a **database schema history** built by replaying DDL. `schema_only` (renamed `no_data` in recent releases) reads the current table definitions first, records them in that history, and then begins streaming from the current end of the binlog — so decoding works immediately. `never` skips that step: the connector has no starting schema and must reconstruct it from DDL statements it finds in the binlog, which is only sound when the binlog still holds the table's entire history. In practice `schema_only`/`no_data` is what you want when you deliberately start from "now", and `never` is a narrow, expert setting.
code
properties · 6 linesconnector.class=io.debezium.connector.mysql.MySqlConnector
topic.prefix=shop
table.include.list=shop.orders
# record table structure, skip the rows, stream from now
snapshot.mode=schema_only
schema.history.internal.kafka.topic=schema-history.shopgo deeper
Recall that both modes skip the existing rows, so a consumer replaying the topic will not see untouched data. Knowing that one of them still records the table structure first is enough at this level.
Explain why a MySQL-style connector keeps a schema history at all: log row events carry values but not column names and types, so the connector decodes against recorded DDL. Then contrast what each mode records before streaming begins.
Demonstrate the operational judgment: choose the no-data start to get off the source's log retention immediately, then plan the backfill explicitly, and know which mode repairs a lost schema history without losing changes.
Own the policy across a fleet: which sources are allowed to start from now, what proves the backfill actually completed, and how you prevent a mode chosen for speed from quietly producing incomplete destination tables.
## The shared behaviour: no rows Start with what the two modes have in common, because it is the source of most incidents. Neither mode produces a single `r` (read) event. A consumer that builds a table by replaying the topic from the beginning will therefore hold only the rows that changed *after* the connector started. Every row that existed and was never touched again is invisible to that consumer forever. If your destination needs a complete table, one of these modes is only ever half the plan — the other half is a backfill, usually an ad-hoc incremental snapshot afterwards. ## Why MySQL needs a schema at all MySQL's row-based binlog records changed rows as ordered lists of column values, together with a table map event carrying types but not the full DDL semantics the connector needs to build a proper event schema. Debezium therefore maintains its own **database schema history**: a record of the DDL applied to the captured tables, kept in a dedicated Kafka topic (or a file, for Debezium Server and the embedded engine). Every binlog event is decoded against the schema as of that point in the log, which is what lets the connector emit correct events for a table whose columns changed six months ago. The same requirement applies to the other connectors that read a log carrying no schema information — SQL Server, Oracle, Db2. The Postgres connector is the notable exception: it reads the current table definitions from the database catalog when it connects, so it has no schema history topic and the `schema_only` distinction does not arise there in the same way. ## schema_only (no_data) With `schema_only`, the connector connects, reads the current structure of every captured table, writes that into the schema history, records the current binlog position, and starts streaming from there. Cost: one cheap metadata pass. Benefit: a connector that is immediately correct and starts producing within seconds of deployment even against a multi-terabyte database. This is the mode behind the standard "stream now, backfill later" pattern: bring the connector up with `no_data` so the log is being consumed from the outset, then issue ad-hoc incremental snapshots table by table to fill in history at a pace you control. ## never `never` instructs the connector to take no snapshot of any kind — not the rows, not the schema. Having no recorded schema, the connector has to derive one from DDL statements found in the log. That is only sound when the binlog retained on the server reaches back to the creation of the tables you capture. Because binlog retention is finite on every real server, that condition is rarely true, which is why `never` is documented as a mode to configure with care rather than a routine choice. The exact starting position and failure behaviour for this mode have been tightened across Debezium releases, so verify against the docs for the version you run rather than assuming. ## The recovery mode people confuse with these A third, closely named mode exists for a different failure: `schema_only_recovery` (renamed `recovery`). Use it when the schema history topic has been lost or corrupted but the connector's *offsets* are intact. It re-reads the current table structures into a fresh history and then continues streaming from the stored offset. Reaching for `schema_only`/`no_data` in that situation is a classic mistake: it would restart from the current end of the binlog and silently drop every change that occurred while the connector was down. ## Picking between them Use `initial` when you want a complete copy. Use `schema_only`/`no_data` when you knowingly want changes from now on — a new analytics feed where history will be loaded separately, a second connector added beside an existing one, or a giant source where the snapshot is impractical and incremental snapshots will do the backfill. Reserve `never` for the rare case where you truly must not run even a metadata pass and you have reasoned about where the connector will begin in the log. And whichever you pick, write down how the pre-existing rows will get to the destination, because none of these modes will do it for you. ## Operational note Because `schema_only` and `never` skip the long data-reading phase, they also skip the operational risk that goes with it: the connector begins consuming the log immediately, so the source stops accumulating unread log segments on the connector's behalf almost at once. On a big database that is often the deciding argument, independent of any preference about backfills.
- Why does the Postgres connector not have the same schema_only consideration?The Postgres connector reads table definitions from the database catalog when it connects and refreshes them as it goes, so it keeps no schema history topic. The MySQL, SQL Server, Oracle and Db2 connectors read logs that do not carry column names and types, so they must maintain a DDL history of their own.
- Your schema history topic was accidentally deleted but the connector's offsets are intact — which mode do you use?The recovery mode (`schema_only_recovery`, renamed `recovery` in newer releases). It rebuilds the history from the current table structures and then resumes streaming from the stored offset. Using `schema_only`/`no_data` instead would restart from the current end of the log and silently lose every change that happened while the connector was down.
- If you start with no_data, how do the pre-existing rows ever reach the destination?Through a separate backfill, most commonly an ad-hoc incremental snapshot triggered per table once the connector is streaming. That reads the table in chunks alongside live changes, so you control when and how fast history arrives, and you can stop it if the source is under pressure.
saying these in an interview costs you the question
- Assumes never begins streaming from right now
- Thinks schema_only backfills the existing rows
- Believes the Postgres connector needs a schema history topic too
- Uses schema_only to repair a deleted schema history topic
- Ships a mode that skips rows with no backfill plan at all