Why does a Debezium MySQL connector keep its own schema history topic, and what breaks if that topic is lost?
answer
- binlog rows carry values, not column names
- structure must match the log position, not today
- the connector keeps private state in Kafka
- compaction on that topic is corruption
- Postgres needs none of this
basics
~20 sMySQL's binlog carries row values without column names, so the connector stores every captured DDL statement in a private schema history topic to decode rows at any log position. Lose that topic and the connector cannot start until its history is rebuilt.
solid answer
~50 sRow events in the MySQL binlog are positional — values and a table id, no column names or types. To turn a row event at some binlog position into a typed change event, the connector must know the table's structure *as of that position*, so it replays every DDL statement it has seen into an in-memory model and persists that stream to `schema.history.internal.kafka.topic`. On restart it reads that topic from the beginning and rebuilds the model before streaming resumes. That topic is the connector's private state: single partition, no compaction, effectively infinite retention, one connector per topic. If it is deleted, compacted away or shared, startup fails with a complaint that the schema for some table is not known. Recovery means rebuilding history from the current DDL — a schema-history recovery snapshot mode — which is only safe if the schema has not changed since the stored offset; otherwise you re-snapshot. Postgres connectors need none of this because logical decoding delivers relation metadata with each change.
code
properties · 5 linestopic.prefix=inventory
schema.history.internal.kafka.bootstrap.servers=kafka:9092
schema.history.internal.kafka.topic=schema-changes.inventory
schema.history.internal.store.only.captured.tables.ddl=true
include.schema.changes=truego deeper
Recall that some Debezium connectors need a dedicated internal Kafka topic to remember table structures, and that it is connector state rather than data anyone downstream should read.
Explain the mechanism: binlog row events are positional, so decoding a row requires the table structure as of that log position, which is why captured DDL is replayed from the history topic on every restart.
Be ready to run the incident. Recognise the unknown-schema startup failure, decide between a schema-history recovery restart and a full re-snapshot, and know that source log retention versus lag decides which option you still have.
Own the platform policy that prevents it: exclude connector state topics from cluster-wide retention and compaction automation, fund the monitoring, and make idempotent keyed sinks a standard so a forced re-snapshot is survivable.
## Two topics that sound alike and are not Candidates conflate them constantly, so separate them first. The **schema history topic** is internal. Its name comes from `schema.history.internal.kafka.topic`, it is written and read only by one connector, and it exists so the connector can reconstruct table structures. Nobody else should consume it, and its contents are an implementation detail. The **schema change topic** is public. It is named after `topic.prefix`, it is controlled by `include.schema.changes`, and it publishes DDL change events so downstream systems can react to a source schema change — auto-evolving a target table, alerting a data team, gating a deployment. Deleting it costs you notifications; deleting the internal one costs you the connector. ## Why MySQL needs history at all A MySQL binlog row event contains a table map and packed values. It does not contain column names, and its type information is minimal. Reading the current `information_schema` is not sufficient either, because the connector may be replaying binlog entries written days ago, before three `ALTER TABLE` statements landed. Decoding a row written under the old structure with today's structure produces garbage or an error. So the connector maintains an in-memory model of every captured table, applies each DDL statement it encounters in the binlog to that model, and appends the statement plus the binlog position to the history topic. On startup it replays the topic up to its stored offset and is then able to decode any row event from that point forward. The same design appears in the SQL Server, Oracle and Db2 connectors for the same reason. Postgres is different. Logical decoding hands the connector relation metadata alongside the changes, and the connector refreshes its schema from the catalog when a relation changes, so there is no history topic in the Postgres connector's configuration at all. Being able to explain *why* the property is absent there is a good senior signal. ## The topic's operational requirements One partition, so the DDL stream has a total order. No compaction and no retention-based deletion, because history is only meaningful read from the beginning — a compacted history is a corrupted history, and this is the single most common way teams destroy it, by pointing the connector at a cluster whose default topic policy compacts or expires everything. Exactly one connector per history topic; two connectors sharing one interleave two unrelated DDL streams. And it must be backed up or reproducible as part of your disaster-recovery plan, because it is genuinely connector state, not derived data. `schema.history.internal.store.only.captured.tables.ddl` limits what gets stored to the tables you actually capture. It keeps the topic small and startup fast, at the cost that if you later widen the include list to a table whose DDL was never stored, the connector has no history for it and needs a fresh snapshot. ## What failure looks like The connector task fails on start, typically with a message that it encountered a change event for a table whose schema is not known to this connector, or that the schema history topic is missing or empty. Streaming does not resume; lag grows for as long as the outage lasts, and on MySQL the retained binlog is finite — so a long outage can end with the required binlog position purged from the server, which turns a recoverable incident into a mandatory re-snapshot. ## Recovery paths If the source schema has not changed since the connector's stored offset, you can rebuild history from the database's current DDL using the connector's schema-history recovery snapshot mode — named `schema_only_recovery` in earlier 2.x and renamed in later releases, so confirm the value for your version. The connector reads today's structure, writes it as the history baseline, and resumes streaming from the stored offset. The safety condition is real and load-bearing: if any DDL landed between the offset and now, the rebuilt history is wrong for the rows in between, and you will decode them against the wrong structure. If the schema did change, or the offsets are also gone, the honest answer is a fresh snapshot: drop the connector's offsets and history, re-snapshot, and accept the duplicate rows downstream — which is why sinks in a CDC pipeline should be idempotent, keyed upserts rather than blind inserts. ## What a senior candidate adds Monitor it: alert on connector task failure, on the history topic's configuration drifting to compacted, and on binlog retention versus current lag, because the last of these is what decides whether an outage is a restart or a re-snapshot. And keep the history topic out of any cluster-wide policy automation that rewrites retention on unlabelled topics.
- How is the schema history topic different from the schema change topic?The history topic is internal connector state: single partition, uncompacted, read only by the connector to rebuild table structures on restart. The schema change topic is a public feed of DDL change events, named after the connector's topic prefix and toggled by include.schema.changes, meant for consumers that need to react to a source schema change.
- Why does the Postgres connector have no equivalent property?Logical decoding supplies relation metadata with the change stream, and the connector refreshes table structure from the catalog when a relation changes. There is no need to replay historical DDL to decode a change, so the connector keeps no DDL history at all. The MySQL, SQL Server, Oracle and Db2 connectors all do need it.
- When is a schema-history recovery restart unsafe?When any DDL was applied between the connector's stored offset and now. Recovery seeds history from the database's current structure, so change events written under the older structure would be decoded against the newer one. In that case you drop offsets and history and take a fresh snapshot instead.
- Why does source log retention decide how bad this outage is?While the connector is down its stored position stays put while the server keeps writing and eventually purging its log. If the required position ages out of retention before you recover, resuming is impossible and a full re-snapshot is the only option. Alerting on retention headroom versus current lag is what buys you the time.
saying these in an interview costs you the question
- Confuses the internal history topic with the public schema change topic
- Says the connector can just read information_schema at startup
- Thinks the history topic is safe to compact or expire like any other
- Lets two connectors share one schema history topic
- Assumes recovery is always safe regardless of DDL since the stored offset