skip to content

What is a Debezium schema change topic (and schema history), and how does Debezium handle DDL/schema evolution in the source database?

level: middleimportance: should knowfreq 42%

answer

  1. binlog row events lack full column metadata -> need DDL history
  2. schema history topic = internal, for connector restart (MySQL)
  3. schema change topic = public DDL events for consumers
  4. include.schema.changes (MySQL default true)
  5. Postgres: no history topic (catalog + self-describing stream)
  6. never delete the schema history topic -> re-snapshot

basics

~20 s

Debezium tracks the source table structure so it can correctly decode log entries that only carry column positions/values. MySQL persists captured DDL to an internal schema history topic, and can also publish DDL to a separate schema change topic for consumers. When columns are added/changed, the event schema evolves accordingly.

solid answer

~50 s

Log entries (especially MySQL binlog rows) often don't carry full column names/types, so Debezium must know the table's structure at the time each change was logged. The **MySQL** connector maintains a **schema history topic** (`schema.history.internal.kafka.topic`) — an internal, compacted topic where it records every captured DDL statement and the binlog position it applied at; on restart it replays this to rebuild the schema before resuming. Separately, the **schema change topic** (named after `topic.prefix`) is an *outward-facing* topic carrying DDL change events for downstream consumers; enabled via `include.schema.changes` (default true for MySQL). When a column is added, the connector emits the new DDL to the schema change topic and subsequent data events use the evolved schema (new field in after). Postgres differs: it gets current schema from the catalog/replication stream and uses a schema change topic but no separate history topic. Schema changes propagate to the (Avro/JSON) schemas registered for each table topic, so a schema registry must allow compatible evolution.

go deeper

for a junior

Know Debezium tracks table structure so it can decode changes, and that schema changes show up in events.

for a middle

Distinguish the internal schema history topic (MySQL restart state) from the public schema change topic, and why MySQL needs it but Postgres doesn't.

for a senior

Reason about schema evolution flowing into registry compatibility and protecting the history topic as critical state.

for a principal

Set registry compatibility policy, plan DDL/online-schema-change handling, and design consumer contracts resilient to schema evolution.

**Why schema tracking is needed.** A MySQL **binlog** in ROW format records changed rows largely by *ordinal column position and value*, not by full self-describing names and types for every event. To turn `(col1=…, col2=…)` into a meaningful event with named, typed fields, Debezium must know the table's **DDL (Data Definition Language) — its CREATE/ALTER structure — as it was at the binlog position of each row event**. Schemas also change over time (ALTER TABLE), so Debezium has to track DDL history, not just current state. **Two different topics — don't confuse them:** 1. **Schema history topic (internal).** For **MySQL** (and SQL Server, Oracle), Debezium writes every DDL it observes — plus the binlog/log position where it applied — to a dedicated **schema history topic** configured by `schema.history.internal.kafka.topic` (older name: `database.history.kafka.topic`). This topic is **internal plumbing for the connector itself**: on restart, Debezium *replays* it to reconstruct the exact table structures up to its resume offset, so it can correctly interpret older binlog entries. **It must not be deleted or its retention misconfigured** — losing it forces a re-snapshot. It should be single-partition and effectively retained forever (compacted/infinite retention). 2. **Schema change topic (public).** Separately, Debezium can publish a **schema change topic** — named after the connector's `topic.prefix` (e.g. `<prefix>`) — containing **DDL change events for downstream consumers** who want to react to schema changes (e.g. evolve a sink table). Controlled by `include.schema.changes` (MySQL default `true`). Each record describes the database, the DDL statement(s), and the position. This is *informational output*, not connector state. **Postgres difference.** The Postgres connector obtains table structure from the database **catalog** and the logical replication stream itself (which is more self-describing than MySQL binlog), so it does **not** need an internal schema history topic. It still can produce a schema change topic for consumers. **How a column add flows through:** 1. DBA runs `ALTER TABLE orders ADD COLUMN coupon VARCHAR(20)`. 2. Debezium captures the DDL: records it in the schema history topic (MySQL) and emits a schema change event to the schema change topic. 3. The internal table model now has the new column. 4. Subsequent row events include `coupon` in `before`/`after` with the new field in the event's value schema. 5. The per-table topic's **value schema evolves** — with Avro + Schema Registry this registers a new schema version, which must be **backward/forward compatible** (adding a nullable column with a default is compatible; dropping/renaming or type-narrowing may break compatibility and require a registry policy decision or consumer changes). **Operational notes:** - Treat the **schema history topic as critical state**: monitor it, never let it be auto-deleted; a corrupted/empty history topic means re-snapshot. - Some DDL (e.g. certain online schema-change tools, gh-ost/pt-osc) needs special handling; Debezium has settings to handle table renames and DDL filtering. - Choose registry compatibility (BACKWARD is common) deliberately so additive schema changes don't break consumers.

  • What happens if the MySQL schema history topic is accidentally deleted?
    Debezium can no longer reconstruct the table structures needed to interpret older binlog events, so the connector fails to resume and typically requires a fresh initial snapshot. The history topic is critical, non-disposable connector state.
  • Why doesn't the Postgres connector need a schema history topic?
    The Postgres logical replication stream plus the live catalog provide enough type/structure information; Debezium reads current schema from the database rather than reconstructing DDL history from a topic.
  • How does adding a nullable column affect the topic's Avro schema and consumers?
    It registers a new schema version. If the column is nullable with a default, it is backward/forward compatible, so existing consumers keep working. Incompatible changes (renames, type narrowing) may violate the registry's compatibility policy and break consumers.

saying these in an interview costs you the question

  • Conflating the internal schema history topic with the public schema change topic — they serve different purposes.
  • Saying the schema history topic can be safely deleted or short-retention.
  • Claiming Postgres needs a schema history topic like MySQL does.
  • Assuming all DDL changes are automatically schema-registry compatible.

context