Which Debezium connector class do you configure for a MySQL, PostgreSQL, MongoDB or SQL Server source?
answer
- One connector per database engine
- Not a driver setting, a class name
- io.debezium.connector.<engine>.<Engine>Connector
- Binlog, replication slot, change streams, CDC tables
basics
~10 sDebezium ships a separate connector class per source: MySqlConnector, PostgresConnector, MongoDbConnector and SqlServerConnector, all under io.debezium.connector. Each reads that engine's own change log and has its own database-side prerequisites.
solid answer
~40 sDebezium is a family of connectors, one per source engine, selected with `connector.class`: `io.debezium.connector.mysql.MySqlConnector`, `io.debezium.connector.postgresql.PostgresConnector`, `io.debezium.connector.mongodb.MongoDbConnector`, `io.debezium.connector.sqlserver.SqlServerConnector` (plus Oracle, Db2 and others). They share a configuration vocabulary — `topic.prefix`, `database.hostname`, `database.port`, `database.user`, `database.password`, the include/exclude lists — but each reads a different mechanism: MySQL the binlog, Postgres a logical replication slot through an output plugin, MongoDB change streams, SQL Server the CDC change tables its own agent populates. The real work of picking a connector is satisfying its prerequisites: `binlog_format=ROW` and a unique `database.server.id` for MySQL; `wal_level=logical` and a replication-capable role for Postgres; CDC enabled per database and table for SQL Server; a replica set for MongoDB.
code
properties · 10 linesconnector.class=io.debezium.connector.postgresql.PostgresConnector
topic.prefix=inventory
database.hostname=pg.internal
database.port=5432
database.user=debezium
database.password=${file:/secrets/pg.properties:password}
database.dbname=shop
plugin.name=pgoutput
slot.name=dbz_shop
publication.name=dbz_shop_pubgo deeper
Be able to name the connector for a given source and point to the connector.class property. Knowing the naming shape io.debezium.connector.<engine>.<Engine>Connector is enough to reconstruct any of them.
Explain what each connector actually reads — binlog, logical replication slot, change streams, CDC change tables — and name the server-side settings each one requires before it can start.
Show that you plan the database-side work first: replication grants, log retention, CDC enablement, and one connector instance per source database with non-colliding identities.
Own the standardisation argument: a single capture technology across heterogeneous engines buys one operational model, and you should be able to say where that uniformity stops and per-engine expertise is still required.
## What `connector.class` selects Debezium is not a single connector with a driver setting. It is a family of independent connector implementations, each packaged as a plugin and each written against one database engine's change mechanism. The `connector.class` property in a connector configuration names the Java class to instantiate, and that single line determines which log Debezium reads, which prerequisites the DBA must satisfy, and which source-specific properties are legal in the rest of the config. ## The per-source classes - **MySQL** — `io.debezium.connector.mysql.MySqlConnector`. Connects as a replication client and reads the binary log. - **PostgreSQL** — `io.debezium.connector.postgresql.PostgresConnector`. Opens a logical replication slot and decodes changes through an output plugin (`plugin.name`, commonly `pgoutput`, which ships with the server). - **MongoDB** — `io.debezium.connector.mongodb.MongoDbConnector`. Consumes change streams from a replica set or sharded cluster, configured with `mongodb.connection.string`. - **SQL Server** — `io.debezium.connector.sqlserver.SqlServerConnector`. Reads the CDC change tables that SQL Server's own capture agent fills from the transaction log, so it is one step removed from the log itself. Others exist — Oracle, Db2, Cassandra, Vitess, Spanner — with the same naming shape. Knowing the shape matters more than reciting the strings: it is always `io.debezium.connector.<engine>.<Engine>Connector`. ## The shared vocabulary Whatever the source, a Debezium configuration looks broadly the same. `topic.prefix` gives the connector a logical name that becomes the namespace for its topics. `database.hostname`, `database.port`, `database.user` and `database.password` describe the connection. `table.include.list` and `table.exclude.list` scope what is captured; `column.include.list` and `column.exclude.list` scope the fields. `heartbeat.interval.ms` keeps offsets moving on quiet sources. Value-shaping options such as `decimal.handling.mode` and `time.precision.mode` behave consistently across engines. That commonality is Debezium's main selling point: one operational model over four or more very different databases. ## Where the connectors genuinely differ The differences that break deployments are on the database side, not in the connector config. **MySQL** needs `binlog_format=ROW` and `binlog_row_image=FULL` — with statement-based logging there are no before/after row images to emit — plus a user granted `SELECT`, `RELOAD`, `SHOW DATABASES`, `REPLICATION SLAVE` and `REPLICATION CLIENT`, and a `database.server.id` that no other replication client uses. Binlog retention on the server sets how long a stopped connector can stay stopped before its position is gone. **PostgreSQL** needs `wal_level=logical`, a role with the `REPLICATION` attribute, and a slot (`slot.name`) plus a publication (`publication.name`). One connector covers one database, named by `database.dbname`. The slot is the operational hazard: it pins write-ahead log until the connector confirms consumption. **SQL Server** needs CDC enabled at the database level and then per table, via the system stored procedures `sys.sp_cdc_enable_db` and `sys.sp_cdc_enable_table`, with SQL Server Agent running to keep the capture job alive. Databases are listed in `database.names`. **MongoDB** needs a replica set (change streams do not exist on a standalone server) and a user allowed to read the oplog/change stream. ## What downstream consumers see The change-event shape is deliberately uniform across connectors, which is why a sink written for one source usually works for another. The engine-specific detail surfaces inside the event's source metadata block: MySQL reports a binlog file and position, Postgres an LSN, SQL Server a change LSN, MongoDB a resume token. If you build tooling that reasons about source position — lag dashboards, replay logic — that block is where the portability stops. ## Practical rules One connector instance covers one source database (one Postgres database, one MySQL server's selected schemas). Two connectors against the same server must differ in every identity-bearing property: distinct `topic.prefix`, distinct slot and publication names on Postgres, distinct `database.server.id` on MySQL. Getting the class right is the easy part; the interview usually moves straight to the prerequisites, because that is where real deployments stall.
- Why does the SQL Server connector behave differently from the MySQL and Postgres ones under load?It does not tail the transaction log itself. SQL Server's own capture agent copies changes into CDC change tables, and Debezium queries those tables. So end-to-end latency depends on the agent's polling and the change-table cleanup job, and if CDC is disabled or the Agent stops, Debezium simply sees no new rows rather than failing loudly.
- Can one connector instance capture two PostgreSQL databases in the same cluster?No. Logical decoding is per-database, and the Postgres connector takes a single `database.dbname`. Two databases means two connector instances, each with its own slot, publication and `topic.prefix`. MySQL is different — one connector covers the whole server and you scope it with `database.include.list`.
Think of the connector classes as different adapters for the same appliance: identical plug on your side, four completely different sockets on the wall, each with its own wiring rules.
saying these in an interview costs you the question
- Thinking one Debezium connector handles all databases via a driver setting
- Assuming SQL Server capture needs no database-side enablement
- Believing MySQL statement-based binlog is enough for CDC
- Expecting MongoDB change streams to work on a standalone server
- Copying a config between engines without changing prerequisites