skip to content

Two Debezium MySQL connectors share the same database.server.id — what happens on the source?

level: seniorimportance: should knowfreq 38%

answer

  1. The connector is not a query client
  2. MySQL identifies replication participants numerically
  3. Uniqueness spans the whole topology
  4. Neither one fails cleanly
  5. Copied configurations are the usual culprit

basics

~20 s

Each Debezium MySQL connector registers as a replication client identified by database.server.id. When two share an id, MySQL drops one of them; both reconnect, kick each other off again, and neither keeps up with the binlog.

solid answer

~50 s

The MySQL connector does not poll — it joins the server's replication topology as a client, and `database.server.id` is the identity it claims. MySQL requires that identity to be unique across every replica and every binlog client attached to the server, so two connectors presenting the same value cause the server to disconnect one of them with an error about a duplicate id. Kafka Connect restarts the failed task, it reconnects, evicts the other, and you get a flapping loop: both connectors alternate between streaming and reconnecting, lag climbs on both, and the logs fill with reconnect noise rather than one clear failure. The classic cause is a staging connector copied from production against the same database, or a second connector added for a different set of tables. Fix it by allocating a distinct id per connector and keeping a registry of allocations, treating it exactly like an IP address on a shared network.

code

properties · 11 lines
properties
-- connector A: orders tables
topic.prefix=orders
database.server.id=184054
database.include.list=shop
table.include.list=shop.orders,shop.order_items

-- connector B: same server, different tables, DIFFERENT id
topic.prefix=customers
database.server.id=184055
database.include.list=shop
table.include.list=shop.customers

go deeper

for a junior

Remember that this property is required, numeric, and must differ for every connector. Recognising it in a configuration and knowing it identifies the connector to MySQL is the expectation.

for a middle

Explain why it exists: the connector attaches as a replication client, and MySQL enforces unique numeric identities across replication participants, so a duplicate gets disconnected.

for a senior

Diagnose from symptoms — two connectors flapping, sawtooth lag, repeated reconnects with no single exception — and distinguish the availability damage from correctness. Then fix it with distinct ids and check binlog retention while you are there.

for a principal

Own server ids as an allocated resource across teams, with a registry and templates that make collision structurally impossible, and set the policy that non-production environments never attach to a production binlog.

## What the property is `database.server.id` is a required numeric property on the Debezium MySQL connector, and it has no safe default. It exists because of how the connector reads changes: rather than querying tables, it opens a replication connection and asks the server to stream binary-log events, exactly as a MySQL replica would. MySQL identifies every participant in its replication topology by a numeric server id, and it enforces uniqueness among them. That means a Debezium connector is, from the server's perspective, another replica. It counts against the same namespace as your real read replicas, any other CDC tool attached to the same server, and every other Debezium connector pointed at it. ## The failure mode When a second client connects claiming an id that is already in use, MySQL does not merge the two streams or refuse the newcomer politely. It drops one connection, reporting that another replica with the same id has connected. The evicted connector's task fails. Kafka Connect, doing its job, restarts it. The restarted connector reconnects, and now *it* evicts the other one. The cycle repeats indefinitely. The symptom set is distinctive once you know it: - Both connectors alternate between `RUNNING` and failed/restarting rather than one failing cleanly. - Consumer lag on both connectors' topics climbs, sometimes in a sawtooth as each gets a brief window of streaming. - Connector logs show repeated disconnects and re-establishment of the binlog stream, with no single explanatory exception. - Nothing is obviously wrong with the database itself; it is serving queries normally. What is *not* happening is data loss. Each connector resumes from its own recorded binlog position when it reconnects, so events are not silently skipped — they are merely late, and duplicated at the boundary in the ordinary at-least-once way. The damage is availability and lag, not correctness. Prolong it, though, and lag can exceed the server's binlog retention, at which point a connector's recorded position is genuinely gone and you are into a rebuild. ## Why it happens Almost always, configuration copying. Someone clones a working production connector configuration to stand up staging, changes the hostname and the topic prefix, and leaves the id. If staging points at a replica of the same server, or at the same server entirely, the collision is immediate. The second common path is capturing two disjoint sets of tables with two connectors against one server — a legitimate design — where the second connector inherits the first one's id. A third, sneakier path: the value is unique among connectors but collides with a *real* MySQL replica, or with a change-capture tool someone else in the organisation attached to the same server months ago. Uniqueness has to hold across the whole topology, not just across your team's connectors. ## Preventing it Treat server ids as an allocated resource with an owner. Practical measures: - Keep a registry — a file in the repository, an entry in the service catalogue — mapping each id to the connector and environment that owns it. - Derive ids deterministically from something already unique, so two people cannot independently pick the same number. - Never copy the id when cloning a configuration; make it a required, uncommented value in your connector templates so it cannot be inherited by accident. - Keep environments genuinely separate: staging should not attach to production's binlog at all, and if it attaches to a restored copy it still needs its own id. ## The neighbouring MySQL prerequisites Since the topic is 'what MySQL requires of the connector', the interviewer often continues. The connector also needs `binlog_format=ROW` and `binlog_row_image=FULL`, or there are no complete row images to publish. Its user needs `SELECT`, `RELOAD`, `SHOW DATABASES`, `REPLICATION SLAVE` and `REPLICATION CLIENT`. And binlog retention on the server sets the outage budget: a connector stopped longer than retention cannot resume from its recorded position and must be re-established with a fresh backfill. That last one is the MySQL analogue of the Postgres slot problem, inverted — MySQL discards old log on a schedule rather than retaining it forever, so the risk is losing your place rather than filling the disk. ## Interview framing A strong answer names the mechanism (the connector is a replication client), predicts the flapping symptom rather than a clean error, distinguishes the availability impact from a correctness impact, and finishes with the allocation discipline that prevents recurrence.

  • Does the collision cause data loss on the affected topics?
    Not directly. Each connector resumes from its own recorded binlog position after every eviction, so events are delayed and duplicated at the boundaries rather than skipped. The real danger is second-order: if the flapping persists long enough that lag exceeds the server's binlog retention, a recorded position is no longer in the log and that connector needs a full re-establishment.
  • Is there a legitimate reason to run two Debezium MySQL connectors against one server?
    Yes — splitting disjoint table sets so a heavy table cannot delay a latency-sensitive one, or isolating a team's capture from another's. It is a supported pattern, provided each connector gets its own `database.server.id`, its own `topic.prefix` and its own schema-history topic. Remember both connectors read the entire binlog and filter, so the source cost roughly doubles.
  • Beyond the id, what MySQL-side settings must be right before capture works?
    Row-based binary logging with full row images, so complete before and after values exist to publish; a capture user granted SELECT, RELOAD, SHOW DATABASES, REPLICATION SLAVE and REPLICATION CLIENT; and a binlog retention window long enough to cover your worst expected connector outage, since a position older than retention cannot be resumed.

saying these in an interview costs you the question

  • Calling it a cosmetic label rather than a replication identity
  • Expecting a clean startup failure instead of a flapping loop
  • Assuming uniqueness only matters among Debezium connectors
  • Copying a production connector config wholesale into staging
  • Believing the collision silently drops change events

context