Why would you set heartbeat.interval.ms on a Debezium Postgres connector capturing a low-traffic table?
answer
- No events means no confirmation
- Something must keep the position moving
- Retention is bounded by consumer progress
- The other database's writes are invisible to you
- One switch emits, another writes to the source
basics
~20 sWith no writes to captured tables Debezium emits nothing, commits no new offset, and the PostgreSQL replication slot never advances, so write-ahead log accumulates until the disk fills. Heartbeats emit periodic events so offsets and the slot keep moving.
solid answer
~50 sA Debezium Postgres connector only confirms its position when it has something to report. Capture a table nobody writes to and the connector sits idle, the slot's confirmed position stays where it was, and Postgres must retain every WAL segment since then — quietly, until the volume fills and the whole database stops accepting writes. `heartbeat.interval.ms` (disabled by default) makes the connector emit a heartbeat event on a schedule to `__debezium-heartbeat.<topic.prefix>`, which gives it a reason to flush an offset and let the slot advance. The harder variant is a busy cluster where the writes land in a *different* database: logical decoding is per-database, so the connector genuinely receives nothing to confirm. For that case `heartbeat.action.query` runs a statement you configure against the captured database each heartbeat — typically an insert into a small heartbeat table — manufacturing local change the connector can consume and acknowledge.
code
properties · 4 linesheartbeat.interval.ms=10000
-- needed when the captured database itself is idle
heartbeat.action.query=UPDATE dbz_heartbeat SET ts = now() WHERE id = 1
table.exclude.list=public.dbz_heartbeatgo deeper
Know that this setting makes the connector emit periodic heartbeat events even when nothing changes, and that the point is to keep the connector's recorded position moving.
Explain the causal chain: no events means no flushed position, an unconfirmed replication slot means the server retains write-ahead log, and retention is bounded by consumer progress rather than time or size.
Show you would set an interval, add an action query when the captured database itself is idle, exclude the heartbeat table, and independently alert on slot lag rather than trusting the mitigation.
Own the policy question: whether a lagging slot is allowed to take the database down or is invalidated and rebuilt, who owns slot cleanup when a connector is decommissioned, and what the recovery budget for a forced backfill is.
## The failure this prevents A PostgreSQL replication slot guarantees that the server keeps every write-ahead log record the consumer has not yet confirmed. That guarantee is the whole point of a slot — it is why a Debezium connector can be stopped for an hour and resume without losing changes — and it is also a loaded gun. Retention is bounded by consumer progress, not by time or size, unless an administrator has explicitly capped it. A slot whose confirmed position stops moving pins WAL indefinitely, and the symptom is a database disk filling steadily until Postgres refuses writes. The database looks healthy right up to the moment it isn't. The surprising part is that a *running, healthy* connector can cause this. ## Why an idle connector stops confirming Debezium advances its recorded position as a side effect of producing change events. Produce nothing and there is nothing to flush, so the last confirmed position stays put. Two shapes of workload trigger it: **A quiet captured table.** You capture `public.exchange_rates`, updated twice a day. Between updates the connector idles. Meanwhile every other table in the same database is busy, WAL keeps being written, and the slot's confirmed position lags further behind the log's head with every transaction. **Writes in a different database of the same cluster.** This one is worse and less obvious. Logical decoding in Postgres is scoped to a single database, so a connector attached to `analytics` is not shown changes made in `orders` at all. It cannot advance past them by decoding them, because it never sees them. From the connector's point of view the source is completely still; from the server's point of view the WAL is growing fast. ## What heartbeats do `heartbeat.interval.ms` defaults to `0`, meaning disabled. Set it to something on the order of seconds-to-tens-of-seconds and the connector emits a heartbeat record to `__debezium-heartbeat.<topic.prefix>` on that schedule. That record is a real produced event, so it carries the connector's current source position and gives it a reason to flush an offset. The slot advances, and Postgres can recycle WAL up to the confirmed point. Heartbeats have two secondary benefits worth naming. They keep the network connection and any intermediate proxies from timing out an idle replication connection, and they give monitoring a positive liveness signal: a topic that has gone quiet distinguishes 'the connector is fine, the source is quiet' from 'the connector is wedged'. Without them those two states look identical from the outside. ## When heartbeats alone are not enough In the cross-database case, emitting heartbeats lets the connector re-flush the position it already has — but that position cannot move past changes it is never given. The fix is `heartbeat.action.query`: a statement Debezium executes against the *captured* database on each heartbeat tick. The idiomatic version inserts or updates a row in a tiny dedicated table: ```sql CREATE TABLE dbz_heartbeat (id int PRIMARY KEY, ts timestamptz); ``` with the action query updating it. That write produces WAL *in the captured database*, which logical decoding does deliver to the connector, which then confirms a fresh position and unblocks retention. Keep the heartbeat table out of your include list, or accept a steady trickle of uninteresting change events on a topic. ## Other engines The setting is not Postgres-only. On SQL Server and Oracle, heartbeats keep an otherwise idle connection alive and keep the recorded position fresh so a restart does not re-read a long stretch of log. Postgres is where the consequence of skipping it is catastrophic rather than merely inconvenient, because of the slot's unbounded retention. ## Operating it Heartbeats are a mitigation, not a monitor. Independently alert on slot lag — the distance between the current log position and each slot's confirmed position — and on the age of the oldest unconfirmed position. Alert on the *derivative* too: a slot that has not moved in ten minutes on a busy cluster is a problem long before the disk is full. Consider bounding retention at the server level so that a runaway slot is invalidated rather than allowed to take the database down, and make sure that trade-off is a deliberate decision: an invalidated slot means the connector loses its place and needs a fresh backfill, which is often still preferable to an outage. Also watch the operational trap of a *stopped* connector. Heartbeats do nothing for a connector that is not running, and a slot left behind by a decommissioned connector is the classic cause of a disk-full incident weeks later. Deleting the connector must include dropping its slot.
- Heartbeats are enabled and the slot still does not advance. What do you check first?Whether the connector is actually running and producing — a paused or failed task emits no heartbeats at all. Then whether the writes are landing in a different database of the same cluster, in which case heartbeats alone cannot help and you need `heartbeat.action.query` to generate change in the captured database. Finally, whether a second, forgotten slot from a decommissioned connector is the one pinning WAL.
- What breaks if you cap replication-slot retention at the server so a lagging slot is invalidated instead of filling the disk?You trade an outage for a rebuild. An invalidated slot means the connector's position no longer exists in the retained log, so it cannot resume and must be re-established with a fresh backfill of the captured tables. That is usually the better failure, but it must be a deliberate choice with an alert attached, not a surprise discovered during recovery.
- Should the heartbeat table be captured by the connector?No — exclude it. Capturing it turns every heartbeat tick into a change event on a data topic, adding steady noise for consumers and inflating volume for no analytical value. Its only job is to produce write-ahead log activity in the captured database so the connector has something to acknowledge.
The slot is a bookmark in a book the library will not reshelve until you say you have passed that page. Heartbeats are you calling in to say 'still on page 900' — and the action query is the librarian writing you a new page to read when nothing else arrives.
saying these in an interview costs you the question
- Assuming an idle connector cannot harm the source database
- Thinking replication-slot retention is capped by time or size by default
- Expecting heartbeats to help when the connector is stopped
- Believing a connector sees changes from other databases in the cluster
- Deleting a connector without dropping its replication slot