Why does a logical replication slot that no CDC consumer reads fill the source primary's disk?
answer
- the server keeps what a consumer has not confirmed
- slots outlive the connector that made them
- the promise survives restarts
- a quiet database can still pin busy log
- cap the retention or lose the primary
basics
~20 sA slot records the oldest log position its consumer still needs, and the server refuses to recycle write-ahead log files at or after that position. If nothing advances the slot, the retained log grows without bound until the disk is full.
solid answer
~60 sA replication slot is a durable promise by the server to keep everything a named consumer has not yet confirmed. It stores a restart position, and the checkpointer will not recycle WAL segments the slot still needs. Stop the connector, lose the consumer, or leave a slot behind after decommissioning a pipeline, and the primary keeps accumulating WAL — for as long as the slot exists, across restarts, because slots are persistent. The end state is a full data disk on the primary, which stops accepting writes: a capture-layer mistake taking down the production database. A logical slot also holds back the catalog `xmin`, so vacuum cannot clean up catalog rows the decoder might still need, adding bloat on top. The Postgres-specific trap is a low-traffic captured database inside a busy cluster: WAL is cluster-wide but a logical slot decodes one database, so the slot's confirmed position barely moves while WAL piles up. Monitor retained bytes per slot, cap retention with a slot size limit, drop slots you have abandoned, and have the capture write a periodic heartbeat so the position keeps advancing.
code
sql · 6 linesSELECT slot_name, active, wal_status,
pg_size_pretty(
pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)
) AS retained_wal
FROM pg_replication_slots
ORDER BY pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn) DESC;go deeper
Remember the one fact: a slot makes the server keep log files until the consumer confirms them, so an abandoned slot grows the source's disk usage.
Explain the mechanism — a durable restart position that blocks segment recycling, persisting across restarts — and that a logical slot also holds back catalog cleanup.
Show you would operate it: alert on retained bytes and inactive slots, cap retention so a slot is invalidated rather than the primary wedged, use a heartbeat for quiet databases, and know the incident options and their costs.
Own the dependency this creates: the transactional primary now depends on a downstream consumer, so decide the retention cap policy, the slot lifecycle in environment teardown, and who is accountable when a pipeline outage threatens a production database.
## What a slot actually promises A replication slot is a small piece of durable server state that says: *a consumer named `x` has confirmed up to position P, so do not throw away anything from P onward.* That promise is exactly what makes log-based capture reliable — a connector can crash, be redeployed, or be offline for a maintenance window and resume without losing a single change, because the server held the log for it. The same promise is what makes it dangerous. The server has no way to know the difference between "this consumer will be back in five minutes" and "this consumer was deleted three weeks ago". A slot is persistent: it survives connector disconnects and server restarts, and it is removed only when someone explicitly drops it. ## The failure sequence 1. A slot exists with a restart position P. 2. The consumer stops confirming — the pipeline is paused, the connector container is gone, credentials expired, the task is failing in a crash loop, or the pipeline was retired and nobody dropped the slot. 3. The application keeps committing, so the log keeps growing beyond P. 4. Checkpoints run, but the server may not recycle segments at or after P, because the slot still needs them. 5. The WAL directory grows monotonically until the filesystem is full. 6. A primary that cannot write WAL cannot commit. Writes fail; the database is effectively down. That last step is why interviewers ask this: it is the case where the *analytics* pipeline takes out the *transactional* database. Nothing else in a typical warehouse stack has that reach. ## The second, quieter cost A logical slot also pins a catalog `xmin` — the decoder needs the catalog as it stood when those changes were made, so vacuum is prevented from removing catalog tuples newer than that horizon. An old slot therefore also produces catalog bloat and degraded planning over time, which is easy to misdiagnose as an unrelated problem. ## The low-traffic-database trap In Postgres, WAL is written for the whole cluster, but a logical slot decodes changes for exactly one database. If the captured database is quiet while other databases in the same cluster are busy, the decoder finds nothing to emit, so the consumer has nothing to confirm and the slot's confirmed position stays where it is — while WAL from the *other* databases piles up behind it. The pipeline looks perfectly healthy (no lag in events, no errors) and the disk fills anyway. The standard remedy is a heartbeat: the capture periodically writes a row to a small dedicated table in the captured database, which produces a change the decoder can emit and the consumer can confirm, dragging the slot's position forward. Most capture tools expose this as a heartbeat interval setting; the mechanism is the same whoever implements it. ## Operating slots properly - **Monitor retained bytes per slot, not connector lag.** Connector-reported lag can be near zero while the slot retains gigabytes. ```sql SELECT slot_name, active, wal_status, pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS retained FROM pg_replication_slots ORDER BY pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn) DESC; ``` - **Alert on `active = false`.** An inactive slot on a busy database is a countdown. - **Cap the retention.** Postgres 13 added `max_slot_wal_keep_size`, which is unlimited unless you set it. Setting it means an over-retaining slot is *invalidated* — the pipeline breaks and must be re-seeded — rather than the primary filling its disk. That is usually the right trade: a broken pipeline is recoverable, a wedged primary is an outage. - **Make slot lifecycle part of decommissioning.** Retiring a pipeline must include dropping its slot. Orphan slots from deleted environments are the single commonest cause of this incident. - **Keep headroom and know the drain path.** During an incident the options are: restart the consumer so it confirms and the slot advances (best), drop the slot and accept re-seeding the pipeline (fast, costly downstream), or add disk to buy time. Deleting WAL files by hand corrupts the database — never an option. ## The framing that matters This risk is the price log-based capture charges for its main advantage. Polling cannot take the source down; it just scans. A slot buys you resumable, delete-accurate, low-latency capture, and in exchange the source database now has a dependency on a downstream consumer staying alive. Owning log-based CDC means owning that dependency explicitly — with monitoring, a retention cap, and a decommissioning checklist.
- Capture lag reported by the connector is near zero, yet retained log keeps growing. What is happening?Most likely the captured database is idle inside a busy cluster: WAL is cluster-wide but the slot decodes one database, so the decoder emits nothing, the consumer confirms nothing, and the slot's position stays put while other databases generate WAL. A periodic heartbeat write into the captured database advances it.
- The primary is minutes from a full disk because of an abandoned slot. What do you do?Confirm which slot is retaining the most, then either restart its consumer so it confirms and the position advances, or drop the slot outright and accept that the pipeline must be re-seeded from a fresh initial load. Buy time with disk if you can. Never delete log files manually.
- What is the tradeoff of setting a maximum slot retention size?An over-retaining slot gets invalidated instead of filling the disk, so the database stays up but the pipeline breaks and needs re-seeding from scratch. You are choosing a recoverable downstream outage over a source-side production outage, which is almost always the right way round.
A slot is a hold on a library shelf: nothing behind it can be re-shelved until the borrower checks in, and the borrower who moved away never does.
saying these in an interview costs you the question
- Thinking a slot disappears when the connector disconnects
- Monitoring connector lag instead of retained log bytes
- Deleting write-ahead log files by hand to free space
- Assuming an idle captured database means an idle slot
- Leaving slots behind when tearing down a pipeline