skip to content

A CDC initial snapshot of a 2 TB table has run eight hours and the source's disk is filling — what is happening?

level: seniorimportance: should knowfreq 58%

answer

  1. nothing has consumed the log yet
  2. the pinned position holds everything after it
  3. an open read view stops version cleanup
  4. every cost here scales with duration
  5. dropping the slot frees disk and loses the baseline

basics

~20 s

Streaming has not started, so every log segment from the pinned position onward must be retained, and the snapshot's open read view blocks reclamation of superseded rows. Both grow with snapshot duration, and the source runs out of space before the scan ends.

solid answer

~50 s

Two accumulations run in parallel while a snapshot is in flight. First, the pinned log position has not been consumed yet, so the source cannot recycle anything after it — in Postgres an unconsumed replication slot pins WAL indefinitely; in MySQL the binary log grows or, worse, is purged on schedule and destroys the position. Second, the snapshot's long read view prevents cleanup of superseded row versions, so table and undo space bloat database-wide. Both scale with **duration**, so the fix is to make snapshots shorter or make them consume the log while they run. Immediate actions: buy space and raise retention, confirm which of the two is actually growing, and check whether the connector can stream concurrently. Structural fixes: chunked or watermark-based incremental snapshots, parallel per-table scans, snapshotting from a replica, or starting from schema only and backfilling history from an existing extract.

code

sql · 11 lines
sql
-- Postgres: how much WAL is the unconsumed slot pinning?
SELECT slot_name, active, wal_status,
       pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS retained
FROM pg_replication_slots;

-- Postgres: is a long read view blocking cleanup?
SELECT pid, state, now() - xact_start AS age, left(query, 60)
FROM pg_stat_activity
WHERE xact_start IS NOT NULL
ORDER BY xact_start
LIMIT 5;

go deeper

for a junior

Understand that until the connector starts reading the log, the database has to keep everything after the position it pinned, and that keeping it costs disk space that grows the longer the snapshot runs.

for a middle

Distinguish the two accumulations — retained log segments and unreclaimed row versions — and name the metric that tells them apart before proposing any fix.

for a senior

Show the triage: stabilise first with space and retention or by streaming concurrently, then choose the structural fix — chunked or incremental snapshot, parallel scans, a replica source — and explain why dropping the slot destroys the baseline.

for a principal

Set the invariant that projected snapshot duration must sit well inside log retention, and decide organisationally whether the source is scanned at all or the baseline comes from an existing extract reconciled by key.

## Diagnose which thing is growing "Disk is filling" has two very different causes during a snapshot, and they need different responses. **Log retention.** The connector pinned a position before it started scanning and has not consumed a byte of the log yet. On Postgres the logical replication slot holds that position; the server will not recycle WAL beyond it, so WAL grows at the source's full write rate for as long as the snapshot runs. On MySQL the binary log is not held by the connector at all — it is purged on its own schedule, so the risk inverts: the disk may be fine and the *pinned position expires*, destroying eight hours of scanning. **Version bloat.** The snapshot's read view is the oldest in the database, so cleanup cannot advance past it. Dead row versions and undo space accumulate across every table, not only the captured ones. Check slot retained bytes or lag, the age of the oldest transaction, dead-tuple counts or undo size, and the free space trend. One number tells you which lever to pull. ## Immediate mitigations - **Add space and raise retention headroom** so the snapshot can finish rather than fail at hour nine and waste the entire scan. - **Guard against the opposite failure on MySQL** by confirming binary-log retention comfortably exceeds the projected snapshot duration. - **Let the connector consume the log while snapshotting** if the implementation supports interleaving. This is the single most effective change: the position advances, WAL is released, and the catch-up phase after the scan shrinks to near zero. - **Do not silently drop the slot to free space.** It frees the disk and destroys the baseline; the connector must start over. Be explicit about the abort decision. If retention will not survive the remaining scan time, killing the snapshot now and restarting with a better strategy costs less than discovering at hour eleven that the position is gone. ## Structural fixes, roughly in order of leverage **Chunk and checkpoint.** Walk the table in primary-key ranges with a bounded fetch size, persisting progress per chunk. A restart resumes at the next chunk instead of at row zero, which turns a catastrophic failure into a pause. **Incremental snapshot with watermarks.** Backfill chunk by chunk *while streaming runs*, reconciling each chunk against log events observed during it. Retention pressure disappears because the log is being consumed the whole time, and the backfill becomes pausable and resumable. **Parallelism.** Multiple tables scanned concurrently, or one large table split by key range across workers, cuts wall-clock duration directly — and duration is what every cost here is proportional to. The limit is source I/O; past a point you are simply moving the outage into the OLTP workload. **Snapshot from a replica.** Moves the scan's read I/O off the primary entirely. Two caveats: the replica has its own read-view pressure and can fall behind while serving a long scan, and the log position you pin must be meaningful on the primary — a GTID set is portable, raw file-and-offset coordinates are not. **Skip the row snapshot.** Capture schema only, start streaming immediately, and reconstruct history from an existing warehouse extract joined by primary key. This eliminates the scan altogether and is the right answer surprisingly often when a full nightly extract already exists. **Trim what you capture.** Excluding blob columns, archival partitions, or soft-deleted rows from the snapshot can remove most of the volume. If the sink genuinely needs 2 TB, it needs it; frequently it does not. ## Prevention, stated as an invariant Measure the projected snapshot duration and the source's log retention window, and require **retention comfortably greater than duration**, with alerts on retained-log bytes, slot lag, oldest-transaction age and free space. A snapshot whose duration is within the same order of magnitude as retention is an incident waiting for a slow day. ## What a strong answer sounds like Name both accumulations, say which metric distinguishes them, take the stabilising action first (space and retention, or concurrent streaming), then propose the structural fix and the invariant that prevents recurrence. Weak answers reach straight for "drop the slot" or "restart the connector", both of which throw away the eight hours already spent.

  • Why is dropping the replication slot to free disk space during a snapshot the wrong move?
    The slot is the only thing holding the log position the snapshot is anchored to. Dropping it releases the space immediately and irreversibly discards the baseline: the connector can no longer resume from that point, so the eight hours of scanning are wasted and the whole snapshot must be repeated. If space genuinely cannot be found, aborting deliberately and restarting with a chunked strategy is the honest version of the same decision.
  • How does snapshotting from a read replica change the log position you pin?
    It moves read I/O off the primary but the pinned position must still be interpretable by the stream you will consume. A GTID set is a global identity and transfers cleanly; raw binary-log file and offset coordinates are replica-local and meaningless on the primary. Postgres additionally needs a version supporting logical decoding on standbys. Get this wrong and the handover silently anchors to the wrong point.
  • What single change most reduces retention pressure without shortening the scan?
    Consuming the log while the snapshot runs. If the connector opens the stream at the pinned position and buffers or interleaves events during the scan, the position keeps advancing, log segments are released, and the post-snapshot catch-up shrinks to almost nothing. Scan duration is unchanged, but the accumulation that was going to fill the disk stops accumulating.
  • What would you monitor so this never becomes a surprise?
    Retained log bytes or slot lag, the age of the oldest open transaction, dead-tuple or undo growth, free space trend, and snapshot progress in chunks. Then hold an explicit invariant: projected snapshot duration must sit well inside the source's log retention window, with an alert when the ratio degrades rather than when the disk is already full.

It is like keeping every receipt since the stocktake began because you have not filed any of them yet — the pile grows for exactly as long as the count takes, and burning the pile means starting the count again.

saying these in an interview costs you the question

  • Suggests dropping the replication slot to reclaim disk space
  • Blames the snapshot for blocking writers rather than pinning versions
  • Assumes MySQL and Postgres fail the same way during a long snapshot
  • Restarts the connector without changing anything about the strategy
  • Ignores that every cost here scales with snapshot duration

context