Your Debezium initial snapshot has run eight hours and the source's replication slot keeps growing — what now?
answer
- snapshotting means the log is not being read
- retention grows with write rate times duration
- the risk lands on the source, not the sink
- stream first, backfill afterwards
- chunked backfills can be paused
basics
~20 sThe connector is snapshotting, so it is not consuming the log yet, and the slot retains every change since the snapshot began until the source disk fills. Triage disk headroom first, then restart with a mode that streams immediately and backfill the tables with incremental snapshots.
solid answer
~50 sDuring the classic snapshot phase a Debezium connector reads tables with SELECTs and does **not** consume the transaction log, so the replication slot it created sits unconsumed and pins log segments on the source. The longer the snapshot runs, the more the source retains — and a full disk takes the production database down, not just the pipeline. First, triage: how much free space is left, how fast the retained volume is growing, and how much of the snapshot remains. If the finish line is close and disk allows, let it run. Otherwise abort and re-approach: bring the connector up with a mode that skips the data (`no_data`/`schema_only`, or `never` on Postgres) so streaming starts within seconds and the slot begins draining, then issue ad-hoc **incremental** snapshots table by table. Those run alongside streaming, checkpoint per chunk, and can be paused or stopped under load.
code
properties · 9 lines# stream immediately, then backfill per table with signals
connector.class=io.debezium.connector.postgresql.PostgresConnector
topic.prefix=inventory
snapshot.mode=never
signal.data.collection=inventory.dbz_signal
table.include.list=inventory.customers,inventory.orders,inventory.dbz_signal
# levers if a classic snapshot is unavoidable:
# snapshot.fetch.size=10240
# snapshot.include.collection.list=inventory.customersgo deeper
Recall the causal chain: while the connector is copying tables it is not reading the log, so the source keeps log segments for it. That is enough to know why a long snapshot is risky.
Explain the arithmetic — retained volume is the source's write rate times the snapshot duration — and name the settings that shorten a snapshot, such as fetch size, parallel snapshot threads and restricting which tables are copied.
Demonstrate triage under pressure: measure free space, growth rate and remaining work before acting; know that aborting frees retention at a cost; and be able to argue for the stream-now, backfill-with-incremental-snapshots restructuring.
Own the standard: how onboarding a new source is sized and approved, what alerting lives on the source database rather than the pipeline, and when a bulk load should bypass the CDC tool altogether.
## What is actually happening A Debezium connector's log position is only advanced by the streaming phase. During a classic snapshot there is no streaming phase running, so nothing acknowledges the log. On Postgres that means the replication slot holds a position from before the snapshot started and the server must retain every WAL segment from there onward; on MySQL it means binlog files the connector still needs must not be purged. The retained volume grows with the source's write rate multiplied by the snapshot's duration — completely independent of how large the tables being copied are. A modest table on a very busy database can be a worse problem than a huge table on a quiet one. The failure mode is severe because the pressure lands on the **source**, not on the pipeline. A full disk on the primary is a production outage. That is why the correct first move is not connector tuning; it is capacity triage. ## Triage, in order Establish the numbers before touching anything. How much free space remains on the volume holding the logs, and how fast is retained volume growing per hour? What is the connector's slot or offset lag now? How far through the snapshot is it — rows read versus rows expected, which the connector's progress logging and the source's own statistics can tell you. From those three you get the only decision that matters right now: will the snapshot finish before the disk does? If it will, and with margin, let it finish and fix the approach next time. If it will not, act: raise disk headroom if the platform allows it, or abort the connector and accept dropping the slot, which immediately releases the retained log — at the price of losing the ability to resume from that point, so any subsequent start must be reasoned about from scratch. ## Making the snapshot itself shorter Several levers reduce snapshot duration directly. `snapshot.fetch.size` controls how many rows are pulled per round trip. Some connectors expose `snapshot.max.threads` to snapshot multiple tables in parallel. `snapshot.include.collection.list` restricts which of the captured tables are snapshotted at all, letting you copy in waves. Splitting a very wide capture set across several connectors gives each a smaller snapshot. And `snapshot.select.statement.overrides` can cut a table down to the rows that matter — recent partitions rather than a decade of history. These all help, but they are optimisations of the wrong shape. They shorten the interval in which nothing consumes the log; they do not remove it. ## The structural fix The modern answer inverts the order. Start the connector in a mode that emits no rows — `schema_only`/`no_data` on the schema-history connectors, `never` on Postgres — so it reaches the streaming phase almost immediately. From that moment the slot advances, the source stops accumulating, and the pipeline is current for all new changes. Then backfill history by sending `execute-snapshot` signals per table. Incremental snapshots read in chunks bounded by window markers, interleave with streaming so the log keeps draining, checkpoint their progress in the connector's offsets so a restart resumes at the next chunk, and accept `pause-snapshot` and `stop-snapshot` so you can yield during peak hours. This converts a single unbounded risk window into a controllable background job. It is the reason incremental snapshots exist. ## Prevention Three habits stop this recurring. Do not create the slot or start the connector days before the pipeline is ready — the clock starts when the slot appears, not when data starts flowing. Monitor retained log volume and connector lag as a first-class alert on the **source database**, with a threshold that fires while there is still time to act. And size the plan before deploying: estimate snapshot duration from row counts and observed read throughput, multiply the source's write rate by that duration, and compare the result with free disk. If the arithmetic is uncomfortable, choose the stream-now-backfill-later path from the start rather than discovering the problem at 3 a.m. ## One more restart hazard Be careful about restarting a connector mid-snapshot to "apply a tuning change". On most connectors a classic snapshot that is interrupted begins again from the beginning, so you pay the whole cost twice and extend the retention window rather than shortening it. Some newer connector versions can resume — confirm for the exact connector and version before assuming either way.
- Why does the size of the retained log depend so little on the size of the table being snapshotted?Retention is driven by the source's write rate multiplied by how long the connector goes without consuming the log. A slow snapshot of a small table on a write-heavy database retains far more than a fast snapshot of a large table on a quiet one. That is why duration, not table size, is the number to estimate.
- Is dropping the replication slot a safe emergency lever?It relieves the disk immediately, which can be the right call during an outage, but it discards the position the connector would have resumed from. Anything written after the slot is dropped and before a new one exists is not captured, so you must plan the restart deliberately — typically a stream-now start plus incremental snapshots to rebuild the affected tables.
- Does restarting the connector to apply a tuning change during a snapshot help?Usually it hurts. On most connectors an interrupted classic snapshot restarts from the beginning, so you pay the read cost again and lengthen the window in which the log is not being consumed. Newer versions of some connectors can resume — verify before relying on it, and prefer changing approach entirely over restarting mid-copy.
saying these in an interview costs you the question
- Treats it as a Kafka problem rather than a source-disk problem
- Assumes the snapshot is consuming the log as it goes
- Restarts the connector mid-snapshot to tune it
- Thinks a bigger fetch size removes the retention risk
- Plans no backfill after starting in a stream-now mode