In Debezium, what do the snapshot-window-open and snapshot-window-close signals do?
answer
- two markers written into the log itself
- they fence one chunk of rows
- rows are held in a buffer meanwhile
- a live change evicts the buffered copy
- the log always wins over the read
basics
~20 sThey mark the boundaries of one incremental-snapshot chunk in the transaction log. Debezium buffers the chunk's rows, drops from that buffer any key it sees change between the two markers, and emits what remains — so a live change always wins over the older snapshot read.
solid answer
~50 sAn incremental snapshot reads a table in chunks while streaming continues, which raises an obvious hazard: a row read into a chunk may also change in the log at the same moment, and emitting the stale snapshot copy afterwards would overwrite the newer value at the sink. Debezium bounds each chunk with **watermarks**. It writes a `snapshot-window-open` row to the signalling table, selects the next chunk of rows by primary key into an in-memory buffer, then writes `snapshot-window-close`. Because both writes travel through the transaction log, the connector knows exactly which streamed events fall inside the window; for any of those, it removes the matching key from the buffer. The surviving buffered rows are emitted as `r` events after the window closes. The result is no duplicated key and no stale value landing on top of a newer one, with no table locks anywhere.
code
text · 6 linesINSERT dbz_signal type=snapshot-window-open <- appears in the log
SELECT next 1024 rows of inventory.customers by PK -> buffer
... streamed from the log inside the window:
UPDATE inventory.customers id=42 -> emitted, and key 42 dropped from buffer
INSERT dbz_signal type=snapshot-window-close <- appears in the log
emit remaining buffered rows as op="r" (id=42 is not among them)go deeper
At minimum, recall that a chunk of an incremental snapshot is fenced by two markers and that a row changing during that fence is taken from the live change, not from the snapshot read.
Explain the sequence concretely: marker, buffered chunk read, marker, then evict from the buffer any key seen in the log inside that fence, then emit the rest as read events. Be able to say why the markers must go through the log.
Show that you can operate it — sizing chunks against connector memory and source load, diagnosing a snapshot that starts and produces nothing because the signalling table is not captured, and knowing the read-only variant for untouchable sources.
Own the guarantee end to end: what the sink must do for this to hold (key-addressed upserts), what it costs the source, and when you would instead backfill outside the CDC tool entirely.
## Why a window is needed During an incremental snapshot two producers are emitting events for the same table at the same time: the chunk reader, running ordinary SELECTs against current data, and the log reader, following live changes. Take a naive interleaving. The chunk reader SELECTs row 42 and holds value `A`. A user then updates row 42 to `B`, and the log reader emits that update. The chunk reader finally flushes its buffer and emits row 42 as a read event with value `A`. A sink applying events in order now holds `A` — a value that is older than what it had a moment ago. The snapshot has corrupted the destination. Ordering the two producers is not possible in general, because the SELECT ran at some unknown point relative to the log. What *is* possible is to establish two known points in the log and reason about the interval between them. ## The mechanism For every chunk, Debezium performs three steps. First it writes a row into the signalling table with type `snapshot-window-open`. Second it runs the chunk SELECT — the next N rows in primary-key order after the last chunk's high-water key — into an in-memory buffer keyed by primary key. Third it writes `snapshot-window-close` into the signalling table. Both of those writes are ordinary DML on a captured table, so both appear in the transaction log, and the connector sees them come back through its own streaming path. That gives it a precisely delimited interval. Streamed change events are always emitted as they arrive; but for any event arriving **between** the open and close markers whose key is present in the buffer, the connector deletes that key from the buffer. When the close marker arrives, whatever is left in the buffer is emitted as `r` events. The invariant that falls out is simple and worth stating in exactly these terms: **the log always wins**. If a row changed during the window, the destination gets the live change and never the older snapshot read. If it did not change, the destination gets the snapshot read. Either way it gets one value, and it is the fresher one. ## What this buys operationally Because correctness comes from log positions rather than from holding the source still, no table lock is required and the streaming feed is never paused. The snapshot can run for days against a huge table without the source database noticing anything beyond the read load, and the pipeline stays current throughout. It also means the backfill is interruptible: the connector records the last completed chunk's high-water key in its offsets, so a restart resumes at the next chunk rather than re-reading the table. ## Tuning `incremental.snapshot.chunk.size` sets how many rows one chunk reads. Larger chunks mean fewer window round-trips — fewer signalling writes, less per-chunk overhead — but a bigger buffer held in connector memory and a longer window during which more keys can be invalidated. Smaller chunks are gentler on memory and on the source, at the cost of far more signalling writes. This is a real knob to reach for when a backfill is either too slow or pushing the connector's heap. Recent releases also expose a watermarking strategy choice: the default writes an open and a close row, while an alternative writes and then deletes the marker row so the signalling table does not accumulate them. That matters when the signalling table lives in a database whose owners object to unbounded growth or to extra log volume. ## The read-only variant Writing to the source is not always permitted. The MySQL connector supports a read-only incremental snapshot that establishes the same window boundaries using log positions (GTIDs) instead of writing marker rows, combined with signals delivered over Kafka rather than through a source table. Same algorithm, same guarantee, no writes to the source. ## Failure modes to know If the signalling table is not captured, the marker writes never come back through the streaming path and the windows never close — the symptom is a snapshot that appears to start and then produce nothing. If the connector's memory is undersized relative to the chunk size, large chunks of wide rows can pressure the heap. And if a table has no usable primary key, chunking cannot proceed at all unless a surrogate key is supplied, because the whole scheme depends on being able to identify the same row in the buffer and in the log.
- Why write the markers into a source table instead of tracking the window in connector memory?Because the boundaries have to be expressible as positions in the transaction log. Writing a row puts a marker into the log at a definite position, letting the connector classify each streamed event as inside or outside the window. An in-memory timestamp cannot be compared against log order in any sound way.
- What does raising incremental.snapshot.chunk.size actually trade off?Fewer chunks means fewer marker writes and less per-chunk overhead, so the backfill finishes sooner. Against that, each chunk holds more rows in the connector's memory and keeps the window open longer, so more keys can be invalidated and the heap footprint grows. Wide rows amplify both costs, so tune with the row width in mind.
- How does the read-only incremental snapshot on MySQL avoid writing to the source?It derives the window boundaries from log positions — GTID ranges — rather than from marker rows, and takes its signals from a Kafka topic instead of a source table. The deduplication logic is unchanged: events observed inside the window evict the matching keys from the chunk buffer.
saying these in an interview costs you the question
- Thinks the window locks the table while the chunk is read
- Says the snapshot read overwrites a concurrent update
- Believes streaming pauses while a chunk is buffered
- Assumes the markers are only for logging or auditing
- Thinks a restart mid-snapshot restarts the whole table