skip to content

How does a CDC connector hand over from its initial snapshot to log streaming without a gap or a duplicate?

level: middleimportance: must knowfreq 76%

answer

  1. two facts must describe the same instant
  2. record the position before you scan
  3. resume from the pinned position, not from now
  4. the sink upserts by key, so re-delivery is free
  5. overlap is acceptable, a gap never is

basics

~20 s

The connector pins a log position before or at the moment it takes its consistent read of the tables, then resumes streaming from exactly that position. Gaps are unacceptable; overlap is fine because sinks upsert by primary key.

solid answer

~50 s

The seam is made safe by pinning a **log position and a read view to the same instant**. The connector records the current position — an LSN in Postgres, binary-log coordinates or a GTID set in MySQL — and opens a snapshot transaction whose visibility matches it, so the rows it scans reflect exactly the state at that position. When the scan finishes, streaming starts from the pinned position, not from wherever the log has reached by then. Everything committed during the hours the scan ran is therefore replayed on top of the baseline. Some of those replayed events restate rows the snapshot already emitted, which is harmless: the sink applies both as upserts keyed by primary key and the later log event wins. The rule to state out loud is **overlap yes, gap never** — a gap silently loses rows, a duplicate does not.

code

text · 12 lines
text
-- CORRECT: pin first, scan second, resume from the pin
t0  record log position P (LSN / GTID set)
t0  open read view consistent with P
t0..t8  scan tables, emit baseline rows (state as of P)
t8  start streaming from P   <- replays 8 hours of changes

-- WRONG: scan first, then ask where the log is now
t0..t8  scan tables, emit baseline rows (state as of t0)
t8  record log position Q
t8  start streaming from Q
    => every change committed between t0 and t8 is lost,
       silently, with no error and no lag alarm

go deeper

for a junior

Know that the connector writes down where it is in the log before it reads the tables, and that streaming then continues from that written-down place rather than from wherever the log has got to.

for a middle

Explain the ordering argument concretely: pin the position, take a read view matching it, scan, then resume from the pin. Be able to say what is lost if the position is recorded after the scan instead.

for a senior

Demonstrate that you rely on at-least-once delivery plus keyed upserts rather than chasing exactly-once at the seam, and that the real sink-side requirement is per-key ordering — a snapshot row applied after a newer log event moves the row backwards.

for a principal

Own the operational envelope around the seam: offset-store durability, log retention headroom against snapshot duration, and whether to buffer the stream during the scan so the source is not forced to retain segments for the whole backfill.

## Two facts that must refer to the same instant A correct handover needs two things captured together: 1. **A consistent read of the captured tables** — one that is not smeared across the hours the scan takes. 2. **The log position corresponding to that read** — the point from which every later change must be replayed. If these two disagree, you get either a gap (rows changed after the read but before the recorded position, never delivered) or unnecessary re-work. The whole design problem is making them agree. ## Why the naive order fails Scan the tables first, then ask the database for its current log position, and you have lost everything committed during the scan. The scan saw `orders` at 09:00; the position you recorded is from 17:00; streaming starts at 17:00; the eight hours of updates in between are in neither the snapshot nor the stream. Nothing errors. The sink simply holds stale rows for every key touched during the window, forever, until that key is written again. The inverse order — record the position first, then scan — is safe, because the scan sees *at least* the state at that position, and everything after it is replayed. That is why real implementations record first. ## How the position and the read view are tied together **Postgres.** Creating a logical replication slot returns a *consistent point* (an LSN) and can export a snapshot name. A second session opens a `REPEATABLE READ` transaction and calls `SET TRANSACTION SNAPSHOT` with that name, so its reads see exactly the database state at the slot's LSN. Streaming later starts from the slot, which by construction begins at that same LSN. **MySQL.** A brief global read lock quiesces writes just long enough to read the current binary-log coordinates (or the executed GTID set) and open a `REPEATABLE READ` transaction with a consistent snapshot. The lock is then released — it was held for milliseconds — while the InnoDB read view keeps the scan seeing that instant for as long as it runs. In both cases the expensive part (the scan) runs without holding anything that blocks writers, and the cheap part (pinning the pair) is what needed the momentary coordination. ## The alternative: buffer the log during the snapshot Some connectors open the log stream at the pinned position immediately and buffer or interleave events while the snapshot is still running, flushing them once the baseline is emitted. This shortens the catch-up phase and starts consuming the log right away, which keeps the source's log from being retained as long. The correctness argument is identical: the events buffered begin at the pinned position, so nothing between the baseline and the live stream is missing. ## Overlap is the safety margin At the seam you will re-deliver rows: a row changed at 11:00 was read by the snapshot at 09:00 in its old state, and its 11:00 change is also replayed from the log. The sink sees the old version, then the new one. Because the sink applies changes as **idempotent upserts keyed by primary key**, applying them in order lands on the correct final state. The same reasoning makes it safe for a connector to crash mid-snapshot and restart a chunk it partly emitted. This is why every CDC design leans on at-least-once delivery plus idempotent application rather than trying to achieve one-and-only-one delivery at the seam. The one thing overlap cannot rescue is **out-of-order** application: if a sink applies the 09:00 snapshot value *after* the 11:00 log event, the row goes backwards. Ordering per key is what actually has to hold, and it is the sink-side property to defend. ## Resuming after a restart Once streaming is underway the connector periodically commits its current position to a durable offset store. On restart it reads that offset and resumes — the snapshot is not repeated. Two things break this: losing the offset store (or changing the connector's identity so it looks like a new connector), and the stored position aging out of the source's retention. In both cases the connector cannot resume and falls back to re-snapshotting or fails outright, which is why offset durability and retention headroom are operational requirements, not niceties. ## What to say in an interview State the invariant first — *the log position must be pinned no later than the read view, and streaming must resume from the pinned position, not from now* — then the consequence: overlap at the seam is expected and is made harmless by keyed upserts, while a gap is silent and permanent.

  • Why is duplicate delivery at the seam tolerable but a gap is not?
    A duplicate is corrected by the next application of the same key: an upsert keyed by primary key converges on the latest version, so re-delivering a row costs only throughput. A gap produces a row that is permanently wrong until that key is written again, with no error, no lag metric and no failed task to alert on. Silent and permanent beats noisy and self-healing every time.
  • What actually has to hold at the sink for the seam to be safe?
    Per-key ordering and idempotent application. The sink must apply changes for a given primary key in source order and must treat every event as an upsert or delete rather than a blind insert. Global ordering across keys is not required. If ordering per key can be violated — for example by parallel writers on the same key — a late snapshot row can overwrite a newer log event and the row goes backwards.
  • A connector restarts mid-snapshot. What determines whether it resumes or starts over?
    Whether it persisted enough state to identify where it was — the pinned log position plus the chunk boundary it had reached. Designs that checkpoint per chunk resume at the next chunk; designs that only checkpoint at snapshot completion must restart the whole scan. Either is correct, because re-emitted rows are absorbed by keyed upserts; the difference is hours of wasted source I/O.
  • Does the connector have to wait for the snapshot to finish before reading any log at all?
    No. Opening the log stream at the pinned position immediately and buffering or interleaving events while the scan proceeds is a common design. Correctness is unchanged because the buffered events start at the pinned position. The operational benefit is real: the source stops having to retain log segments from the pin onward, which is often the binding constraint on a long snapshot.

It is like photographing a moving crowd and then filming from a tripod: you must start the film from the moment of the photograph, not from whenever you finished developing it, or the people who moved in between are in neither the photo nor the film.

saying these in an interview costs you the question

  • Records the log position after the snapshot finishes instead of before
  • Claims exactly-once delivery is required at the seam
  • Treats duplicates as a correctness bug rather than an expected cost
  • Thinks streaming should resume from the log's current end after a snapshot
  • Forgets that the read view and the log position must describe one instant

context