skip to content

Why can a CDC consumer see an order row before its order_items, and what fixes it?

level: seniorimportance: should knowfreq 42%

answer

  1. the source committed both tables at once
  2. capture emits per row, not per transaction
  3. the two tables travel separate routes
  4. the wrong state repairs itself shortly after
  5. markers carry a transaction id and counts

basics

~20 s

Capture emits one event per row change, so a single source transaction that wrote both tables becomes independent events on independent routes. Atomicity stops at the source; the consumer sees an intermediate state that converges seconds later.

solid answer

~60 s

The source committed both tables atomically, but CDC delivers row changes, not transactions: the order insert and the item inserts are separate events, keyed by different rows, usually routed and applied independently. So a reader can observe an order with no items — briefly, and self-correcting. Four responses, in order of how often they are the right one. **Tolerate it**: make consumers eventually consistent, and never join across CDC-fed tables in a way that treats a missing child as a business fact. **Buffer by transaction**: some connectors emit begin/end markers carrying the transaction id and per-table event counts, so a consumer can hold events until it has the whole transaction — correct, but it costs memory proportional to the largest transaction, adds latency to the end marker, and serialises what you had parallelised. **Apply to a committed watermark**: let the sink ingest freely but expose a view only up to the highest fully-applied commit position. **Change the shape at the source**: publish one business event carrying the whole aggregate — the outbox pattern — so there is nothing to reassemble.

code

text · 5 lines
text
tx 8841 BEGIN
  orders       id=500  op=insert   total=90.00
  order_items  id=900  op=insert   order_id=500
  order_items  id=901  op=insert   order_id=500
tx 8841 END   (event counts: orders=1, order_items=2)

go deeper

for a junior

Understand that change events describe single row changes, so two tables written by one source transaction arrive as separate events and can land a moment apart in the target.

for a middle

Explain why keying events per row is what splits a transaction, and that the resulting partial state is transient. Describe what transaction begin and end markers carry and what a consumer could do with them.

for a senior

Weigh the options out loud: tolerate and converge, buffer by transaction id, or expose a consistent cut at a committed watermark. Quantify the buffering cost for a bulk transaction and name the monitoring the watermark needs.

for a principal

Decide what consistency the platform promises consumers — row-level convergence by default, transaction-level only where a consumer cannot act on partial state — and push producers toward publishing whole business events where the shape is the real problem.

## Where the atomicity is lost A source transaction is atomic in the database: either both the `orders` row and its `order_items` rows are visible, or neither is. A change-data-capture pipeline does not carry that property forward. It reads the log and emits one event per *row* change, and those events are then keyed by their own row's primary key so that each row's history stays ordered. Keying by row is exactly what makes per-key ordering work, and it is exactly what splits a transaction: the order's event and the items' events now travel independent routes, are batched independently, and are applied independently. A consumer reading them can therefore observe a state the source never had — an order with no lines, or a payment with no corresponding order. The important qualifier is that this state is **transient**. The missing events are already in flight, and within a batch interval or two the sink converges on the source's state. The bug is never "the data is wrong forever"; it is "a reader who looks at the wrong moment draws a wrong conclusion". ## Response 1: tolerate it, deliberately For most analytical consumers this is the correct answer and the cheap one. Batch models run on a schedule, over a target that has long since converged; the window of inconsistency is smaller than the scheduling interval and nobody sees it. What you must do is make sure no consumer treats an intermediate state as a business fact: an alert that fires on "orders with zero line items" will page someone every few minutes forever. Give such checks a grace period, or run them against a consistent cut rather than the live table. ## Response 2: transaction-boundary markers Some connectors can emit metadata around each transaction: a begin marker, an end marker, and on the end marker the transaction identifier plus a count of events emitted per table. Each data event also carries the transaction id and its ordinal within the transaction. That is enough for a consumer to reassemble atomicity — buffer everything for a transaction id, wait for the end marker, check the counts, then apply the whole set in one sink transaction. ```text tx 8841 BEGIN orders id=500 op=insert order_items id=900 op=insert order_id=500 order_items id=901 op=insert order_id=500 tx 8841 END (event counts: orders=1, order_items=2) ``` The costs are real and are what the interview is actually probing. Buffering needs memory proportional to the largest transaction, and a bulk job that updates ten million rows in one transaction will exhaust it. Latency now waits for the end marker rather than the event. And because the consumer must see a transaction's events together, you have serialised work that per-key routing had parallelised — the throughput ceiling drops. Reach for this when a downstream system genuinely cannot show partial state, not by default. ## Response 3: apply freely, expose a consistent cut A middle path that keeps parallelism: let the sink apply events as they arrive, but track the highest source commit position for which *all* events are known to be applied — a committed watermark — and expose only data up to that position to readers, through a view or a published snapshot marker. Consumers see a consistent point-in-time state that lags slightly, while the loader stays fast. The engineering cost is computing the watermark honestly across all lanes: it is the minimum of the per-lane progress, not the maximum, and one stalled lane holds the whole cut back, which is a monitoring requirement in itself. ## Response 4: fix the shape at the source If the consumer needs a business event rather than a set of row changes, the durable answer is for the producing service to write one record describing the whole change — the outbox pattern, where the application inserts an event row inside the same transaction as its business writes and CDC captures that single row. There is then nothing to reassemble, because atomicity was preserved by construction, and the consumer is decoupled from the source's table layout. This is a source-side design change, so it is available only when you own the producing application. ## Choosing between them Ask what the consumer does with the data. A dashboard refreshed hourly: tolerate. A downstream service that emits a customer-facing notification on seeing an order: it must not act on a partial state, so either buffer by transaction or move it to an outbox event. A replica used for operational reads: expose a consistent cut. A fraud rule that joins three CDC-fed tables in real time: this is the case where partial state produces false positives, and it is the strongest argument for transaction-level delivery. ## What a strong answer includes Name the cause precisely — row-level events keyed per row, so transaction atomicity ends at capture — say that the inconsistency is transient rather than permanent, and then give at least two responses with their costs rather than jumping straight to "buffer the transaction". Volunteering the memory and latency cost of buffering large transactions is what separates someone who has run this from someone who has read about it.

  • What does buffering a whole transaction before applying it cost you?
    Memory proportional to the largest transaction — a bulk job touching millions of rows in one transaction will blow the buffer — plus latency, since nothing applies until the end marker arrives, and throughput, since events that per-key routing had parallelised must now be gathered and applied together. Reach for it only where partial state is genuinely unacceptable.
  • How would you expose a consistent point-in-time view without buffering transactions?
    Let the loader apply events as they arrive, but track a committed watermark: the highest source commit position for which every lane has applied everything. Serve readers through a view or snapshot pinned to that position. Compute it as the minimum of per-lane progress, and monitor it, because one stalled lane freezes the whole cut.
  • An alert fires constantly on 'orders with zero line items' in the CDC-fed target. What do you change?
    The alert, first: it is observing a transient state that exists between two lanes applying, so give it a grace period or evaluate it against a consistent cut rather than the live tables. Only if genuine orphans persist beyond that window is there a real pipeline problem worth investigating.

saying these in an interview costs you the question

  • Assumes CDC preserves source transaction atomicity downstream
  • Treats a briefly missing child row as permanent data loss
  • Proposes buffering every transaction without mentioning memory or latency
  • Thinks ordering all events by commit timestamp restores atomicity
  • Joins CDC-fed tables in real time and alerts on partial state

context