skip to content

Compare a polling-publisher relay versus a CDC/Debezium relay for moving outbox rows to Kafka. What are the trade-offs?

level: middleimportance: must knowfreq 70%

answer

  1. Polling: SELECT unsent, FOR UPDATE SKIP LOCKED, mark/delete
  2. CDC: tail WAL/binlog via Debezium on Kafka Connect
  3. Debezium Outbox Event Router SMT reshapes the row
  4. Polling = simple but latency + DB load
  5. CDC = real-time, commit-order, but heavier ops + replication slot

basics

~20 s

A polling publisher repeatedly queries the outbox table for new rows and sends them to Kafka. CDC (e.g. Debezium) instead tails the DB's transaction log and streams outbox inserts with no polling. CDC is lower-latency and lower DB load; polling is simpler and needs no extra infrastructure.

solid answer

~50 s

Both relays move outbox rows to Kafka but differ in mechanism. A **polling publisher** is a job that periodically `SELECT`s unsent rows (often `ORDER BY id LIMIT n FOR UPDATE SKIP LOCKED`), publishes them, and marks/deletes them. It's trivial to operate and needs no extra components, but it adds query load, has latency bounded by the poll interval, and you must manage marking, contention, and table growth. **CDC with Debezium** tails the database write-ahead log (Postgres WAL, MySQL binlog), so new outbox inserts are streamed to Kafka in near real time with no polling queries and naturally in commit order. Debezium even has an **Outbox Event Router** SMT that reshapes the change event into a clean domain message. The cost is operational: you run Kafka Connect + Debezium, configure logical replication/binlog, and handle connector failover and schema/log-position management. Rule of thumb: polling for simple/low-throughput systems; CDC when you need low latency, low DB overhead, and strict ordering at scale.

go deeper

for a junior

Know that polling repeatedly queries the table, while CDC/Debezium tails the DB log and is faster.

for a middle

Compare latency, DB load, ordering, and infra; know SKIP LOCKED for pollers and the Outbox Event Router SMT.

for a senior

Reason about ordering guarantees, replication-slot retention risks, and when each relay fits the throughput/latency budget.

for a principal

Decide relay strategy across teams: standardize on CDC at scale vs polling for simple services, accounting for ops maturity and failure domains.

## Two ways to get rows out of the outbox The outbox table guarantees events are recorded atomically with state changes, but something must publish those rows to **Kafka**. That something is the **relay** (a.k.a. message relay / publisher). There are two dominant designs. ### 1. Polling publisher A background worker runs on a timer: ```sql SELECT * FROM outbox WHERE published = false ORDER BY id LIMIT 100 FOR UPDATE SKIP LOCKED; -- lets multiple workers run without colliding ``` For each row it publishes to Kafka, then marks it published (`UPDATE ... SET published = true`) or deletes it. **Pros**: dead simple; no extra infrastructure; works with any database; easy to reason about and test. **Cons**: - **Latency** is bounded by the poll interval (e.g. up to 1s if you poll every second). - **DB load**: constant queries, plus index maintenance on `published`/`created_at`. - **Ordering** must be enforced by you (e.g. order by a monotonic id, single-threaded per key, or partition-aware claiming) — naive multi-worker polling can reorder. - **Table growth / cleanup**: you must purge or archive sent rows. - **Contention**: multiple pollers need `SKIP LOCKED` or partitioning to avoid double-claiming. ### 2. CDC (Change Data Capture) with Debezium **CDC** reads the database's **transaction log** — the internal, ordered record of every committed change (Postgres **WAL** via logical replication, MySQL **binlog**). **Debezium** is a CDC platform that runs as a **Kafka Connect** connector and emits a Kafka message for every insert into the outbox table, in commit order. **Pros**: - **Near-real-time**: no poll interval; events flow as soon as the transaction commits and the log record is read. - **No query load** on the table; reading the log is cheap and doesn't compete with OLTP. - **Ordering for free**: the log is the commit order; per-aggregate ordering is preserved when you route by a key. - **Outbox Event Router SMT**: Debezium's built-in Single Message Transform maps outbox columns (aggregate id, type, payload) into a clean topic/key/value, so consumers see a domain event, not a raw row-change. **Cons**: - **Operational weight**: you run Kafka Connect + Debezium, enable logical replication/binlog, manage connector offsets (log positions), failover, and slot/retention (e.g. an unconsumed Postgres replication slot can pin WAL and fill disk). - **Heavier to test/develop** than a polling loop. - Schema/log-format coupling to the specific database. ### Choosing | Concern | Polling | CDC/Debezium | |---|---|---| | Latency | poll interval | near real-time | | DB overhead | queries + index churn | log read only | | Ordering | you enforce it | commit order, natural | | Infra needed | none | Connect + Debezium | | Cleanup | you purge rows | rows can be tombstoned/short-lived | | Operational risk | low | replication slot/binlog management | Both are **at-least-once** (a relay crash mid-publish can resend), so consumers must be idempotent regardless of which you pick.

  • What is the Debezium Outbox Event Router and why is it useful?
    It's a Single Message Transform (SMT) that maps raw outbox-table change events into clean domain messages: routing to a topic by the aggregate type, setting the Kafka key from the aggregate id (preserving per-aggregate ordering), and using the payload column as the message value. It saves consumers from parsing raw CDC row-change envelopes.
  • What operational hazard is specific to Postgres logical-replication CDC?
    An unconsumed or lagging replication slot pins the WAL: Postgres can't recycle those log segments, so disk fills up and the database can stall. You must monitor slot lag and ensure the connector keeps consuming (or drop stale slots).

saying these in an interview costs you the question

  • Saying CDC eliminates duplicates / gives exactly-once (still at-least-once).
  • Claiming polling cannot preserve ordering at all (it can, with a monotonic id and key-aware claiming).
  • Forgetting that multi-worker polling needs SKIP LOCKED or partitioning to avoid double-publishing.
  • Ignoring replication-slot/binlog retention as an operational risk of CDC.
  • Believing Debezium reads the outbox table by querying it (it reads the transaction log, not the table).

context