skip to content

A team implements the outbox pattern using a database polling relay that queries SELECT * FROM outbox WHERE published = false every second. A colleague suggests replacing it with Debezium reading the database's write-ahead log via change data capture instead. What changes, and what are the trade-offs?

level: middleimportance: should knowfreq 60%

answer

  1. polling = repeated SELECT, added query load + latency
  2. CDC = tails WAL/binlog, near-real-time, no extra query load
  3. Debezium = Kafka Connect source connector
  4. CDC needs replication slot / binlog access + ops overhead
  5. both remain at-least-once

basics

~20 s

Polling repeatedly asks the database 'anything new?' which adds load and delay. CDC tools like Debezium instead tap directly into the database's internal change log and stream new rows out in near real time, without hammering the table with queries - but it needs more setup and access to that low-level log.

solid answer

~30 s

A polling relay issues periodic SQL queries against the outbox table, adding query load proportional to poll frequency and introducing latency roughly equal to half the poll interval. Debezium instead attaches to the database's transaction log (Postgres logical replication slot, MySQL binlog) and streams committed row changes as they happen, giving lower latency and near-zero added query load on the primary table. The trade-off is operational: CDC requires enabling and maintaining logical replication/binlog access, deploying and monitoring a Kafka Connect plus Debezium pipeline, and handling connector failover/rebalancing - meaningfully more infrastructure than a simple poller, though both remain at-least-once.

go deeper

for a junior

Should grasp that polling repeatedly asks the database and CDC instead streams changes as they happen, in plain terms.

for a middle

Should compare latency and query-load trade-offs and name Debezium/WAL or binlog as the CDC mechanism.

for a senior

Should discuss operational cost of running a CDC pipeline (Kafka Connect, replication slot management, connector rebalancing) versus a simple poller, and know both remain at-least-once.

for a principal

Should reason about when the added CDC infrastructure investment is justified by scale/latency requirements versus when it's premature complexity, and how CDC choice affects blast radius during database failover or schema migrations.

## The same job, read two different ways Both approaches solve the same problem - getting rows out of the outbox table and onto the broker - but they read that table in fundamentally different ways, and the difference has real consequences for latency, load, and operational surface area. | | Polling relay | Change data capture | |---|---|---| | **Latency** | bounded by the poll interval | as soon as the transaction commits | | **Query load** | a query against the live table on every poll cycle | a stream the database is already producing | | **Operational surface area** | an application process with a loop | a replication slot or binlog access, and a connector | ## How a polling relay reads the table A polling relay is the simplest possible implementation: a loop that periodically issues a SQL query like `SELECT * FROM outbox WHERE published = false ORDER BY id LIMIT 100`, publishes each returned row to the broker, and marks it published (or deletes it) on success. - **Load.** Every poll cycle is a query against the live table, so at high poll frequency (every 100-500ms) this adds sustained read load, and even with a good index on the published/id columns, that's still periodic contention on a table that's also being written to by every incoming request. - **Latency.** Latency is bounded by the poll interval - on average, an event sits in the table for about half the interval before the next poll picks it up, so a 1-second interval means roughly 500ms of added delivery latency on top of everything else, worse in the tail when a poll cycle happens to just miss a fresh insert. ## How change data capture reads it instead Change data capture flips the direction of information flow: instead of the relay asking the database 'anything new?', the database tells the relay 'here's what changed,' by exposing its internal transaction log. - **Postgres** does this via logical replication slots (using an output plugin like `pgoutput` or `wal2json`). - **MySQL** exposes its binlog. `Debezium` is a Kafka Connect source connector that attaches to that log, decodes committed row-level changes, and emits them as Kafka messages essentially as soon as the transaction commits - typically single-digit-to-low-double-digit milliseconds of added latency, and no repeated `SELECT` load on the table at all, since it's reading a stream the database is already producing for replication purposes. ## What that improvement costs The cost of that improvement is operational, not conceptual. Standing up CDC means: 1. enabling logical replication or binlog access on a production database (a configuration change with its own resource and retention implications - unconsumed WAL segments can accumulate if the connector falls behind or goes down, risking disk pressure on the primary), 2. deploying and monitoring a Kafka Connect cluster running the Debezium connector, 3. handling connector restarts and offset management so it resumes from the right point in the log after a crash, and 4. understanding connector rebalancing behavior in a distributed Connect deployment. None of that exists with a polling relay, which is just an application process with a database connection and a loop - trivial to write, deploy, and reason about, at the cost of higher latency and steady query overhead. ## What CDC does not change A point worth being precise about: CDC does not remove the need for the outbox table itself, and it does not upgrade delivery to exactly-once. - **The outbox table.** Teams still capture from a dedicated outbox table rather than business tables directly, because business tables have internal-shape columns and high-frequency unrelated updates that would leak implementation detail and generate noisy, irrelevant change events; the outbox table exists specifically to be a clean, intentional, append-only public event schema, regardless of which relay mechanism reads it. - **Delivery.** CDC-based relays are still at-least-once in practice - connector restarts, offset resets after a rebalance, and downstream retry logic can all produce redelivery, so consumers still need to dedupe. ## Choosing between them The practical decision is a cost-benefit call tied to scale and existing infrastructure. - **A well-indexed poller.** A low-to-moderate volume service where sub-second latency doesn't matter, or a team without an existing Kafka Connect footprint, is often better served by a well-indexed poller at a reasonably short interval - simpler to debug, fewer moving parts, one less thing that can silently stall. - **Debezium.** A team already running Kafka Connect for other services, or one whose downstream consumers genuinely need near-real-time freshness (e.g., search indexing, fraud detection), gets a much better marginal return from adopting Debezium. This mirrors how it's used at real scale: companies like Shopify and Airbnb have documented using Debezium-based CDC on outbox tables specifically to get low-latency, reliable event propagation out of sharded MySQL/Postgres fleets without hand-rolling pollers per service.

  • Does using Debezium/CDC still require an outbox table, or can you CDC the business table directly?
    You still want a dedicated outbox table rather than CDC-ing the business table directly, because business tables have internal-shape columns and frequent unrelated updates that would leak implementation detail and generate noisy change events. The outbox table gives you a stable, intentional public event schema and a clean append-only stream to capture.
  • What happens to already-published outbox rows under a CDC-based relay - do you still need cleanup?
    Yes. CDC captures the insert event once, but the row still sits in the table afterward, so you need the same periodic reaper/TTL cleanup as a polling relay to keep the table from growing unbounded, since CDC doesn't delete rows for you.
  • If polling is simpler, when would a team reasonably stick with it instead of adopting CDC/Debezium?
    Low-to-moderate event volume where sub-second latency doesn't matter, a team without the operational appetite to run Kafka Connect, or an early-stage system where added infra risk outweighs the latency win. Polling with a reasonably short interval and proper indexing is often good enough and much simpler to reason about and debug.

Polling is like refreshing a webpage every few seconds to see if there's news; CDC is like subscribing to the newsroom's internal wire feed that pushes every story to you the instant it's filed - faster and less wasteful, but you need a direct line into the newsroom's infrastructure to get it.

saying these in an interview costs you the question

  • Thinks CDC eliminates the need for an outbox table entirely
  • Believes polling can't work reliably at all
  • Doesn't know CDC reads a log (WAL/binlog) rather than querying the table
  • Assumes CDC gives exactly-once delivery

context