skip to content

Walk through, step by step, how the transactional outbox pattern uses an outbox table plus a relay/poller process to publish an OrderPlaced event reliably after an order is saved.

level: middleimportance: must knowfreq 80%

answer

  1. outbox table in same TX as business write
  2. relay/poller reads + publishes + marks done
  3. at-least-once, not exactly-once
  4. consumers must dedupe/be idempotent
  5. reaper job cleans up old rows

basics

~20 s

The service writes both the order row and a row describing the event into the same database transaction. A separate background process then reads the new outbox rows, sends them to the message broker, and marks them as sent - so the event is only ever 'in flight' if the original database write actually succeeded.

solid answer

~40 s

In the same local ACID transaction as the business write (INSERT INTO orders), the service also INSERTs a row into an outbox table containing the event type, payload, and an unpublished flag. Because it's one transaction, both rows commit together or neither does - eliminating the dual-write risk. A relay (a polling job, or a CDC connector like Debezium reading the DB's write-ahead log/binlog) picks up new outbox rows, publishes them to the broker, and then marks them published or deletes them. Delivery is at-least-once: the relay may crash after publishing but before marking-published, causing a duplicate on restart, so downstream consumers must be idempotent.

go deeper

for a junior

Should describe that the event gets written to a table alongside the business data, and that 'something' later sends it - detail on relay internals not required.

for a middle

Should describe the same-transaction insert, name the relay/poller as a separate process, and know delivery is at-least-once requiring idempotent consumers.

for a senior

Should discuss relay implementation choices (polling vs CDC), per-row ordering guarantees, and outbox table cleanup/retention.

for a principal

Should reason about relay throughput/lag under load, how to scale the relay safely (single-writer vs partitioned polling to avoid duplicate publishes), and the operational cost of an extra moving part in the deployment topology.

## Two halves, one database The transactional outbox pattern solves the dual-write problem by refusing to treat 'save state' and 'publish event' as two independent writes in the first place. Instead, it collapses the atomicity requirement down to a single system - the application's own database - which already gives you **ACID** transactions for free. The mechanism has two distinct halves: - a **synchronous write half**, and - an **asynchronous delivery half**, and understanding why they're split is the key to understanding the whole pattern. ## The write half The write half happens inside the original request. When the order service handles 'place order,' it opens one database transaction and issues two `INSERT`s: one into the `orders` table with the business data, and one into an `outbox` table containing enough to describe the event - typically - an `id` (often a UUID generated at insert time, useful later for consumer deduplication), - an `aggregate_id` (the order id, used for partitioning/ordering), - an `event_type` ('OrderPlaced'), - a JSON payload, - a `created_at` timestamp, and - a status or published flag. Both `INSERT`s commit or roll back together, because they're the same transaction. This is the entire trick: there is no longer a window where the order exists but no durable record exists anywhere that an event needs to go out, because the 'record that an event needs to go out' lives in the same table, guarded by the same commit, as the business fact itself. ## The delivery half The delivery half is deliberately decoupled from the request and runs as an independent process - **the relay**. 1. **In its simplest form** this is a polling loop: every N milliseconds it runs something like `SELECT * FROM outbox WHERE published = false ORDER BY id`, sends each row to the broker (Kafka, RabbitMQ, SNS), and on successful publish updates that row's status to published (or deletes it). 2. **Change data capture.** A more sophisticated relay uses change data capture instead of polling: a tool like `Debezium` attaches to the database's write-ahead log or binlog and streams newly-committed outbox rows out with lower latency and no repeated query load (this trade-off is its own leaf topic). Either way, the relay is what actually talks to the broker; the original request thread never does, and therefore the original request's latency and success/failure no longer depend on the broker being reachable at all. ## Why the split is what makes it work This split is exactly why the pattern works: the risky, two-system operation (talk to the broker) has been moved entirely out of the transactional boundary and into a component whose job is specifically to retry forever against a durable backlog. If the relay crashes mid-publish, it just resumes from wherever the unpublished rows still are - nothing is lost, because 'needs to be published' is a row in a table, not an in-memory retry counter. The cost of this reliability is that delivery becomes **at-least-once** rather than exactly-once: if the relay publishes successfully to the broker but crashes before writing back the published status, on restart it sees the row as still unpublished and sends it again. Every consumer of outbox-relayed events must therefore be **idempotent** - either - deduplicating by the event's stable `id`, or - structuring the downstream write as an upsert rather than an append-only insert. ## Housekeeping the table An important practical detail often missed: the outbox table needs housekeeping. Rows accumulate indefinitely if nothing removes them, so production systems run a periodic **reaper** that deletes or archives rows past a retention window once they're confirmed published (e.g., anything older than a few days), keeping the table small and the polling query's index scan cheap. Debezium-based teams often use a Kafka Connect single-message-transform to also emit a tombstone or simply let the reaper handle cleanup separately, since CDC captures the insert event but does not remove the row itself. ## Where it shows up A well-known concrete instance of this exact shape appears in Debezium's own documentation and in Chris Richardson's microservices.io write-up of the pattern: an e-commerce order service writes to `orders` and `outbox` in one transaction, and a Debezium connector on the outbox table streams events into Kafka, from which a shipping service, a notification service, and an analytics pipeline each independently consume - none of them ever query the order service's database directly, and none of them can see an event for an order that didn't actually commit.

  • How does the relay know which outbox rows are new without re-scanning the whole table every time?
    It typically tracks a monotonically increasing column - an auto-increment id or a sequence - and queries WHERE id > last_processed_id ORDER BY id, or WHERE published = false with an index on that column. It persists the last processed offset (its own state store, or an update on the outbox row itself) so a poller restart resumes roughly where it left off rather than rescanning everything.
  • What happens to the outbox table over time, and how do teams manage that?
    Rows accumulate indefinitely if never cleaned up, bloating the table and slowing index scans used by the poller. Teams typically run a periodic reaper that deletes or archives rows older than a retention window (e.g., a few days) once they're confirmed published, or track a published_at timestamp with a TTL-based cleanup job.
  • Why must the outbox insert and the business-entity insert happen in the exact same database transaction?
    If they're two separate transactions, you've just moved the dual-write problem one level down - now it's 'commit the order' vs 'commit the outbox row' instead of 'commit the order' vs 'publish to Kafka,' with the same crash-in-between risk. The whole point of the pattern is that a single local ACID transaction is the only atomicity boundary you get for free.

Like a restaurant kitchen writing every completed dish onto a pass-through ticket rail in the same motion as plating it - a runner (the relay) continuously grabs tickets off the rail and carries them to the dining room, so a dish is never plated without eventually being announced, and if the runner drops a ticket and re-grabs it, the diner just hears the order called twice, not zero times.

saying these in an interview costs you the question

  • Describes the outbox insert happening in a separate transaction from the business write
  • Doesn't mention the relay/poller step at all - thinks the outbox table alone solves publishing
  • Assumes exactly-once delivery to the broker
  • No mention of what marks a row as published or how re-delivery is handled

context