skip to content

Transactional Outbox Pattern

Writing to the database and an outbox table in one transaction, then relaying to Kafka by polling or CDC. The standard answer to the dual-write problem, and a very frequent design question.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

6

What is the Transactional Outbox pattern, and what problem does it solve when you need to update a database and publish a Kafka message in one operation?

level: juniorimportance: must knowfreq 78%

answer

  1. Dual write = two systems, no shared transaction
  2. One DB txn: business row + outbox row
  3. Relay reads outbox -> Kafka
  4. At-least-once -> consumers idempotent
  5. Kafka txn can't enlist your DB

basics

~20 s

Instead of writing to the database AND sending to Kafka separately (which can half-fail), you write the business row and an 'outbox' row in the same DB transaction. A separate process later reads the outbox and publishes to Kafka.

solid answer

~40 s

The Transactional Outbox solves the dual-write problem: a service that must both persist state and emit an event has no single transaction spanning the database and Kafka, so one can succeed while the other fails, leaving them inconsistent. The pattern makes the write atomic by inserting the event into an 'outbox' table within the same local DB transaction as the business change. Because both rows commit or roll back together, the event is durably recorded if and only if the state change happened. A separate relay (a polling publisher or a CDC tool like Debezium) then reads outbox rows and publishes them to Kafka, marking or deleting them once sent. This guarantees the event is eventually published at least once, decoupling the atomic local commit from the unreliable remote send.

go deeper

for a junior

Know the one-liner: write business row + outbox row in one DB transaction, a separate process publishes to Kafka.

for a middle

Explain the dual-write failure modes (lost event vs phantom event) and why a shared local transaction fixes them.

for a senior

Articulate the at-least-once consequence and that consumer idempotency is mandatory; distinguish polling vs CDC relays.

for a principal

Frame it as decoupling atomic local commit from unreliable remote delivery, and discuss when to prefer outbox vs alternatives (event sourcing, listen-to-yourself).

## The problem: dual writes A common requirement is: when something happens in a service, (1) save the new state to a database, and (2) tell other services by publishing a message to **Kafka** (a distributed event log). The naive approach does these as two separate calls: ``` db.save(order) // call 1 kafka.send(orderEvent) // call 2 ``` This is a **dual write**: two independent systems updated by two independent operations with no shared transaction. There is no way to make both happen atomically, so failures leave you inconsistent: - DB commit succeeds, then the app crashes before `kafka.send` -> state changed but **no event** (downstream never learns). - `kafka.send` succeeds, then the DB transaction rolls back -> event published for a change that **never happened** (a phantom). Kafka transactions don't help here because they are internal to Kafka (atomic across Kafka topics/offsets); they cannot enlist your relational database. ## The solution: an outbox table The **Transactional Outbox** turns the dual write into a single local write. You create an `outbox` table in the **same database** as your business data. When you change state, you insert the event row into `outbox` **inside the same database transaction**: ``` BEGIN; INSERT INTO orders (...); INSERT INTO outbox (id, aggregate_id, type, payload, created_at) VALUES (...); COMMIT; ``` Because both inserts share one ACID transaction, they commit together or roll back together. After commit, the event is **guaranteed to be durably recorded** exactly when the business change is. No phantom events, no lost events. ## Getting it to Kafka: the relay The outbox row isn't in Kafka yet. A separate **message relay** moves outbox rows to Kafka: - **Polling publisher**: a background job periodically `SELECT`s unsent outbox rows, publishes each to Kafka, and marks them sent (or deletes them). - **CDC (Change Data Capture)**, e.g. **Debezium**: tails the database's transaction log (WAL/binlog) and streams new outbox inserts to Kafka with no polling. The relay can crash and retry, so it publishes **at least once** — a row may be sent more than once. Consumers must therefore be **idempotent** (handle duplicates safely). What you get end-to-end is *atomic local commit + at-least-once delivery*. ## Key terms - **Atomic**: all-or-nothing. - **Dual write**: updating two systems without a shared transaction. - **At-least-once**: every event is delivered, possibly more than once. - **Idempotent**: applying the same event twice has the same effect as once.

  • Why can't you just use a Kafka transaction to fix the dual-write problem?
    Kafka transactions are atomic only across Kafka resources (topic-partitions and consumer offsets). They cannot enlist an external relational database, so a DB write and a Kafka send still span two systems with no common commit.
  • Does the outbox give exactly-once delivery to Kafka?
    No. It guarantees the event is recorded atomically with the state change and published at least once. Duplicates are still possible (relay retries), so consumers must be idempotent. Exactly-once end-to-end requires dedup/idempotency on top.

saying these in an interview costs you the question

  • Claiming the outbox gives exactly-once delivery by itself (it's at-least-once).
  • Saying a Kafka transaction can make the DB write and Kafka send atomic.
  • Putting the outbox table in a different database than the business data (then it's still a dual write).
  • Thinking the outbox row IS the Kafka message — it still needs a relay to publish it.

context

open as a page

Outbox delivery is at-least-once, so consumers may see duplicate Kafka messages. How do you make a consumer idempotent?

level: middleimportance: must knowfreq 68%

basics

~20 s

Give each event a unique id. The consumer records processed ids (e.g. in a 'processed_messages' table) inside the same transaction that applies the effect, and skips any id it has already seen. That way reprocessing a duplicate does nothing.

open as a page

Compare a polling-publisher relay versus a CDC/Debezium relay for moving outbox rows to Kafka. What are the trade-offs?

level: middleimportance: must knowfreq 70%

basics

~20 s

A polling publisher repeatedly queries the outbox table for new rows and sends them to Kafka. CDC (e.g. Debezium) instead tails the DB's transaction log and streams outbox inserts with no polling. CDC is lower-latency and lower DB load; polling is simpler and needs no extra infrastructure.

open as a page

What columns belong in a well-designed outbox table, and what does each one enable?

level: juniorimportance: should knowfreq 48%

basics

~20 s

Typical columns: a unique id (UUID), aggregate type (which topic), aggregate id (the Kafka key/ordering), event type, payload (the message body, often JSON), a timestamp, and optionally a published flag or sequence. These let the relay route, key, order, and dedup events.

open as a page

How do you preserve per-aggregate message ordering when publishing outbox events to Kafka, and where can ordering break?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Kafka only orders messages within a single partition. Use the aggregate id (e.g. order id) as the Kafka message key so all events for one entity land on the same partition in order. Publish them in commit/sequence order from the relay.

open as a page

When would you choose the Transactional Outbox over alternatives like Kafka transactions, two-phase commit (XA), or the listen-to-yourself pattern? What are the systemic trade-offs?

level: principalimportance: should knowfreq 38%

basics

~20 s

Use the outbox when you must update a relational DB and publish to Kafka atomically without distributed transactions. Kafka transactions only cover Kafka, XA/2PC is slow and fragile across DB+broker, and listen-to-yourself changes your read model's source of truth. The outbox trades simplicity for at-least-once delivery and a relay to operate.

open as a page