What is the Transactional Outbox pattern, and what problem does it solve when you need to update a database and publish a Kafka message in one operation?
answer
- Dual write = two systems, no shared transaction
- One DB txn: business row + outbox row
- Relay reads outbox -> Kafka
- At-least-once -> consumers idempotent
- Kafka txn can't enlist your DB
basics
~20 sInstead of writing to the database AND sending to Kafka separately (which can half-fail), you write the business row and an 'outbox' row in the same DB transaction. A separate process later reads the outbox and publishes to Kafka.
solid answer
~40 sThe Transactional Outbox solves the dual-write problem: a service that must both persist state and emit an event has no single transaction spanning the database and Kafka, so one can succeed while the other fails, leaving them inconsistent. The pattern makes the write atomic by inserting the event into an 'outbox' table within the same local DB transaction as the business change. Because both rows commit or roll back together, the event is durably recorded if and only if the state change happened. A separate relay (a polling publisher or a CDC tool like Debezium) then reads outbox rows and publishes them to Kafka, marking or deleting them once sent. This guarantees the event is eventually published at least once, decoupling the atomic local commit from the unreliable remote send.
go deeper
Know the one-liner: write business row + outbox row in one DB transaction, a separate process publishes to Kafka.
Explain the dual-write failure modes (lost event vs phantom event) and why a shared local transaction fixes them.
Articulate the at-least-once consequence and that consumer idempotency is mandatory; distinguish polling vs CDC relays.
Frame it as decoupling atomic local commit from unreliable remote delivery, and discuss when to prefer outbox vs alternatives (event sourcing, listen-to-yourself).
## The problem: dual writes A common requirement is: when something happens in a service, (1) save the new state to a database, and (2) tell other services by publishing a message to **Kafka** (a distributed event log). The naive approach does these as two separate calls: ``` db.save(order) // call 1 kafka.send(orderEvent) // call 2 ``` This is a **dual write**: two independent systems updated by two independent operations with no shared transaction. There is no way to make both happen atomically, so failures leave you inconsistent: - DB commit succeeds, then the app crashes before `kafka.send` -> state changed but **no event** (downstream never learns). - `kafka.send` succeeds, then the DB transaction rolls back -> event published for a change that **never happened** (a phantom). Kafka transactions don't help here because they are internal to Kafka (atomic across Kafka topics/offsets); they cannot enlist your relational database. ## The solution: an outbox table The **Transactional Outbox** turns the dual write into a single local write. You create an `outbox` table in the **same database** as your business data. When you change state, you insert the event row into `outbox` **inside the same database transaction**: ``` BEGIN; INSERT INTO orders (...); INSERT INTO outbox (id, aggregate_id, type, payload, created_at) VALUES (...); COMMIT; ``` Because both inserts share one ACID transaction, they commit together or roll back together. After commit, the event is **guaranteed to be durably recorded** exactly when the business change is. No phantom events, no lost events. ## Getting it to Kafka: the relay The outbox row isn't in Kafka yet. A separate **message relay** moves outbox rows to Kafka: - **Polling publisher**: a background job periodically `SELECT`s unsent outbox rows, publishes each to Kafka, and marks them sent (or deletes them). - **CDC (Change Data Capture)**, e.g. **Debezium**: tails the database's transaction log (WAL/binlog) and streams new outbox inserts to Kafka with no polling. The relay can crash and retry, so it publishes **at least once** — a row may be sent more than once. Consumers must therefore be **idempotent** (handle duplicates safely). What you get end-to-end is *atomic local commit + at-least-once delivery*. ## Key terms - **Atomic**: all-or-nothing. - **Dual write**: updating two systems without a shared transaction. - **At-least-once**: every event is delivered, possibly more than once. - **Idempotent**: applying the same event twice has the same effect as once.
- Why can't you just use a Kafka transaction to fix the dual-write problem?Kafka transactions are atomic only across Kafka resources (topic-partitions and consumer offsets). They cannot enlist an external relational database, so a DB write and a Kafka send still span two systems with no common commit.
- Does the outbox give exactly-once delivery to Kafka?No. It guarantees the event is recorded atomically with the state change and published at least once. Duplicates are still possible (relay retries), so consumers must be idempotent. Exactly-once end-to-end requires dedup/idempotency on top.
saying these in an interview costs you the question
- Claiming the outbox gives exactly-once delivery by itself (it's at-least-once).
- Saying a Kafka transaction can make the DB write and Kafka send atomic.
- Putting the outbox table in a different database than the business data (then it's still a dual write).
- Thinking the outbox row IS the Kafka message — it still needs a relay to publish it.