A service updates its own database and then needs to publish an event so the next saga step can run. What is the 'dual write' problem this creates, and how does the transactional outbox pattern solve it?
answer
- dual write = DB commit + broker publish aren't atomic
- outbox row in SAME db transaction as business row
- relay: poll table OR CDC (Debezium tailing WAL)
- at-least-once delivery -> consumers need idempotency
- purge sent outbox rows
basics
~20 sIf a service saves to its database and separately sends a message, one can succeed while the other fails, leaving things out of sync. The outbox pattern saves the message in the same database transaction as the data change, then a separate process reliably delivers it, so both always happen together or neither does.
solid answer
~60 sThe dual-write problem: a service commits a DB transaction and then calls the message broker to publish an event — these are two independent operations with no shared transaction, so a crash or broker outage between them leaves the DB updated but the event never sent, or, if publish happens first, the event sent but the DB write later fails and rolls back, silently breaking the saga. The outbox pattern fixes this by writing the event as a row in an 'outbox' table in the same local database transaction as the business data change — since it's the same transaction, both commit or neither does, with the database's own ACID guarantee doing the work. A separate relay process — either polling the outbox table or using change-data-capture (e.g., Debezium tailing the DB's write-ahead log) — reads new outbox rows and publishes them to the broker, then marks them sent. This gives at-least-once delivery of events that are guaranteed to correspond to a committed state change, at the cost of an extra table, a relay component, and downstream consumers needing to handle duplicate delivery (idempotency).
go deeper
Should recognize that saving to a database and sending a message are two separate actions that can fail independently, and that the outbox pattern's basic idea is to save the message alongside the data.
Should explain that the outbox row is written in the same local transaction as the business row, and name at least one way the event gets relayed to the broker (polling).
Should discuss CDC-based relay (e.g., Debezium) as the more robust alternative to polling, and connect at-least-once delivery to the need for idempotent consumers.
Should discuss operational concerns — outbox table growth/purging, ordering guarantees under multiple relay instances, and how to audit outbox vs. business tables to prove no dual-write gap exists in production.
## The dual-write problem The problem the outbox pattern solves starts from a very ordinary piece of code: a service handles a request, writes a row to its database, and then publishes an event to a message broker so downstream saga participants know to proceed. The trouble is that "write to the database" and "publish to the broker" are two entirely separate systems with no shared transaction manager between them. Either way, there is a gap: - **Commit first, publish second.** If the code commits the database write first and then the process crashes, or the broker is unreachable, before the publish call completes, the database now reflects a change that nobody downstream will ever hear about — the saga silently stalls. - **Publish first, commit second.** If the code publishes first and the subsequent database commit then fails or the process crashes before committing, downstream services react to an event describing a state change that never actually happened. Retrying doesn't fully fix this either: if you publish first and then write, and the write fails, you can't simply un-publish the event once other services may have already consumed it. This is the **"dual write"** problem — two writes to two independent systems that need to be atomic together but have no way to be, because a two-phase commit across a database and a message broker is exactly the kind of distributed transaction the saga pattern exists to avoid in the first place. ## What the outbox makes atomic The **transactional outbox pattern** resolves this by changing what actually gets committed atomically. Instead of writing to the business table and calling the broker, the service writes to the business table and inserts a row into a second table — the **outbox** — in the same local database transaction. The outbox row typically contains: - the event type; - a serialized payload; - metadata like an aggregate ID and timestamp. Because both writes are part of one ACID transaction against a single database, the database's own commit/rollback guarantee makes them atomic: either the business row and the outbox row are both persisted, or neither is. This turns a cross-system atomicity problem into an ordinary single-database transaction, which every relational (and many NoSQL) database already knows how to do reliably. ## Getting the event onto the broker Getting the event out of the outbox table and onto the broker is a second, separate concern, handled by a **relay** process. There are two common implementations. | Relay | Mechanism | Trade-off | |---|---|---| | Polling — the simpler one | a background job periodically queries the outbox table for unsent rows (ordered by insertion time or an autoincrement ID), publishes each to the broker, and marks it sent or deletes it | cheap to build, but it adds polling latency and load on the table | | Change-data-capture (CDC) — the more robust one | a tool like Debezium tails the database's write-ahead log (the same mechanism the database uses internally for replication) and streams new outbox rows to the broker without querying the table at all | giving lower latency and less load, at the cost of an extra infrastructure component (Kafka Connect plus Debezium, or an equivalent) that itself needs to be operated and monitored | ## At-least-once, not exactly-once The trade-off this introduces is that the relay only guarantees at-least-once delivery, not exactly-once: if the relay crashes after publishing to the broker but before marking the outbox row as sent, it will republish the same event on restart. Consumers of these events must therefore be idempotent — able to safely process the same event twice without double-effects, typically by tracking already-processed event IDs (this is where the companion **inbox pattern** comes in on the consumer side). The outbox table also needs housekeeping — old sent rows must be purged or archived — or it grows unbounded. ## Failure modes Failure modes show up as either: - **"ghost" delays** — a stuck relay means events pile up in the outbox and downstream saga steps stall, even though the source data is correct; - **duplicate processing** — a relay retry after a network blip between it and the broker causes a consumer to see the same event twice, and if that consumer isn't idempotent, it double-charges or double-reserves. Because the outbox row and the business row commit together, an operator can always audit the outbox table against the business table to confirm no dual-write gap exists — that auditability is one of the pattern's practical benefits beyond just correctness. ## A reference implementation A concrete widely used example is Debezium's outbox event router, purpose-built to read outbox rows from services (written with plain JDBC/JPA) and republish them onto Kafka topics per aggregate type, which is the reference implementation many teams start from rather than hand-rolling their own polling relay.
- Why can't you just retry the broker publish in a try/catch after the database commit and call it good enough?A retry loop still has a gap: if the process crashes entirely (not just a transient broker error) after committing the DB write but before the retry succeeds, the event is lost with no record it was ever supposed to be sent. The outbox pattern removes this gap by making the event's existence itself durable and transactional, independent of whether the process is alive to retry it.
- How does the inbox pattern relate to the outbox pattern?The outbox guarantees an event was reliably sent at least once; the inbox pattern is the consumer-side counterpart that guarantees processing that event exactly once in effect, by recording processed event IDs in an 'inbox' table within the same local transaction as the consumer's own business-data update, so a redelivered duplicate is detected and skipped.
- What happens to event ordering if you use a polling relay with multiple relay instances running for scale?Multiple concurrent pollers can race and publish rows out of order or double-publish the same row unless you add locking (e.g., SELECT ... FOR UPDATE SKIP LOCKED) or partition the outbox by aggregate ID so only one poller owns a given aggregate's rows at a time; naive multi-instance polling without this breaks per-aggregate ordering guarantees.
It's like mailing a signed contract: instead of signing the contract and separately, unreliably, hoping the mail carrier remembers to pick it up, you seal both the contract and the stamped envelope into the same filing action — the envelope only exists because the contract was filed, so a mail carrier (the relay) can come by later and deliver it, and if they lose it, you still have the file to know it needs redelivering.
saying these in an interview costs you the question
- Proposes publishing to the broker and writing to the DB as two calls wrapped only in a try/catch, without recognizing the crash-between-them gap
- Doesn't know the outbox row must be written in the SAME transaction as the business data
- Assumes the outbox pattern gives exactly-once delivery
- Forgets consumers still need idempotency even with an outbox in place
- Confuses the outbox pattern with a plain message queue