When a service updates its own database and then needs to publish an event about that change so other services can update their read models, what can go wrong if it does these as two separate steps, and what's a common pattern to fix it?
answer
- dual-write = two systems, no shared transaction
- outbox table in same local transaction
- relay: poll or CDC forwards it
- at-least-once delivery -> idempotent consumers
- silent drift if event never published
basics
~20 sIf a service saves to its database and then separately sends a message, one of those two steps can fail after the other succeeds, for example the app crashes right between them, leaving the database and the message out of sync. A common fix is to write the event into the same database transaction as the data change, then have a separate process reliably forward it.
solid answer
~50 sThis is the 'dual write' problem: writing to the database and publishing to a message broker are two independent operations that can't be wrapped in one atomic transaction across two different systems, so a crash between them leaves either an event that was never sent, data changed but no one told, or an event sent for a write that then rolled back, an announcement of something that didn't happen. The standard fix is the transactional outbox pattern: write the event as a row in an outbox table inside the same local database transaction as the business data change, so both succeed or fail together atomically, then a separate relay process, polling the outbox or CDC tooling reading the database's transaction log, reads new outbox rows and publishes them to the broker, retrying until acknowledged, then marking them sent.
go deeper
Can recognize in plain terms that 'save to DB' and 'send a message' are two separate steps that can fail independently, causing things to get out of sync.
Can explain the outbox table concept: write the event with the data change in one local transaction, then forward it separately.
Compares polling vs CDC-based relays, explains why delivery is at-least-once, and designs idempotent consumers to match.
Weighs outbox and CDC infrastructure cost across the whole service fleet, sets standard event/outbox schema and tooling org-wide, and decides when the added complexity is or isn't justified for a given service.
## Two writes that cannot be made one Consider a service handling a request that both changes its own data and needs to tell the rest of the system about that change: it opens a local database transaction, updates its own tables, and commits — then, as a separate step, calls a message broker to publish an event describing what happened. The problem is that these are two independent operations against two different systems that don't share a transaction coordinator, so there's no way to guarantee both happen or neither does. - **Commit, then publish.** If the process crashes, the network hiccups, or the broker is briefly unreachable in the gap between the commit and the publish call, you get a database that changed with no event ever sent — other services relying on that event to update their own read models never find out, and their copies silently fall behind with no error raised anywhere. - **Publish before commit.** The reverse ordering is just as broken: if the publish succeeds but the subsequent database write then fails or rolls back due to a constraint violation, you've announced something to the rest of the system that never actually happened. ## Why it happens This is called the **dual-write problem**, and it exists because a relational database transaction and a message broker publish are fundamentally different kinds of operations with no shared all-or-nothing guarantee between them. In principle you could use a distributed transaction protocol like two-phase commit across both systems, but in practice this is rarely used: - most modern message brokers don't support distributed-transaction protocols well or at all; - even where it's available it adds real latency and couples the availability of the write path to the availability of the broker, which defeats much of the point of decoupling services with asynchronous messaging in the first place. ## The standard fix The standard fix is the **transactional outbox pattern**. Instead of publishing to the broker directly, the service writes the event as a row into an 'outbox' table, inside the exact same local database transaction as the business data change. Because both writes go to the same database in the same transaction, they are atomic by construction: either both the business change and the outbox row commit together, or neither does, using nothing more exotic than the local ACID guarantee every relational database already provides. Delivery to the broker is then handled by a separate process: 1. either a lightweight **relay** that polls the outbox table for unpublished rows and forwards them, 2. or, increasingly common, a **change-data-capture** tool that tails the database's transaction log for new outbox rows and streams them to the broker without the application needing to poll anything itself. Once a row is confirmed delivered, the relay marks it published, or deletes and archives it. ## What the fix still costs This fix has its own costs and its own remaining failure mode to design around. It's inherently at-least-once delivery, not exactly-once: if the relay crashes after successfully publishing but before marking the row as sent, it will republish that same event on restart. Consumers of these events therefore have to be idempotent, typically by tracking the highest event ID or version they've already applied per source and ignoring anything not newer, because delivered exactly once isn't a guarantee this pattern, or most real message brokers, actually provides. There's also new operational surface: - the outbox table needs monitoring for a growing backlog; - a stuck relay is now a silent failure mode of its own; - it needs periodic cleanup or archival so it doesn't grow unbounded, since every published business event otherwise sits there forever. ## Putting it together A concrete walk-through: in an order-placement flow, the Order service inserts a new order row and an `OrderPlaced` row into its outbox table in one local transaction, so both succeed or fail together. A change-data-capture tool reads the database's transaction log, notices the new outbox row, and forwards it to the message broker. The Inventory service and the Shipping service each independently consume that topic to update their own local read models, reserving stock and scheduling a shipment respectively, and each one is written to be idempotent, keyed on the event's unique ID, so that if the relay ever redelivers the same `OrderPlaced` event after a crash, Inventory doesn't double-reserve the stock and Shipping doesn't schedule two shipments for one order.
- Why not just use a distributed transaction, two-phase commit, across the database and the message broker instead of an outbox table?Two-phase commit requires both systems to support the protocol, most modern message brokers don't support it well or at all, and even when available it adds significant latency and availability coupling — if the coordinator or either participant is slow, the whole transaction blocks. The outbox pattern gets the same atomicity guarantee using only the database's native local transaction, which every service already has, at much lower operational cost.
- The outbox relay delivers a given event twice because it crashed after publishing but before marking the row as sent. Is this a problem?Only if the downstream consumer isn't idempotent. Because outbox-based delivery is inherently at-least-once, consumers need to detect and ignore an event they've already applied, typically by tracking the last-processed event ID or version per source and skipping anything not newer. This is a standard, expected part of the pattern, not a bug in the outbox itself.
- How do you keep the outbox table from growing forever?A cleanup job periodically deletes or archives rows once they're confirmed published and old enough that redelivery is no longer a concern, often after a retention window that also supports debugging or replay. Some teams move published rows to a cold archive table or object storage instead of hard-deleting them, in case they're needed for auditing.
Like mailing a signed contract and calling to confirm receipt as two separate actions — if you get interrupted right after signing but before the call, the other party never finds out the deal happened, even though your copy is legally binding. The outbox pattern is like sealing a copy of the notification in the same envelope as the contract itself, so they're guaranteed to travel together.
saying these in an interview costs you the question
- Thinks writing to the database and publishing to the broker can be made atomic just by calling publish() right after commit() with no other mechanism
- Doesn't recognize that outbox-based delivery is at-least-once and consumers must be idempotent
- Proposes publishing the event before the database commit as a fix
- Has never heard of CDC or log-tailing as an alternative to polling the outbox table
- Assumes this is a rare edge case rather than something that happens routinely under crashes, restarts, and deploys