A command handler for `PlaceOrder` first commits the new order state to its database, and then, as a second separate step, publishes the resulting `OrderPlaced` event to a message broker. What failure mode does this two-step design create, and what are the common ways teams close that gap?
answer
- dual-write problem
- transactional outbox pattern
- CDC / Debezium-style relay
- event sourcing = no second step
- state and event must not diverge
basics
~20 sIf the app crashes right after saving to the database but before sending the event, the save happened but nobody downstream finds out. The write and the event need to succeed or fail together, not as two separate bets.
solid answer
~50 sThis is the dual-write problem: the database commit and the broker publish are two independent operations, with no distributed transaction spanning both, so a crash or broker outage between them leaves the two out of sync - most dangerously, state committed with its event silently lost, so no consumer ever hears about an order that genuinely exists. The standard fix is the transactional outbox pattern: write the event as a row in an outbox table in the same database transaction as the state change, then a separate, dependable process, a poller or a change-data-capture connector, reads unpublished outbox rows and forwards them to the broker, retrying until acknowledged. Event sourcing sidesteps the problem differently: the events are the write, appended atomically as the only source of truth, so there's no second step to lose. Both approaches trade extra infrastructure for the guarantee that state and event never diverge.
go deeper
Not expected to know this pattern by name; can be prompted to notice that save-then-send has two steps that could fail independently.
Can articulate that a crash between the two steps loses the event, even without naming the outbox pattern.
Names the transactional outbox pattern and can sketch how the outbox table and relay process work together.
Can compare outbox-with-relay against event sourcing as two structurally different solutions to the same class of problem, and discuss the downstream idempotency burden the outbox's at-least-once relay pushes onto consumers.
## The dual-write problem The dual-write problem arises whenever a single logical operation needs to durably affect two independent systems — here, a relational database holding the order state and a message broker that will notify downstream consumers — and there is no shared transaction coordinator spanning both. Concretely, the handler first commits the order row inside its own database transaction; that commit either fully succeeds or fully rolls back, and the database's own guarantees end right there. The handler then, as a wholly separate operation, calls the broker's publish API. Between those two steps sits a window during which: - **the process can crash**, - **the network to the broker can fail**, or - **the broker itself can be unreachable**. If anything goes wrong in that window, the database commit has already happened and cannot be un-done just because the publish failed — the order genuinely exists — yet no `OrderPlaced` event was ever delivered, so every downstream consumer that should have reacted to it never will, silently and with no error surfaced anywhere obvious. ## Why it matters, and why it exists at all This failure mode matters because CQRS and event-driven systems depend on the event stream being a **faithful, complete record** of everything that happened to the write side; a single silently dropped event breaks that contract in a way that's very hard to detect after the fact, since there's no error, no exception, nothing in a log unless someone specifically instrumented the gap. The problem exists fundamentally because most relational databases and most message brokers are separate systems with separate commit protocols, and genuinely distributed two-phase-commit transactions across the two — even where technically possible — are rare in practice because of the operational complexity, latency cost, and the fact that many popular brokers, such as Kafka, don't support XA-style distributed transactions with arbitrary external resource managers in the first place. ## The transactional outbox The standard resolution, the transactional outbox pattern, works by turning the second write into a write against the same database as the first. Instead of calling the broker directly inside the handler, the handler writes a row describing the event — its type, payload, and an ordering key — into an 'outbox' table, in the exact same local database transaction that commits the order state change. Because both writes are now against the same database in the same transaction, they succeed or fail together with full ACID guarantees; there is no window where one happens and the other doesn't. A separate, independent process then takes responsibility for actually getting that row to the broker: - **a poller** that periodically queries for unsent outbox rows, or - **a change-data-capture connector such as Debezium** that tails the database's write-ahead log and streams new outbox rows to the broker as they're committed. That relay process retries until the broker acknowledges receipt, and only then marks the row as sent or deletes it. ## The trade-off The trade-off is added infrastructure and a shift in where complexity lives: - **an outbox table** to maintain and eventually prune, - **a relay process** that itself needs monitoring, and - **critically, a new obligation pushed onto every consumer of the resulting events** — because the relay guarantees only at-least-once delivery, not exactly-once, a crash between publishing to the broker and marking the outbox row as sent will cause a redelivery on restart, so every downstream consumer must now be idempotent to duplicate `OrderPlaced` events, which is the same idempotency discipline command handlers need for duplicate commands, just moved to the event-consuming side. ## Event sourcing answers it structurally Event sourcing offers a structurally different answer to the same class of problem: if the events themselves are the durable state, with no separate 'current state' row committed independently, then there is only ever one write to make durable — appending to the event store — and publishing to consumers can be done by tailing that same append-only log, so there's nothing left to diverge in the first place. That said, event sourcing brings its own substantial costs, like rebuilding read models from event streams and handling event schema evolution, so it's not a drop-in replacement for the outbox pattern in a conventional state-stored system. ## A well-known instance A well-known real-world instance of this exact pattern is Debezium's outbox event router, built specifically to consume outbox-table rows via CDC and republish them to Kafka, which several production systems adopt precisely because it decouples the reliability guarantee, at-least-once delivery from a committed database row, from the application code's own retry logic, letting the database's own durability do the heavy lifting instead of hoping an in-process publish call never fails at the wrong moment.
- Why not just publish the event first, before committing the database write?That flips which side can go missing: if the publish succeeds but the database commit then fails or rolls back, consumers react to an event describing something that never actually happened in the system of record, which is often worse, since you can't easily un-notify downstream systems. Publishing after commit at least fails toward under-notification rather than over-notification, but neither raw ordering alone solves the problem without an outbox or equivalent.
- What does the outbox relay process need to guarantee, and what happens if it crashes mid-relay?It needs at-least-once delivery of each outbox row to the broker, tracked by marking rows as sent, or deleting them, only after a broker acknowledgment. If it crashes after the broker ack but before marking the row sent, it will resend on restart, so consumers must be able to handle duplicate events, which pushes the idempotency requirement downstream to event consumers as well.
- How does event sourcing avoid this problem structurally, rather than just papering over it?In event sourcing, the events themselves are the durable state - there is no separate state row that gets committed first; appending the event to the event store is the commit. Since publishing to other consumers can be done by tailing that same append-only log, there's only ever one write to make durable, not two independent writes that can diverge.
It's like paying a contractor and mailing them a signed confirmation letter as two separate trips to the post office. If something interrupts you after paying but before mailing the letter, the payment is real but the contractor has no proof and no notification - from their side, nothing happened. The fix is to put the payment and the letter in the same envelope handed to one courier in one trip, so either both arrive or the whole trip is retried.
saying these in an interview costs you the question
- Doesn't recognize that a database commit and a broker publish are two separate points of failure
- Proposes 'just retry the publish' as a complete fix without addressing what happens if the process crashes before retrying
- Confuses the outbox pattern with simply logging events for audit purposes, missing that it's the mechanism that makes the publish reliable
- Assumes a message broker publish and a database commit can be wrapped in one ACID transaction across both systems