skip to content

questions

6

A service needs to save an order to its database and publish an OrderPlaced event to a message broker. Why is calling db.save(order) followed by broker.publish(event) as two separate steps risky, and what could go wrong?

level: juniorimportance: must knowfreq 75%

answer

  1. two systems, one transaction boundary
  2. crash-between-calls
  3. DB commits, publish fails (or vice versa)
  4. no XA across DB+broker in practice
  5. silent failure, no error anywhere

basics

~20 s

If the app writes to the database and then sends a message as two separate steps, one can succeed while the other fails - the app crashing between them means the database and other services disagree about what happened.

solid answer

~30 s

This is the 'dual write' problem: two independent systems (database, message broker) are updated in two separate network calls with no shared transaction. If the DB commit succeeds but the publish fails (crash, network blip, broker down), the order exists but no downstream service ever hears about it. If you publish first and the DB write then fails, you've told the world about something that never actually happened. Retrying blindly risks duplicate publishes. There's no atomic 'commit both or neither' without a special mechanism like the transactional outbox.

go deeper

for a junior

Should recognize the problem exists and describe the crash-in-the-middle scenario in plain terms; doesn't need to name the outbox pattern yet.

for a middle

Should name the term 'dual write', explain why 2PC isn't practical here, and gesture toward 'write the intent in the same DB transaction' as the fix direction.

for a senior

Should be able to sketch the outbox-table solution unprompted and explain why it converts a two-system problem into a single-system transaction plus an at-least-once relay.

for a principal

Should discuss this as an instance of a general distributed-systems problem (can't atomically commit across two independent resources without a shared coordinator) and identify when dual writes are actually tolerable (idempotent, reconcilable, low-stakes events).

## What the dual write actually is The dual-write problem arises whenever a single logical operation requires committing changes to two independent systems that share no transaction coordinator - most commonly an **OLTP database** and a **message broker**. Concretely: a service handling 'place order' needs to - (1) persist the order row so the business fact is durable, and - (2) tell the rest of the system about it by publishing an event such as `OrderPlaced` to Kafka, RabbitMQ, or SNS. These are two separate network calls to two separate pieces of infrastructure, each with its own failure modes, its own latency profile, and nothing built in that makes them succeed or fail together. ## Why it matters Why does this matter? Because a huge number of real systems are built around the assumption that 'if it's in the database, other services eventually find out.' - A **shipping service** waits for `OrderPlaced` to schedule a warehouse pick. - A **notifications service** waits for it to email the customer. - An **analytics pipeline** waits for it to update a dashboard. If the event silently never arrives, none of that happens, and there is often no error anywhere - the order just sits there, correctly saved, invisibly orphaned from the rest of the system. This is worse than a loud failure because nothing pages anyone; it surfaces days later as 'why didn't this customer get their confirmation email,' and by then the causal thread is cold. ## Neither ordering is safe The mechanics of the failure are simple to state: | Ordering | What the failure looks like | |---|---| | **save-then-publish** | means a crash, deploy restart, OOM kill, or network partition between the two calls leaves the DB committed and the event never sent | | **publish-then-save** | inverts the risk: the broker successfully has the message, but the subsequent DB write then fails (constraint violation, connection pool exhaustion, deadlock), and now consumers process an event describing an order that doesn't exist | Neither ordering is safe; you've just chosen which side of the story goes missing. ## The naive fixes, and why they collapse Naive 'fixes' don't hold up either. - **Wrapping both calls in a try/catch** and retrying the publish on failure doesn't help if the process itself dies before the retry logic ever runs - the retry loop lives in the same doomed process as everything else. - **Blindly retrying without deduplication** trades missing events for duplicate ones, pushing the problem downstream onto every consumer instead of solving it. ## Why the textbook fix is not used The theoretically 'correct' fix - a distributed transaction (two-phase commit, XA) spanning the database and the broker - is rarely used in practice. Most modern brokers (Kafka, SQS, SNS, most managed RabbitMQ setups) don't support XA participation at all, and even where **2PC** is technically available, it - introduces a transaction coordinator as a new single point of failure, - holds locks/resources open across both systems for the duration of the coordination protocol, and - interacts badly with autoscaled, ephemeral service instances that come and go. The performance and operational cost of 2PC across a database and a broker is high enough that essentially no mainstream microservices architecture uses it for this purpose. ## The practical resolution The practical resolution, which this leaf's sibling questions build on, is to convert the two-system problem into a one-system problem: 1. Write the event's intent to publish into the same database, in the same local **ACID** transaction as the business change (a dedicated **outbox table**). 2. Let a separate, decoupled process (a polling relay or a change-data-capture connector like `Debezium`) pick up that durably-recorded intent and deliver it to the broker asynchronously, with retries, independent of the original request's lifecycle. Because the outbox row and the order row commit or roll back together as one atomic unit, there is no window in which the database has the order but no record exists anywhere that an event needs to go out. The remaining relay step is decoupled and can safely retry indefinitely without risking the original business transaction, because it operates entirely after that transaction is already durably committed. ## Where it shows up A concrete real-world instance of this: Chris Richardson's microservices.io catalog documents the transactional outbox pattern explicitly as the standard answer to 'how to reliably send events when using a database and a message broker' in a microservices architecture, precisely because ad hoc dual writes were a recurring, hard-to-debug source of silent data loss across early microservice migrations.

  • Why not just use a distributed transaction (two-phase commit) across the database and the message broker?
    Most brokers (Kafka, RabbitMQ, SQS/SNS) don't support XA/2PC, and even where a broker technically does, 2PC is slow, adds a transaction coordinator as a single point of failure, and doesn't play well with autoscaled, ephemeral services. It also holds resources open across both systems for the duration of the coordination protocol, hurting throughput. In practice almost nobody wires 2PC between an OLTP database and a message broker in production.
  • What if the service just retries the publish until it succeeds?
    Retrying on its own doesn't fix atomicity - the process could crash after the DB commit but before any retry logic even runs, silently losing the event forever since the retry state lived only in memory. You need the intent to publish captured durably in the same transaction as the state change, which is exactly what the outbox table provides.

Like mailing a wedding invitation and updating your guest list in two separate trips to two separate buildings - if you get hit by a bus after mailing but before updating the list (or vice versa), your records and the world's knowledge of what happened disagree, and nothing tells you that happened.

saying these in an interview costs you the question

  • Says 2PC/XA transactions are the standard production fix
  • Doesn't recognize that publish-then-save has the same problem as save-then-publish, just inverted
  • Thinks retrying the publish call alone solves atomicity
  • Can't state which order (DB first vs broker first) is 'safer' and why neither is safe alone

context

open as a page

Walk through, step by step, how the transactional outbox pattern uses an outbox table plus a relay/poller process to publish an OrderPlaced event reliably after an order is saved.

level: middleimportance: must knowfreq 80%

basics

~20 s

The service writes both the order row and a row describing the event into the same database transaction. A separate background process then reads the new outbox rows, sends them to the message broker, and marks them as sent - so the event is only ever 'in flight' if the original database write actually succeeded.

open as a page

A team's outbox relay publishes events out of order across different aggregates, and occasionally a consumer receives the same OrderPlaced event twice. What guarantees does the transactional outbox pattern actually provide here, and what must consumers/producers do to handle it correctly?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The outbox pattern guarantees an event eventually gets published if the database write succeeded, but not that it's published exactly once or perfectly ordered across everything - so consumers must be built to safely handle duplicate or occasionally reordered messages.

open as a page

A team implements the outbox pattern using a database polling relay that queries SELECT * FROM outbox WHERE published = false every second. A colleague suggests replacing it with Debezium reading the database's write-ahead log via change data capture instead. What changes, and what are the trade-offs?

level: middleimportance: should knowfreq 60%

basics

~20 s

Polling repeatedly asks the database 'anything new?' which adds load and delay. CDC tools like Debezium instead tap directly into the database's internal change log and stream new rows out in near real time, without hammering the table with queries - but it needs more setup and access to that low-level log.

open as a page

Instead of writing to an outbox table, a service publishes an OrderPlaced event to Kafka first, and a consumer inside that same service subscribes to its own topic to then update its local 'orders' read table. What is this approach called, and how does it avoid the dual-write problem?

level: seniorimportance: should knowfreq 30%

basics

~20 s

This is called 'listen to yourself' - the service treats publishing the event as the single source of truth, and only updates its own database in reaction to consuming that event back, so there's just one write path instead of two independent ones.

open as a page

Under what circumstances would a senior engineer argue AGAINST introducing the transactional outbox pattern for a service that currently does a direct, unguarded dual write (save to DB, then publish), and what would they propose instead?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

If the event being published isn't critical - losing or duplicating it occasionally causes no real harm, or the same information can be recovered another way - then adding an outbox table and a relay process might be more complexity than the problem is worth.

open as a page