skip to content

In a service that persists to a relational database and also calls external systems (a payment provider, a message broker, a search index), how do you decide where transaction boundaries begin and end, and how do you keep the database and those systems consistent without stretching a transaction across them?

level: principalimportance: should knowfreq 30%

answer

  1. Boundary = the invariant, not the request
  2. External calls never inside; they can't roll back
  3. Outbox row in the same transaction, dispatch after commit
  4. At-least-once ⇒ idempotency key from the business operation
  5. Avoid 2PC; reconcile asynchronously and monitor outbox age

basics

~20 s

Put the boundary around the smallest unit that covers a database invariant, and keep external calls out of it. Record intent inside the transaction (an outbox row) and dispatch after commit; make external calls idempotent so retries are safe. Reconcile asynchronously rather than pretending remote systems can join the transaction.

solid answer

~60 s

Two rules do most of the work. **The boundary equals the invariant.** Whatever set of writes must be all-or-nothing shares one transaction; anything else does not belong inside. That places the boundary at the use case, not at the repository call and not at the whole HTTP request. **Nothing that can hang belongs inside it.** A payment call, a broker publish, an index update, or a wait on a user all stretch lock hold time and connection occupancy to the latency of a system you do not control, and none of them roll back when the transaction does. So: gather and validate inputs before opening the transaction; inside it write the state change plus an **outbox row** describing the intended external effect; commit; then a dispatcher reads the outbox and performs the call, retrying until acknowledged. That gives at-least-once delivery, so the receiver must be idempotent — a stable key derived from the business operation, not a per-attempt id. Where an effect must precede the write (an authorization hold), do it first and record the reference in the transaction, with a reconciliation job for orphans.

code

sql · 8 lines
sql
BEGIN;
  INSERT INTO orders (id, customer_id, total, status)
  VALUES (?, ?, ?, 'CONFIRMED');

  INSERT INTO outbox (id, aggregate_id, type, payload, created_at)
  VALUES (?, ?, 'OrderConfirmed', ?, now());
COMMIT;
-- dispatcher publishes outbox rows after commit and marks them sent

go deeper

for a junior

Say the transaction should wrap only the database writes that must happen together, and that calls to other systems belong outside it.

for a middle

Explain lock hold time and connection occupancy, that external effects cannot roll back, and the write-intent-then-dispatch-after-commit sequence.

for a senior

Cover the outbox mechanics, at-least-once delivery and idempotency keys, ordering of effects that must precede the write, and reconciliation for crash windows.

for a principal

Frame the availability-versus-consistency trade explicitly, reject two-phase commit with reasons, define the consistency window and its monitoring, and set the codebase-wide convention for boundary placement.

## The two questions a boundary answers A transaction boundary is a claim about atomicity ("these writes happen together or not at all") and a claim about resources ("these locks and this connection are held for this long"). Most boundary mistakes come from optimizing one and forgetting the other. **Where it begins:** at the first write that participates in an invariant. **Where it ends:** at the last one. Everything else — validation, computation, fetching reference data, serializing a response — should sit outside, before or after. That naturally makes the unit of work the *use case* ("place order", "cancel subscription") rather than the repository call (too narrow, breaks the invariant) or the whole request handler (too wide, includes work with no business need to be atomic). ## Why external calls must be outside Putting a remote call inside the boundary fails on three counts. 1. **Resource cost.** Lock hold time and connection occupancy become the remote system's p99 latency. A provider having a bad minute turns into pool exhaustion and a service-wide outage, because every in-flight transaction is holding a connection while waiting on the network. 2. **No shared atomicity.** A rollback cannot un-send a payment or un-publish a message. The transaction gives you nothing over the remote system, so keeping it open buys no consistency — only exposure. 3. **Ambiguous failure.** If the call times out you do not know whether it happened, and now you must decide the transaction's fate with incomplete information while holding locks. Distributed two-phase commit is the textbook alternative and is worth naming — and rejecting for most systems: it requires participant support, adds a coordinator whose failure blocks participants holding locks, and its availability characteristics are worse than the asynchronous alternative. ## The pattern: commit locally, dispatch afterwards The durable arrangement is the **transactional outbox**. Inside the one transaction that changes your state, also insert a row describing the effect to be performed — the message to publish, the index update, the call to make — with a stable business key. Commit once. A separate dispatcher polls or tails that table and performs the effect, marking the row done when acknowledged, retrying with backoff otherwise. What this buys: the state change and the *intent* to do the external work are atomic, because they are one local transaction. What it costs: the effect happens shortly *after* the commit, and may happen more than once if the dispatcher crashes between performing it and marking the row. That is at-least-once semantics, and it puts a hard requirement on the receiver — idempotency keyed by something derived from the business operation (order id, event id), never a value generated per attempt. Most real integrations already offer this, which is why the pattern is common. The mirror-image failure to avoid is publishing *before* commit, or in an after-write hook that runs while the transaction may still roll back: consumers then act on state that never became durable, which is far harder to detect than a duplicate. ## When the external effect must come first Some effects cannot be deferred — you will not create an order until the card authorization succeeds. Then the order is: perform the external call outside any transaction, then open a short transaction that records the result and its reference. The hazard is the crash between the two, leaving a hold with no local record. Handle it with a pre-recorded intent row (written and committed before the call), a reconciliation job that queries the provider for references you never linked, and a stable idempotency key so a retry of the same operation cannot double-charge. ## Scope discipline inside the boundary Two further rules keep boundaries honest. - **All writes go through the same connection.** Anything that fetches its own connection or spawns an asynchronous task inside the boundary escapes it; those writes commit independently and survive a rollback. Background work is started *after* commit, not inside. - **Order lock acquisition consistently** across paths that touch the same rows, so concurrent use cases do not deadlock — and note that if a transaction must be retried after a deadlock or serialization failure, the retry replays everything inside the boundary. That is another reason external effects must not be in there. ## What to say about the tradeoff The honest summary is that you are trading strong cross-system consistency for availability and bounded resource usage, and paying for it with eventual consistency plus idempotency and reconciliation. Say what the window looks like (typically sub-second, bounded by dispatcher lag), say how you detect when it is exceeded (outbox age, undispatched count), and say what reconciles it when something is lost. A principal-level answer names the failure modes and the monitoring, not just the pattern.

  • Why not use two-phase commit across the database and the message broker?
    It is technically possible where all participants support it, but the coordinator becomes a critical component whose failure leaves participants in-doubt while holding locks, so availability is worse than the components individually. It also requires distributed-transaction support from every participant, which most HTTP APIs and many brokers do not offer, and it adds latency to every write. The outbox plus idempotent consumers gives adequate guarantees with far better failure behaviour, at the price of eventual rather than immediate consistency.
  • How do you keep the outbox from becoming a bottleneck or an unbounded table?
    Monitor two signals: the age of the oldest undispatched row and the undispatched count — both indicate dispatcher lag or a failing downstream. Dispatch in batches with claim-and-skip semantics so multiple dispatchers can run in parallel, index the table on the pending predicate, and delete or archive completed rows aggressively so the working set stays small. If a downstream is persistently failing, move repeatedly-failing rows to a dead-letter state so one bad message cannot stall the rest.

Posting a letter: you write it and drop it in your own outbox atomically with the rest of your work, then the courier delivers it. You do not hold the whole office frozen while waiting for the recipient to sign.

saying these in an interview costs you the question

  • Calling a payment provider or publishing to a broker inside an open transaction
  • Believing a rollback can undo an external side effect already performed
  • Publishing events before commit, so consumers observe state that never became durable
  • Proposing distributed two-phase commit as the default without acknowledging its availability cost
  • Starting asynchronous work inside the transaction, where it may read state not yet committed
  • Assuming the outbox needs no idempotency because delivery 'usually works'

context