skip to content

You moved message publishing to AFTER_COMMIT so you never publish on rollback. A colleague says the system is now fully consistent. What failure can still occur, and what is this problem called?

level: seniorimportance: must knowfreq 55%

answer

  1. two directions: false events vs lost events
  2. commit then crash before publish = lost
  3. dual-write = two systems, no shared tx
  4. AFTER_COMMIT ≈ at-most-once
  5. outbox → at-least-once + idempotent consumers

basics

~20 s

The DB can commit and then the app crash (or the broker be down) before the message is published, so the event is lost forever. AFTER_COMMIT prevents false events but not lost ones. This is the dual-write problem: two systems updated without a shared transaction.

solid answer

~50 s

AFTER_COMMIT closes only one direction of inconsistency: it stops you publishing events for data that later rolls back. It does not guarantee the message is ever sent. Between the successful commit and the publish call, the JVM can crash, the broker can be unavailable, or the send can throw — and because the transaction is already committed, nothing retries it. Now the database says the order exists but no OrderPlaced event was ever emitted: a permanently lost message. This is the dual-write problem — writing to two independent systems (the DB and the message broker) that share no atomic transaction, so any interleaving of failures can leave them inconsistent. AFTER_COMMIT gives you at-most-once-ish semantics, not at-least-once. The standard fix is the transactional outbox: persist the message into an outbox table within the same DB transaction, then a relay reads that table and publishes, turning the dual write into a single atomic write plus reliable delivery.

code

java · 14 lines
java
@Transactional
public void place(Order o) {
    orderRepository.save(o);
    publisher.publishEvent(new OrderPlaced(o.getId()));
}

@TransactionalEventListener // AFTER_COMMIT
public void publish(OrderPlaced e) {
    // COMMIT already happened. If the JVM dies or Kafka is down
    // right here, this send never completes and the event is
    // lost forever -- the DUAL-WRITE problem. AFTER_COMMIT does
    // not make this atomic with the DB write.
    kafkaTemplate.send("orders", e.orderId().toString());
}

go deeper

for a junior

May only know the rollback direction; likely misses the lost-event direction.

for a middle

Recognizes the crash window exists but may not name it the dual-write problem.

for a senior

Names the dual-write problem, frames delivery semantics, and points to the outbox — target level.

for a principal

Discusses XA trade-offs, at-least-once + idempotency, CDC vs polling relays, and where Spring Modulith's registry sits.

## The two directions of inconsistency When a service both writes the database **and** notifies the outside world, there are two ways to become inconsistent: 1. **Notify but don't persist** — publish an event, then the transaction rolls back. Downstream acts on data that doesn't exist. *AFTER_COMMIT fixes this.* 2. **Persist but don't notify** — commit the database, then fail to publish (crash, broker down, exception). Downstream never learns about data that does exist. *AFTER_COMMIT does NOT fix this.* Moving side effects to AFTER_COMMIT only addresses direction #1. Your colleague is wrong: direction #2 is still wide open. ## Why direction #2 persists with AFTER_COMMIT The AFTER_COMMIT callback runs **after** the commit returns. Consider the timeline: ``` T1 BEGIN T2 INSERT order T3 COMMIT <-- durable here <== JVM crash / broker outage / send() throws T4 publisher.publish(OrderPlaced) // never runs, or fails ``` At T3 the order is durably in the DB. If anything goes wrong before T4 completes, the message is **lost with no compensation** — Spring does not persist or retry the callback. You now have an order with no corresponding event, forever. ## The name: the dual-write problem This is the **dual-write problem** (a.k.a. write-to-two-systems / the 'no distributed transaction' problem). You are performing **two writes** — one to the database, one to the message broker — that are **not covered by a single atomic transaction**. Any failure between them, in either order, breaks consistency. XA/two-phase commit could in theory span both, but it is heavyweight, brittle, poorly supported by modern brokers (Kafka), and generally avoided. ## Delivery-semantics framing - Publish **before** commit → risk of **phantom** events (over-delivery relative to committed data). - Publish **after** commit (this pattern) → risk of **lost** events (under-delivery). Roughly at-most-once. - **Transactional outbox** → **at-least-once**: the event intent is committed atomically with the data, and a relay retries delivery until it succeeds (consumers must be idempotent to absorb duplicates). ## The fix in one line Turn two non-atomic writes into **one atomic write + asynchronous reliable delivery**: write the message to an **outbox** table in the same transaction as the business data, then a separate poller or CDC (change-data-capture) relay reads the outbox and publishes. ## Where Spring fits Spring Modulith's **event publication registry** implements exactly this: an `@ApplicationModuleListener` event is stored in an `event_publication` table inside the transaction; incomplete publications are republished on restart — an outbox for in-process module events. For cross-service brokers, teams use Debezium/CDC or a polling relay over an outbox table. ## Interview takeaway AFTER_COMMIT is necessary but not sufficient. It eliminates false positives (events for rolled-back data); it does not eliminate false negatives (missing events for committed data). Naming the dual-write problem and the outbox remedy is the senior-level answer.

  • Why not just use an XA / two-phase-commit transaction spanning the DB and the broker to make both atomic?
    XA is heavyweight, hurts throughput, has fragile recovery, and many modern brokers (notably Kafka) don't participate well. The industry standard is to avoid distributed transactions and use the outbox pattern with idempotent consumers instead.
  • Does publishing BEFORE commit instead solve it?
    No, it just flips the failure mode: you risk phantom events for data that then rolls back. Neither pure before nor pure after commit is reliable; you need the outbox to get atomicity plus at-least-once delivery.

saying these in an interview costs you the question

  • Claiming AFTER_COMMIT makes the system fully consistent / guarantees delivery.
  • Not recognizing the crash-after-commit window.
  • Proposing XA/2PC as the obvious correct fix without noting its practical problems.
  • Confusing the two failure directions (phantom vs lost events).

context