Walk through the read-process-write ordering that makes a dedup store actually safe under consumer crashes. Where exactly do you record the key, side effect, and offset?
answer
- 3 actions: key write, side effect, offset commit
- key write + side effect = ONE transaction
- commit offset AFTER the DB commit
- enable.auto.commit=false, manual ack
- crash after DB commit/before offset commit => dedup absorbs redelivery
basics
~20 sMake the side effect and the dedup-key write part of the same atomic transaction, so either both happen or neither does. Commit the Kafka offset only after that transaction succeeds. On redelivery, the existing key tells you to skip.
solid answer
~50 sThe trap is ordering. You have three actions: (1) record the dedup key, (2) perform the side effect, (3) commit the Kafka offset. The danger windows: if you commit the offset before the side effect, a crash loses the work (at-most-once); if the side effect and the key-write aren't atomic, a crash between them leaves the side effect done but the key missing, so redelivery re-does it. The safe pattern when your side effect is a DB write: put the side-effect row AND the dedup-key insert in ONE database transaction (the UNIQUE constraint on the key is what enforces dedup). Commit that transaction, THEN commit the Kafka offset (typically with enable.auto.commit=false and a manual commit after processing). On redelivery, re-inserting the key violates the UNIQUE constraint, you catch it, skip the side effect, and just commit the offset. The offset commit is allowed to be non-atomic with the DB because the dedup key makes reprocessing idempotent.
code
kotlin · 18 lines@KafkaListener(topics = ["orders"])
fun onMessage(record: ConsumerRecord<String, OrderEvent>, ack: Acknowledgment) {
val key = record.value().orderId // business idempotency key
try {
// single DB transaction: dedup key + side effect together
processInTransaction(key, record.value())
} catch (e: DuplicateKeyException) {
// UNIQUE violation => already processed, safe to skip
}
// commit Kafka offset only after the DB transaction succeeded
ack.acknowledge()
}
@Transactional
fun processInTransaction(key: String, event: OrderEvent) {
processedRepo.insert(key) // UNIQUE(key) enforces dedup
orderRepo.insert(event.toOrder()) // the actual side effect
}go deeper
Know the rule of thumb: do the work, then tell Kafka you're done — never the reverse.
Order the three steps correctly and know to disable auto-commit and commit offsets after processing.
Co-commit dedup key + side effect in one DB transaction, then commit the offset; analyze each crash window.
Decide when local dedup suffices vs pushing idempotency downstream or using transactional outbox/EOS, and codify the ordering as a team pattern.
## The three actions and why order matters Processing one record involves: 1. **Dedup-key write** — record that this message was handled. 2. **Side effect** — the actual business write (insert an order row, decrement inventory). 3. **Offset commit** — tell Kafka you're done, so you won't re-read this record. Kafka offset commits and your external store (DB/Redis) are **two different systems**; you cannot make a single atomic commit across both without exactly-once machinery. So the strategy is: make reprocessing **safe** (idempotent) via the dedup key, and choose an order where the worst-case crash is recoverable. ## The failure windows - **Commit offset first, then side effect**: crash after offset commit but before side effect → record never reprocessed → **lost work** (effectively at-most-once). Never do this for non-idempotent work. - **Side effect, then dedup-key write (separate transactions)**: crash between them → side effect done, key absent → redelivery re-does the side effect → **duplicate**. The dedup store didn't help. - **Dedup-key write, then side effect (separate transactions)**: crash between them → key present, side effect missing → redelivery sees the key, skips → **lost work**. The lesson: the **dedup-key write and the side effect must be atomic** with respect to each other. ## The safe pattern (DB side effect) When the side effect is itself a write to the same relational database: 1. Begin a DB transaction. 2. `INSERT` the dedup key (UNIQUE constraint enforces it). If it throws a unique-violation → this message was already processed → roll back / no-op, skip to offset commit. 3. Perform the business write(s). 4. **Commit the DB transaction** (key + side effect land together, atomically). 5. **Then commit the Kafka offset** (manual commit; `enable.auto.commit=false`). Now analyze crashes: - Crash before step 4 commits → nothing persisted → redelivery redoes cleanly. - Crash after step 4 but before step 5 → key + side effect persisted, offset NOT committed → redelivery re-reads the record → the dedup-key INSERT fails (already present) → side effect skipped → offset committed. **Idempotent.** This is the **"process-then-commit-offset"** discipline combined with **dedup-key-and-side-effect-in-one-transaction**. The Kafka offset is allowed to lag the DB because the dedup key absorbs the redelivery. ## When the side effect is NOT in the same DB (e.g. Redis SETNX, or a remote call) You lose the single-transaction atomicity. Mitigations: - **Dedup key in the same store as the effect when possible** (e.g. effect is also in Redis). - For a remote, non-transactional side effect (HTTP call), the cleanest fix is to make the downstream itself idempotent (pass the idempotency key to it) rather than rely solely on a local dedup store. - A **claim-check** variant: SETNX the key first with a short 'in-progress' marker; if the process crashes mid-effect, a redelivery sees the marker, and you need a reconciliation/timeout policy — more complex and still not perfectly atomic. ## Offset-commit details - Set `enable.auto.commit=false` so offsets don't advance on a timer independent of your processing. - Commit after the transaction, per-batch or per-record. Per-record commits are simpler to reason about but costlier; per-batch is fine because the dedup key handles partial-batch reprocessing. - Spring Kafka users: `AckMode.MANUAL`/`MANUAL_IMMEDIATE` and `ack.acknowledge()` after the DB commit express exactly this ordering. ## Why this isn't full exactly-once It's at-least-once delivery + idempotent processing = **effectively-once side effects**, scoped to what the dedup store covers. (The exactly-once *concept* itself is a sibling topic.)
- Why is committing the Kafka offset before the side effect dangerous?A crash after the offset commit but before the side effect means the record is never re-read, so the work is permanently lost — that's at-most-once, unacceptable for non-idempotent side effects.
- If the dedup key and the side effect are in separate transactions, what goes wrong?A crash between them breaks atomicity: either the key exists without the effect (work skipped on retry = lost) or the effect exists without the key (re-done on retry = duplicate). Only co-committing them is safe.
- How do you express this ordering in Spring Kafka?Set enable.auto.commit=false, use a manual ack mode (MANUAL/MANUAL_IMMEDIATE), wrap the dedup-key insert plus side effect in a @Transactional method, and call ack.acknowledge() only after that method returns successfully.
saying these in an interview costs you the question
- Committing the Kafka offset before performing the side effect.
- Writing the dedup key and the side effect in two independent transactions.
- Leaving enable.auto.commit=true so offsets advance on a timer regardless of processing.
- Claiming this gives true cross-system exactly-once (it gives effectively-once via idempotency).