skip to content

Outbox delivery is at-least-once, so consumers may see duplicate Kafka messages. How do you make a consumer idempotent?

level: middleimportance: must knowfreq 68%

answer

  1. At-least-once -> duplicates inevitable
  2. Unique event id (outbox row id / UUID)
  3. processed_messages PK in SAME txn as effect = inbox
  4. Or natural idempotency: upsert, set-to-state, version check
  5. Offset commit alone = another dual write

basics

~20 s

Give each event a unique id. The consumer records processed ids (e.g. in a 'processed_messages' table) inside the same transaction that applies the effect, and skips any id it has already seen. That way reprocessing a duplicate does nothing.

solid answer

~50 s

Because the relay publishes at-least-once, a consumer can receive the same event more than once (relay retry, consumer redelivery after a crash before commit). Idempotency means processing a duplicate has no extra effect. The standard technique: every outbox event carries a stable unique id (the outbox row id / event id). The consumer applies the business effect and inserts that id into a dedup store (e.g. a `processed_messages` table with the id as primary key) in the **same local DB transaction**; a duplicate hits a unique-key violation and is skipped. Alternatives: make the operation **naturally idempotent** (upserts, set-to-state rather than increment, conditional updates with version/optimistic locking), or dedup by a deterministic key. You generally cannot rely on Kafka offset commits alone for exactly-once, because the effect (DB write) and the offset commit are themselves a dual write — hence the inbox/processed-id table that ties effect + dedup into one transaction.

go deeper

for a junior

Know that duplicates can arrive and the fix is a unique event id the consumer checks before acting.

for a middle

Implement the processed-id/inbox table inside the same transaction as the effect; cite naturally idempotent operations.

for a senior

Explain why offset commits alone fail, manage dedup-store growth, and combine idempotency with per-key ordering.

for a principal

Standardize event-id conventions and idempotency contracts across services; weigh inbox table vs natural idempotency vs Kafka EOS per pipeline.

## Why duplicates are unavoidable The outbox + relay path is **at-least-once**: the relay may crash after sending to Kafka but before marking the row published, so it resends; and a consumer may crash after applying an effect but before committing its Kafka **offset**, so Kafka redelivers. Therefore every consumer in this pattern **must tolerate duplicates**. 'Tolerate' means **idempotent**: handling the same event N times yields the same result as handling it once. ## Technique 1 — dedup table (the 'inbox' / processed-ids) Every outbox event carries a **stable unique id** — typically the outbox row's primary key (a UUID is ideal). The consumer, in **one local transaction**, both applies the effect and records the id: ```sql BEGIN; INSERT INTO processed_messages(message_id) VALUES (:eventId); -- PK; dup -> conflict UPDATE account SET balance = balance - :amt WHERE id = :acct; -- the effect COMMIT; ``` If the event was already processed, the `INSERT` violates the primary key and the whole transaction aborts — the effect is **not** re-applied. Because effect + dedup commit together, there's no window where one happens without the other. This is sometimes called the **inbox pattern** (mirror of the outbox). ## Technique 2 — naturally idempotent operations Design the effect so repetition is harmless: - **Upsert / set-to-state**: `SET status = 'SHIPPED'` is idempotent; `status = status + 1` is not. - **Conditional / optimistic update**: `UPDATE ... WHERE version = :expected` — a stale duplicate matches nothing. - **Deterministic derived keys**: write to a row keyed by the event so re-writes overwrite identically. ## Why offsets alone aren't enough A tempting idea: just commit the Kafka offset after processing. But the effect (DB write) and the offset commit live in **two systems** — that's another dual write. If you commit the offset and then crash before the DB write (or vice versa), you either skip or reprocess. The processed-id table folds dedup **into the same DB transaction as the effect**, sidestepping that gap. (Kafka's own exactly-once/transactional producer helps for Kafka-to-Kafka pipelines, but a consumer writing to an external DB still needs application-level idempotency.) ## Edge cases - **Dedup-store growth**: prune processed ids by time window or partition; size the window to exceed maximum possible redelivery lag. - **Ordering vs idempotency**: idempotency stops double-apply; it does not reorder. Pair it with per-key ordering. - **Cross-partition rebalance**: after a consumer group rebalance, redelivery from the last committed offset is normal — the dedup table absorbs it.

  • Why isn't committing the Kafka offset after processing enough to guarantee exactly-once effects on a database?
    The DB write and the offset commit are in two different systems with no shared transaction — a dual write. A crash between them either reprocesses (offset not committed) or skips (offset committed, effect lost). A processed-id row committed with the effect closes that gap.
  • Give an example of a naturally idempotent operation versus a non-idempotent one.
    Idempotent: 'SET balance = 100' or 'UPDATE ... WHERE version = 7' or an upsert keyed by the event. Non-idempotent: 'balance = balance - 10' (increment/decrement) — applying it twice double-charges.

saying these in an interview costs you the question

  • Assuming the outbox alone gives exactly-once so consumers need no dedup.
  • Relying only on Kafka offset commits for DB-effect exactly-once (still a dual write).
  • Using a non-stable/derived id that differs across redeliveries (dedup fails).
  • Recording the processed id in a separate transaction from the effect (reintroduces a gap).
  • Confusing idempotency (no double-apply) with ordering (no reorder) — you need both.

context