What are the three delivery semantics in Kafka (at-most-once, at-least-once, exactly-once), and what does each guarantee about message delivery?
answer
- 0/1, 1+, exactly 1
- loss vs duplicate vs neither
- at-least-once = default
- commit-then-process vs process-then-commit
- EOS = idempotent + transactions
basics
~20 sAt-most-once: each message is delivered zero or one time (loss possible, no duplicates). At-least-once: delivered one or more times (duplicates possible, no loss) — Kafka's default. Exactly-once: delivered once and only once (no loss, no duplicates).
solid answer
~40 sKafka defines three delivery guarantees. At-most-once means a message may be lost but never duplicated — you favor speed/no-dupes over completeness. At-least-once means a message is never lost but may be processed more than once — duplicates are possible, so consumers must be idempotent; this is Kafka's default. Exactly-once means each message affects state once and only once — no loss and no duplicates. The differences arise from the ORDERING of two actions on the consumer side: committing the offset versus processing the record. Commit-before-process gives at-most-once (a crash after commit loses the record). Process-before-commit gives at-least-once (a crash after processing but before commit reprocesses the record). Exactly-once needs extra machinery (idempotent producer + transactions) to make processing and offset commit atomic.
go deeper
Memorize the three labels and their loss/duplicate property: at-most (loss, no dupes), at-least (dupes, no loss, default), exactly (neither).
Connect each semantic to the commit/process ordering and to producer retry behavior; know to make at-least-once consumers idempotent.
Explain the exact crash windows that create loss vs duplicates, and why ordering alone can't yield exactly-once.
Frame the trade-off for a system: cost/latency of EOS vs idempotent at-least-once, and when each is the right architectural choice.
## The problem In any distributed messaging system a producer sends data, a broker stores it, and a consumer reads and processes it. Crashes, retries, and network failures can occur at any step. The **delivery semantic** describes what happens to each message when those failures happen: can it be lost, can it be seen more than once, or neither? ## Key terms - **Offset**: a monotonically increasing integer that identifies each record's position within a partition. A consumer tracks how far it has read by **committing** an offset (storing 'I have consumed up to here' in the `__consumer_offsets` topic). - **Process**: the consumer's business work for a record (write to a DB, call an API, etc.). - **Ack**: the acknowledgement a broker sends a producer confirming a write was persisted. ## The three semantics 1. **At-most-once** — every message is delivered **0 or 1 times**. Messages may be **lost** but are **never duplicated**. You accept loss in exchange for never seeing a record twice and for lower latency. 2. **At-least-once** — every message is delivered **1 or more times**. Messages are **never lost** but may be **duplicated**. This is Kafka's **default** behavior. Because dupes can occur, downstream processing should be **idempotent** (re-applying it has no extra effect). 3. **Exactly-once** — every message effectively affects state **once and only once**: **no loss, no duplicates**. The strongest and most expensive guarantee. ## Where the semantic comes from On the **consumer side** the semantic is determined by the order of two operations: - **Commit the offset, then process** → at-most-once. If the consumer crashes after committing but before finishing processing, the record is skipped on restart (it looks already-consumed) → **loss**. - **Process, then commit the offset** → at-least-once (the default). If the consumer crashes after processing but before committing, on restart it re-reads the record and processes it again → **duplicate**. On the **producer side** the semantic is determined by retries and acks: a producer that retries after a network glitch (where the broker actually persisted the first attempt but the ack was lost) creates a **duplicate** — at-least-once. Disabling retries avoids duplicates but risks **loss** — at-most-once. ## Exactly-once You cannot get exactly-once by ordering alone, because there is always a window where a crash duplicates or loses work. Kafka achieves it by combining the **idempotent producer** (dedups producer retries on the broker) with **transactions** that atomically commit both the produced output records AND the consumed input offsets. If the transaction aborts, neither is visible. That makes the read-process-write cycle atomic, eliminating both the duplicate and the loss windows. ## Practical takeaway Most teams run **at-least-once** and make consumers idempotent — it is simpler and cheaper than full exactly-once and avoids the data loss of at-most-once.
- Which of the three is Kafka's default and why?At-least-once. Auto-commit and the typical process-then-commit flow mean a crash before the offset commit causes reprocessing — duplicates, never loss.
- If you can only have at-least-once, how do you make the end result correct?Make the consumer's processing idempotent (e.g. upsert by a business key, or dedup by record key/offset) so reprocessing a duplicate has no extra effect.
saying these in an interview costs you the question
- Saying at-least-once means messages can be lost (it means the opposite — possible duplicates, never loss).
- Claiming Kafka is exactly-once by default (default is at-least-once).
- Thinking exactly-once means each record is physically transmitted only once on the wire (it means each record affects state once; retries/dupes still occur underneath and are deduped).