skip to content

questions

5

In event-driven messaging systems, what do at-most-once, at-least-once, and exactly-once delivery mean, and which one do most production systems actually rely on by default?

level: juniorimportance: must knowfreq 85%

answer

  1. commit before vs after processing
  2. loss vs duplication
  3. two independent actions problem
  4. dual-write problem
  5. auto-commit interval trap

basics

~20 s

At-most-once can lose a message but never repeats it. At-least-once never loses a message but can repeat it. Exactly-once means every message is processed once, no loss and no duplicates. Most real systems use at-least-once plus deduplication.

solid answer

~30 s

These describe what happens to a message when a consumer crashes mid-processing. At-most-once: the offset/ack is recorded before processing, so a crash loses the message but never reprocesses it. At-least-once: the offset is committed only after processing succeeds, so a crash after processing but before commit causes reprocessing — safe against loss, but consumers must tolerate duplicates. Exactly-once combines at-least-once delivery with deduplication or transactional writes so the net effect looks like each message was applied once. Most production systems default to at-least-once because it's the easiest correct baseline and push idempotency to the consumer.

go deeper

for a junior

Should recite the three definitions correctly and know at-least-once is the common default, without necessarily explaining the commit-ordering mechanism.

for a middle

Should explain the commit-before/commit-after mechanism that produces each guarantee and give one concrete example, such as auto-commit vs manual commit in Kafka.

for a senior

Should connect the taxonomy to the dual-write problem, articulate the cost side of each choice, and describe at least one production failure mode like a rebalance duplicate or an auto-commit loss.

for a principal

Should reason about which guarantee to choose per data-flow in a heterogeneous system, weighing operational cost (dedup storage, transaction overhead) against business cost of loss/duplication, and design boundary contracts between components accordingly.

## Two actions per message Delivery guarantees describe the relationship between two actions a consumer must perform for every message: - **(1) Do the work the message triggers** — write to a database, call an API, update in-memory state. - **(2) Record that the message was handled** — typically by committing an offset (Kafka), acknowledging a message (RabbitMQ/JMS), or deleting it from a queue (SQS). These two actions cannot be made atomic for free because they usually touch two different systems: the message broker and whatever the consumer's side effect touches. The order in which you do them determines which guarantee you get. - **Commit/ack before processing** gives at-most-once: if the consumer crashes after the commit but before finishing the work, the message is gone forever as far as the broker is concerned, because the broker already believes it was delivered successfully. - **Commit/ack after processing** gives at-least-once: if the consumer crashes after finishing the work but before the commit, the broker still holds the message (or an uncommitted offset) and will redeliver it on restart or rebalance, so the work is done twice. **Exactly-once** is not a third position in that ordering — it is at-least-once combined with a mechanism that makes reprocessing invisible, either by writing the output and the offset atomically in one transaction, or by making the side effect idempotent so replays don't change the outcome. ## Why the taxonomy exists The taxonomy exists because 'process a message' is not a single operation — it is at minimum two operations (mutate state, record progress) that live on different failure domains and cannot share a transaction unless you deliberately build one. This is the same fundamental issue as the **dual-write problem** in distributed systems: whenever two systems must agree on the outcome of one logical unit of work and they don't share a transaction manager, you must choose which one you trust to fail safe. Message systems formalized three named points on that spectrum because engineers kept re-discovering the same three failure modes — silent loss, harmless-if-handled duplication, and the expensive all-guarantees option — and needed shared vocabulary, especially for financial and inventory systems where 'processed twice' and 'processed never' have very different costs. ## What each one costs - **At-most-once** is cheap and fast — no need to persist processing state, no idempotency logic — but it is only acceptable where losing an occasional event is tolerable, like a live metrics gauge or a best-effort push notification. - **At-least-once** is the workhorse default: it costs you having to make downstream operations tolerate duplicates (an insert becomes an upsert, a 'charge card' call becomes 'charge if not already charged for this order'), but that is usually a one-time design investment. - **Exactly-once** costs the most: it either requires a transactional coordinator spanning the broker and the sink, which adds latency and throughput overhead and only works when both ends are transaction-aware, or it requires persistent dedup state that must be looked up on every message and grows with the number of distinct keys you need to remember, which is itself an availability and storage liability. ## Failure modes - The most common production symptom of at-least-once is a **duplicate charge, email, or row** that surfaces only under a real crash or a consumer-group rebalance, not under normal load, so it's routinely missed in testing and discovered in an incident review. - The most common symptom of accidentally-at-most-once is **silent data loss**: a team enables auto-commit with a short interval to reduce duplicate work, a consumer starts throwing exceptions mid-batch, and offsets already advanced past the failed messages, so nobody notices until a downstream reconciliation job reports missing records days later. - A subtler failure mode is **broker-side redelivery** even when the consumer did everything right: a network partition can cause a broker to redeliver a message whose ack was in flight and actually received, because the broker cannot distinguish 'ack lost in transit' from 'consumer never processed it.' ## Walking it through in Kafka Concretely, a Kafka consumer with `enable.auto.commit=true` and a five-second commit interval is at-most-once-ish under crash conditions, because the offset can advance on a timer independent of whether the record finished processing. Switching to manual `commitSync()` called only after the business logic returns successfully turns it into at-least-once: a crash before the commit simply replays the record on restart. A payment-processing consumer built this way must then either wrap the debit and the commit in one transaction, or make the debit idempotent by keying it on the order ID so a replay is a no-op — the exact boundary this leaf sits on before handing off to consumer-side idempotency patterns.

  • Why can't a consumer just make processing and offset-commit a single atomic step by default?
    Because they usually live in two different systems — the broker's log and whatever the consumer's side effect touches — and atomicity across two independent systems normally requires an explicit distributed transaction or protocol, which most brokers and sinks don't support out of the box. Kafka's transactional producer is one of the few mechanisms that makes this atomic, but only when both the read and the write happen through Kafka itself.
  • If a team wants at-most-once on purpose, when is that actually the right call?
    When the cost of losing an occasional message is lower than the cost of building and maintaining duplicate-tolerant logic — e.g., streaming live telemetry gauges, best-effort cache invalidation, or fire-and-forget analytics pings where the next event supersedes the lost one anyway.
  • How does a Kafka consumer-group rebalance make duplicates more likely even under a correct commit-after-processing design?
    A rebalance can reassign a partition to a different consumer instance mid-flight; if the original consumer had finished processing but not yet committed when it lost the partition, the new owner starts from the last committed offset and reprocesses the same records the first instance already handled.

Like a mail carrier deciding whether to mark a letter 'delivered' before or after actually placing it in the mailbox — mark it first and a dropped letter is lost forever; mark it after and a distracted carrier interrupted mid-delivery might deliver two copies when told to start over.

saying these in an interview costs you the question

  • Says exactly-once means no duplicates can ever reach the consumer, full stop
  • Doesn't know the loss/duplication direction is set by commit-before vs commit-after processing
  • Treats at-least-once as free/default with no mention of needing duplicate-tolerant consumers
  • Thinks enabling a broker feature like Kafka transactions alone makes an entire pipeline exactly-once
  • Can't name a concrete scenario where at-most-once is an acceptable, deliberate choice

context

open as a page

When a Kafka consumer processes a batch of 10 records and commits only the offset of the last record after the whole batch succeeds, what happens on a crash after record 6 finishes, and why is committing after every single record not simply 'safer'?

level: middleimportance: must knowfreq 75%

basics

~20 s

Committing the offset once per batch is faster but replays the whole batch if a crash happens partway through, since Kafka only tracks one checkpoint per partition, not per message. Committing after every single message shrinks that replay window but is much slower.

open as a page

Why do practitioners often say Kafka's 'exactly-once semantics' are really 'effectively-once' once you look at the whole pipeline end-to-end, from source system through to the final side effect?

level: seniorimportance: must knowfreq 70%

basics

~20 s

Because 'exactly-once' guarantees usually only cover the messaging system itself. The moment a message triggers an outside effect — an email, an API call, a non-transactional database write — that effect can still happen more than once, so the real guarantee is 'looks like exactly once if downstream systems are built to tolerate replay.'

open as a page

How does Kafka's transactional producer combine the idempotent producer with a transaction coordinator to give exactly-once semantics for a read-process-write pipeline, and what exactly does that atomicity cover?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Kafka can bundle writing output records and committing the input offset into one all-or-nothing transaction, using a producer ID and sequence numbers to avoid duplicate writes on retry, so a crash mid-write either commits everything or nothing — but only within Kafka.

open as a page

As a principal engineer designing a payment pipeline that spans a Kafka topic, an external payment gateway, and a legacy non-transactional data warehouse sink, how would you decide where to enforce which delivery guarantee, and what's the risk of applying the same guarantee uniformly everywhere?

level: principalimportance: should knowfreq 40%

basics

~20 s

Different parts of the pipeline need different guarantees based on how costly a duplicate or a loss is there — you can't just pick one setting for the whole system. Use real exactly-once only where it's cheap, like Kafka-to-Kafka, and layer at-least-once plus deduplication everywhere the guarantee has to reach outside Kafka, matching the mechanism to each sink's actual risk and cost.

open as a page