skip to content

In distributed systems, what is the difference between at-most-once, at-least-once, and exactly-once message delivery, and which of these can actually be guaranteed over an unreliable network?

level: juniorimportance: must knowfreq 85%

answer

  1. at-most vs at-least vs exactly-once
  2. Two Generals Problem - ack of the ack
  3. effectively-once = at-least-once + idempotency
  4. timeout is ambiguous: lost request vs lost response

basics

~20 s

At-most-once means a message might get lost but never duplicated. At-least-once means it might arrive more than once but never lost. Exactly-once (arrives precisely one time) can't be truly guaranteed when networks can drop or delay messages - only faked using retries plus deduplication.

solid answer

~30 s

At-most-once delivery sends a message once and gives up on failure - fast but lossy. At-least-once retries until acknowledged, guaranteeing delivery but risking duplicates. True exactly-once is impossible over an unreliable network because a sender can never distinguish 'my message was lost' from 'my acknowledgment was lost' - so it can't know whether to retry. What systems call 'exactly-once' in practice is really at-least-once delivery plus idempotent processing on the receiver, producing an effectively-once outcome. The choice among the three is a trade-off between simplicity, latency, and correctness under failure.

go deeper

for a junior

Should state the three categories correctly and know that at-least-once is the common real-world default, even without a deep explanation of why exactly-once is impossible.

for a middle

Should explain the timeout ambiguity (lost request vs. lost response) as the reason exactly-once delivery can't be guaranteed, and name idempotency as the practical fix.

for a senior

Should connect this to the Two Generals Problem by name, explain 'effectively-once' precisely, and give at least one real system (Kafka EOS, Stripe idempotency keys) that implements it correctly.

for a principal

Should reason about where the 'effectively-once' boundary leaks in multi-system pipelines (e.g., a Kafka EOS pipeline writing to an external non-transactional API) and how to design idempotency boundaries at each hop.

## What a delivery guarantee promises Delivery semantics describe the guarantee a messaging system makes about how many times a message sent by a producer will be observed by a consumer, given that networks, processes, and machines can fail at any point. There are three canonical categories, and understanding why only two of them are actually achievable is foundational to distributed systems design. ## At-most-once: simplicity paid for with silent loss **At-most-once delivery** means the sender transmits a message a single time and does not retry, regardless of whether it receives confirmation. If the message or its acknowledgment is lost in transit, the message is simply gone. This is the semantics you get 'for free' from a bare fire-and-forget call, or from a system that explicitly disables retries. - **Its virtue is simplicity and low latency:** no retry logic, no deduplication bookkeeping, no risk of duplicate side effects. - **Its cost is silent data loss** - a metrics sample or a log line can vanish and nobody notices. - **It is appropriate only when occasional loss is tolerable** (best-effort telemetry, heartbeats) and never when correctness depends on every event landing. ## At-least-once: the practical default **At-least-once delivery** means the sender keeps retrying until it receives confirmation that the message was received. It retries on: - a timeout - a missing acknowledgment - an explicit failure signal This guarantees no message is silently dropped, but it introduces the possibility of duplicates: if the message actually arrived and was processed, but the acknowledgment got lost on the way back (or the consumer crashed after processing but before acking), the sender has no way to tell the difference between 'nothing happened' and 'everything happened, only the ack got lost.' From the sender's point of view, both look identical - a timeout. So it retries, and the receiver may see the same logical message twice. Almost every practical durable messaging system (default **Kafka** producers, standard **SQS** queues, **HTTP** retry-on-timeout logic composed with application-level acking) is built on at-least-once semantics, because guaranteeing 'not lost' is achievable and cheap, while eliminating duplicates at the transport layer is not. ## Why exactly-once delivery is out of reach **Exactly-once delivery** - a message observed by the consumer precisely one time, no more, no less - is the semantics everyone wants and, as a pure delivery guarantee, cannot be achieved over an unreliable network. The reasoning traces to the **Two Generals Problem**: two parties trying to reach agreement over a channel that can lose messages can never be fully certain the other side received the last message, because the acknowledgment of that message could itself be lost, requiring an acknowledgment of the acknowledgment, ad infinitum. Concretely: the sender transmits, then waits. If nothing comes back, it cannot distinguish three cases: 1. the request was lost before reaching the receiver; 2. the request arrived and was processed but the response was lost on the way back; 3. the request arrived and is still being processed slowly. The first case calls for a safe retry; the second calls for not reprocessing; the sender cannot tell them apart from a timeout alone, and no finite number of extra round trips removes this ambiguity. ## What production systems actually sell you What production systems market as 'exactly-once processing' is not exactly-once delivery at the network layer at all: - **Kafka's** transactional plus idempotent producer with `read-committed` consumers - **Flink's** exactly-once checkpointing - **Stripe's** idempotency-key API It is at-least-once delivery at the transport layer, combined with idempotent processing at the application or storage layer, so that even though the same message may be delivered and handled multiple times, its side effects are applied only once. This composite guarantee is usually called **'effectively-once'**: the wire can duplicate freely, but the receiver recognizes a duplicate and discards or no-ops on the repeat. This is the crucial mental shift for engineers: exactly-once is not a delivery property you can buy from your message broker - it's an engineering outcome you build by pairing at-least-once delivery with idempotency on the consumer side. ## The concrete case, and where it leaks A concrete example: **Kafka's exactly-once semantics (EOS)** for stream processing combines 1. an idempotent producer - each message carries a producer ID and sequence number so the broker can drop duplicate retries at write time; 2. transactional writes across `input-offset-commit` and `output-produce`. Even here, the guarantee is scoped to within the Kafka cluster - the moment that pipeline writes to an external system without its own idempotency mechanism, the 'exactly-once' property leaks and duplicates can reappear, which is why idempotency keys and dedup stores matter even in ecosystems that advertise exactly-once out of the box.

  • Why can't we just add a third round trip - sender asks 'did you get it?', receiver confirms, sender confirms the confirmation - to nail down exactly-once?
    Because that confirmation message can be lost too, and now the receiver doesn't know whether the sender got ITS confirmation. Each additional round trip just relocates the ambiguity to a later message rather than eliminating it; there is no finite protocol over a lossy channel that removes the uncertainty entirely, which is the core insight of the Two Generals Problem.
  • If TCP already guarantees ordered, non-duplicated byte delivery, why do we still worry about duplicate messages at the application layer?
    TCP's guarantee stops at the transport connection - it prevents duplicate/out-of-order bytes on a single live connection, but says nothing about what happens when the connection drops and the application retries at a higher level. The retry is a brand-new transmission from the application's perspective, so duplication reappears one layer up.
  • Is at-most-once ever the right default for a system, or is it always a bug?
    It's a deliberate, correct choice for high-volume, loss-tolerant data like metrics, non-critical logs, or best-effort presence pings, where the cost of a retry outweighs the cost of occasionally missing one data point. It becomes a bug when applied to anything with business consequence, like payments or order state transitions.

Mailing a letter with no tracking is at-most-once - if it's lost, you never know. Re-mailing a certified letter every week until you get a signed receipt is at-least-once - the recipient might get several identical copies before your receipt arrives. Getting the letter delivered exactly once, with total certainty, would require a postal system that can never lose a truck - which doesn't exist, so instead the recipient just throws away the duplicate letters they recognize.

saying these in an interview costs you the question

  • Claims a message broker or API can deliver 'exactly-once' at the network/delivery layer without mentioning idempotency
  • Thinks TCP alone solves application-level duplicate delivery
  • Can't explain why a sender-side timeout is ambiguous
  • Doesn't know at-least-once is the practical default for durable systems

context