skip to content

A client retries an append call to an event store after a network timeout, but the original append had actually already succeeded server-side before the timeout. How do idempotency keys prevent this from creating a duplicate event?

level: seniorimportance: should knowfreq 45%

answer

  1. timeout is ambiguous: sent? lost? committed?
  2. client generates event ID once, reused on retry
  3. store dedups by ID -> no-op on repeat
  4. duplicate event != self-correcting like a row overwrite
  5. ID must be fixed before retryable I/O, not regenerated per attempt

basics

~20 s

Each write carries a unique ID. If the same ID shows up twice because of a retry, the store recognizes it already processed that exact write and just confirms success again instead of adding a second copy.

solid answer

~50 s

A timeout only tells the client it didn't receive a response — it doesn't tell it whether the server-side append actually committed or not, so a naive retry that just calls append again risks writing the same logical event twice. Idempotent append handles this by having the client attach a stable, unique identifier to the write (often the event's own ID, or a client-generated request ID) before the first attempt. The store persists that identifier alongside the event and, on any subsequent append carrying the same identifier, detects the duplicate and returns success without inserting a second event — instead of appending or erroring. This turns 'at-least-once delivery from an unreliable network' into an effective 'exactly-once outcome' at the storage layer, which is critical in event sourcing because a duplicated event (e.g., a second MoneyDeposited) directly corrupts the derived state, not just a log entry.

go deeper

for a junior

Should grasp that retrying after a timeout could accidentally send the same thing twice, and that some ID is used to catch that.

for a middle

Should explain that the client generates the ID once and the server checks for it, and know why the timeout is ambiguous.

for a senior

Should articulate why duplicate events are worse than duplicate table rows (non-self-correcting, propagate to every projection) and know the common bug of regenerating the ID inside a retry loop.

for a principal

Should design idempotency across multi-stream operations (outbox/saga-level keys) and reason about retention/bookkeeping trade-offs for dedup tracking at scale.

## Why a timeout is ambiguous The core problem idempotent append solves is that a network timeout is fundamentally ambiguous from the client's point of view. When a client calls `append(streamId, expectedVersion, events)` and the connection times out before a response arrives, the client cannot distinguish between three very different outcomes: - **(a)** the request never reached the server, - **(b)** the request reached the server and was rejected, or - **(c)** the request reached the server, was committed successfully, and only the response was lost on the way back. In cases (a) and (b) it's safe (even necessary) to retry the exact same append. In case (c), retrying naively means appending the same logical event a second time — and because an event store's default behavior on a fresh append call is simply to append, without any built-in notion of 'have I seen this exact write before,' a naive retry produces a genuine duplicate: two `MoneyDeposited(amount: 100)` events in the stream where there should be one, silently inflating the derived balance by 100 the next time the stream is replayed. ## How idempotent append closes it Idempotent append closes this ambiguity by attaching a stable identity to the write itself, independent of the network transport. 1. The most common mechanism is to have the client generate the event's ID (often a UUID) **once, before the first attempt**, rather than letting the server assign it — so a retry of the same logical operation carries the identical event ID as the original attempt, not a new one. 2. The event store persists this ID alongside the event when it commits the append, typically indexed for fast lookup. 3. On any append call, the store checks whether it has already committed an event with that exact ID in that stream; if so, it treats the call as a **no-op duplicate** and returns the same success response the original commit would have produced (often including the position/version the event actually landed at), rather than inserting a second event or throwing an error. From the client's perspective, retrying a timed-out append is now always safe: either the retry is genuinely the first successful attempt (server never got it), or it's a harmless duplicate detection that confirms 'yes, that already happened, here's where.' ## Why a duplicate event is worse than a duplicate row This matters more in event sourcing than in a typical CRUD system because the event log is the literal source of truth that all derived state is computed from — a duplicate row in a mutable table is often self-correcting (the next `UPDATE` just overwrites it again), but a duplicate event is **not self-correcting**: it's a permanent, immutable extra fact that every future replay will fold in, permanently double-counting the deposit, permanently double-shipping the order, forever, unless someone notices and appends a manual correction. The blast radius of a duplicate event is also wider than a single table row: every projection/read-model built by consuming that stream inherits the duplication independently, so fixing it after the fact means correcting the event store and every downstream materialized view. ## The trade-off The trade-off idempotent append introduces is bookkeeping cost and a design constraint: - the store (or a companion table) needs to **track seen-IDs per stream** indefinitely or for some bounded retention window, - and clients must be disciplined about **generating the ID once and reusing it verbatim across retries** rather than accidentally minting a fresh ID on each attempt — which is a surprisingly easy bug to introduce if the event ID is generated inline at call time instead of being fixed earlier in the request's lifecycle (e.g., at the point the command was received, before any retryable I/O). There's also a scope question: idempotency keyed per-stream is enough to prevent double-appends to that stream, but if the same logical business operation could legitimately touch multiple streams (a transfer debiting one account and crediting another), true end-to-end idempotency needs a higher-level mechanism — often an outbox/saga pattern with its own dedup key — spanning both appends. ## A concrete failure mode A concrete failure mode: a payment service calls append with a freshly-generated UUID for the event ID on every call, including retries, because the UUID generation sits inside the retry loop instead of outside it. Every network hiccup then produces a genuinely distinct event ID on retry, so the store's duplicate-detection has nothing to match against, and the 'idempotent' append silently isn't — the team only discovers this when reconciliation between ledger totals and bank statements starts drifting under load. ## Where it shows up A concrete real-world example: EventStoreDB deduplicates appends within a stream based on the event's ID field specifically for this reason — client SDKs are documented to generate the event ID client-side, once, precisely so that transport-level retries (which the client library issues automatically on transient network errors) land on the same ID as the original attempt and get deduplicated server-side rather than double-appended.

  • Why can't the event store just dedupe based on the event's content instead of a separate ID?
    Content-based dedup would incorrectly collapse two genuinely distinct events that happen to have identical payloads, like two separate $50 deposits made minutes apart, which are legitimate separate business facts, not duplicates. A caller-assigned unique ID is the only reliable signal that two append attempts represent the same logical operation.
  • What happens if idempotency is only enforced per-stream, but a business operation writes to two streams?
    Per-stream dedup only protects each individual stream from a duplicate append to itself; it doesn't guarantee the two writes happen together atomically or that a retry can't duplicate one while correctly deduping the other. That case needs a higher-level pattern, like an outbox/saga with its own operation-level idempotency key, on top of per-stream dedup.
  • Where should the retryable event ID be generated to make idempotency actually work?
    It must be generated once, before the first network attempt — for example when the command is first received — and reused verbatim on every retry of that same attempt. If it's generated fresh inside the retry loop, every retry gets a new ID, defeating dedup entirely.

Like putting a tracking/confirmation number on a mailed form before you send it: if the post office loses the confirmation reply and you're not sure your form arrived, you resend it with the same number, and the receiving office recognizes the number and discards the duplicate instead of processing your request twice.

saying these in an interview costs you the question

  • thinks a timeout means the write definitely failed and is always safe to blindly resend
  • proposes deduping on event payload content instead of an explicit ID
  • generates the idempotency/event ID inside the retry loop instead of before the first attempt
  • doesn't realize a duplicate event corrupts every downstream projection, not just the store

context