skip to content

How do you design a case that proves a write request is idempotent when the client replays it?

level: seniorimportance: nice to knowfreq 33%

answer

  1. Same key, sent twice
  2. Count the effects, not the bytes
  3. Safe and idempotent are different properties
  4. A lagging read view fakes a pass
  5. The real replay follows a client timeout

basics

~20 s

Send the identical request twice with the same caller-supplied idempotency key, then assert the observable state changed once — one record, one downstream effect — reading it back through an authoritative path rather than a cached view.

solid answer

~50 s

The oracle is state, not the second response. I send the same request twice with the same idempotency key and then assert that exactly one record exists and exactly one downstream effect was produced — one hold, one notification — reading state back through a path the write is guaranteed to have updated, not through a cached view that may lag. The second response does not have to be byte-identical: a contract may legitimately answer the replay with a different status, so I assert what the contract states rather than what feels intuitive. I also replay the case that matters in practice — the first attempt timed out at the client after the write applied — because that is the situation idempotence exists for. Keys are generated per run so the case never depends on data another run left behind.

code

pseudocode · 13 lines
pseudocode
key = "hold-" + runId + "-" + caseId
request = { flight: "OP417", seat: "22C", passengerRef: runId }

first  = client.post("/seat-holds", request, headers = { "Idempotency-Key": key })
second = client.post("/seat-holds", request, headers = { "Idempotency-Key": key })

assert first.status == 201
assert second.status in contract.replayStatuses      # 200 or 409, per contract
assert second.body.holdId == first.body.holdId

holds = client.get("/seat-holds?passengerRef=" + runId).body.items
assert holds.size == 1
assert notifications.countFor(holdId = first.body.holdId) == 1

go deeper

for a junior

Be able to state the property plainly: performing the operation once or several times leaves the same state. Knowing that a read is safe while a delete is merely idempotent is enough at this level, along with the idea that a retry can duplicate a create.

for a middle

Explain the mechanism — a caller-supplied key the service records with the first outcome — and describe the case shape: send twice with one key, then assert the number of records and downstream effects rather than comparing the two responses.

for a senior

Show the judgement: the oracle is scoped state read through a path with a real consistency guarantee, the contract decides what the replay's status may be, and the replay worth simulating is the one after a client timeout. Be able to name why such a case goes false-green.

for a principal

Own where the property is required at all: which operations carry a retrying caller, whether keys are a platform convention or reinvented per service, how long outcomes are retained, and whether idempotence is verified once at the edge or demanded of every downstream consumer.

Idempotence is the property that performing an operation once and performing it several times leave the system in the same state. It is not the same as safety: a safe operation makes no intended state change at all, so a read is both safe and idempotent, a delete is idempotent but not safe, and a create is usually neither — unless the service is given something that lets it recognise a repeat. That something is normally a caller-supplied idempotency key: a value the client generates for a logical attempt and repeats on every retry of it, which the service stores alongside the first outcome so a later request with the same key returns that outcome instead of doing the work again. ## The oracle is state, not the response The most common broken design asserts that the two responses are identical and calls it proved. They may be identical while two records were created; they may differ while the operation was perfectly idempotent. The evidence that settles it is the observable state afterwards: how many records exist, how many downstream effects were emitted, what the aggregate totals now say. So the case shape is: perform, replay, then assert a scoped state count — scoped to the identifiers this case created, never an absolute count against shared data that other cases are also writing to. ## Read state through a path that cannot lie On an airline seat-map service, a replay case looked healthy for months. It held seat 22C twice with the same key, then read the flight's seat map back and asserted one hold. The seat map was served from a projection refreshed on a schedule, so both reads returned the same stale view — the one taken before either write. The case would have passed even if the second request had created a duplicate hold, because the read could not see either. A stale-cache read is the classic false green in replay cases, and the defence is to assert through a path with a defined consistency guarantee: the authoritative record endpoint, a query the contract states is read-your-writes, or the same projection after the condition it exposes has actually changed. ## What the second response is allowed to be Contracts vary, legitimately. Some return the original success and status on the replay; some return the created resource with a plain success status where the first returned a creation status; some answer a conflict status and expect the caller to treat it as "already done". All three can be idempotent. The case must assert the contract's stated behaviour and say in its name which it is — a case that silently assumes the response repeats verbatim will fail the day the service adds a correct, documented conflict answer. ## The replay that actually happens in production The interesting replay is not a tidy second call. It is the first call whose response never arrived: the client timed out, the write had already applied, and the retry went out with the same key. A case worth writing simulates that — induce the timeout on the client side while letting the request complete, then replay — because it exercises the window where the key is recorded but the outcome is not yet final. Two further dimensions are worth a case each when the contract claims them: two identical requests issued concurrently, where the service must serialise on the key rather than racing; and a replay sent after the key's retention window has expired, where the contract may correctly permit a second effect and the case must assert that rather than the opposite. ## Keeping the case honest Generate the key per run, from the run's own identifier, so a case never inherits an outcome another run recorded — a suite that reuses a fixed key passes the first time and then passes for the wrong reason forever. Scope assertions to the entities this case created so parallel runs do not read each other's rows. Clean up what the case created, or make the identifiers unique enough that leftovers are harmless. And be explicit about what the case does not prove: idempotence at the service interface says nothing about whether every internal consumer of the resulting event is also idempotent, and a duplicate downstream effect is a defect the state assertion should be extended to catch rather than assumed away. ## Where the value lands This is not an everyday screening question, but it separates people who have run a retrying client from people who have not. Any caller with a retry policy — a mobile client on a flaky connection, a queue consumer, a gateway that retries a timeout — will send duplicates eventually, and the cost of getting it wrong is money or inventory rather than a cosmetic defect. Being able to design the replay case, name its oracle and explain why the second response is the weakest possible evidence is a strong signal at senior level.

  • How is an idempotent operation different from a safe one?
    A safe operation is not intended to change state at all, so repeating it is trivially harmless — a read is the example. An idempotent operation may change state, but repeating it leaves the same state as performing it once: deleting a record that is already deleted, or setting a field to a fixed value. A create is normally neither, which is why it needs a caller-supplied key before a replay can be recognised.
  • What makes a replay case flaky, and how do you keep it honest?
    Three things: a fixed idempotency key that makes every run after the first pass for the wrong reason; absolute state counts that other runs are concurrently changing; and a read-back path that lags the write, which can pass while a duplicate exists. Generate the key from the run identifier, scope every count to entities this case created, and read through a path whose consistency the contract actually guarantees.
  • Why is asserting only that the replay returned no error a weak oracle?
    Because the failure mode is a duplicate effect, and a duplicate is usually created successfully. A second hold, a second charge and a second notification all return a perfectly healthy response. Without a state assertion the case cannot distinguish "the service recognised the repeat" from "the service happily did it twice", which is the exact distinction the case exists to make.

saying these in an interview costs you the question

  • Asserting the two responses are byte-identical
  • Calling an operation idempotent because the retry did not error
  • Reading state back through a view that lags the write
  • Confusing idempotent with safe — no state change at all
  • Reusing one fixed idempotency key across every run
  • Testing only the tidy replay, never one after a client timeout

context