Two retries carrying the same Idempotency-Key HTTP header arrive at your API at the same moment, before the first has finished. What must the server guarantee, and how do you implement it?
answer
- At most one execution per key, even simultaneously
- INSERT-first under unique index = the claim
- Loser: 409 in-progress (or 425 / short wait-and-replay)
- Lease on in_progress rows, else the key is poisoned by a crash
- Effect + completion in one transaction
basics
~20 sExactly one may execute. Claim the key with an INSERT under a unique constraint before doing any work — the database picks the winner. The loser sees an in-progress record and returns 409 Conflict (or waits briefly), never executing. In-progress rows need a lease so a crash doesn't block the key forever.
solid answer
~50 sThe guarantee is **at most one execution per key**, even under simultaneous arrival. A `SELECT`-then-execute check is a check-then-act race: both requests find no record and both charge the card. The fix is to make the write the claim — `INSERT` an `in_progress` row keyed by `(scope, key)` with a unique index, before any business work. Exactly one insert succeeds; the other gets a unique-violation and knows a sibling owns the key. The loser's options: return **409 Conflict** with a "request already in progress, retry later" error (Stripe's behavior), return **425 Too Early**, or briefly poll for completion and then replay the stored response. 409 is simplest and honest — the client already retries. Two further requirements. **Lease the in-progress row**: if the winner crashes mid-flight, an expired lease must let a later retry take over, otherwise the key is poisoned permanently. And **commit the effect and the completed record together**, so a replay never reports success for work that rolled back.
code
http · 8 linesPOST /v1/payments HTTP/1.1
Idempotency-Key: 8f14e45f-ceea-467a-9f5a-3c3f2f0d1b77
HTTP/1.1 409 Conflict
Content-Type: application/json
{"error":{"code":"idempotency_key_in_progress",
"message":"A request with this Idempotency-Key is currently being processed. Retry shortly."}}go deeper
State the guarantee — only one of the two may execute — and that the database's unique constraint should decide the winner rather than an application check.
Explain the insert-first claim, what the loser returns (409 in-progress, retryable), and why select-then-execute races.
Own the failure modes: leases for crashed in-progress rows, takeover semantics, transactional completion with the effect, and the trade between returning 409 and waiting to replay under load.
Reason about it as a system property — where the mutual-exclusion primitive lives, reconciliation with downstream providers via key pass-through, and how the choice behaves during a fleet-wide retry storm.
## The scenario A client times out and retries. Or a mobile app fires two requests from a double tap. Or a load balancer duplicates a request across two backends. Two requests with the same `Idempotency-Key` are now being handled simultaneously by different processes, and the first has not written a result yet. If both execute, the customer is charged twice — the exact outcome the pattern promised to prevent. This is where naive implementations fail. Storing keys is easy; handling the concurrent window is the real engineering. ## Why check-then-act loses ``` row = SELECT * FROM idempotency WHERE scope=? AND key=? if row is null: result = chargeCard() INSERT INTO idempotency (...) ``` Both requests run the SELECT before either runs the INSERT. Both see null. Both charge. The window is small, which is exactly why this bug survives testing and surfaces in production during a network blip that makes clients retry en masse. ## Claim by writing Make the **insert** the claim: ``` INSERT INTO idempotency (scope, key, fingerprint, state, lease_expires_at) VALUES (?, ?, ?, 'in_progress', now() + interval '60 seconds') -- unique index on (scope, key) ``` The database's unique index is the arbiter. One request gets a row; the other gets a unique-constraint violation, which is not an error condition but information: *someone else owns this key*. No distributed lock service is required — the constraint you already have is the mutual exclusion primitive. Only the winner proceeds to execute. ## What the loser returns Three defensible designs: **409 Conflict** with a specific error code such as `idempotency_key_in_progress`. This is what Stripe does. It is honest, cheap, and holds no server resources. The client is already a retrying client, so it backs off and tries again, and by then the record is completed and it gets a replay of the real response. Downside: the client sees an "error" for what is actually a success in flight, so the error must be documented as retryable. **425 Too Early.** Semantically closer to "come back later" and less likely to be logged as a failure by naive clients, but far less common, so clients may not handle it specially. **Block and then replay.** The loser waits (polling the row, or on a condition variable / advisory lock) up to a short deadline, then returns the completed stored response. The client experience is best — it gets the real answer on the first call. The costs are real though: you hold a request thread or connection for the duration, which is a availability hazard under a retry storm, and you need a hard deadline after which you fall back to 409 anyway. Reasonable when the operation is fast and the client fleet is large and unsophisticated. What is **not** acceptable: executing the operation, or returning a synthetic success. ## Crashed in-progress rows If the winner's process dies after inserting `in_progress` and before writing the result, the key is now claimed by nobody. Without recovery, every future retry sees `in_progress` and gets 409 forever — a permanently poisoned key, and an operation the client can never complete or confirm. The fix is a **lease**: store `lease_expires_at` on the in-progress row. A later request that finds an expired lease may take over the key (atomically, with a conditional update that re-stamps the lease so only one taker wins). Set the lease longer than the maximum plausible execution time, because taking over a key whose original execution is still running reintroduces double execution. The deeper safety net for money operations is **reconciliation with the downstream system**: pass your key through to the payment provider as their idempotency key, so a takeover that re-calls them cannot double-charge — the provider deduplicates on the same key. ## Atomicity of completion Write the business effect and the transition to `completed` in one transaction whenever they share a database. Otherwise: - Mark completed first, then the effect rolls back → replays report a success that never happened. - Effect commits, then completion write fails → a retry re-executes, double effect. Co-locating the dedupe table with the business data buys you this for free and is the main argument against putting the dedupe store in a separate cache. ## What to say in an interview "At most one execution per key. I claim the key with an INSERT under a unique index before any work, so the database arbitrates instead of an application-level check-then-act. The loser returns 409 with a retryable in-progress error rather than executing. In-progress rows carry a lease so a crashed request doesn't poison the key, and the effect and the completed record commit in the same transaction so a replay never reports work that rolled back."
- Why not just take a distributed lock in Redis for the key?You can, but it adds a second system with its own failure modes — lock expiry during a slow execution, split brain on failover — and it still does not give you atomicity between the effect and the dedupe record. If the dedupe table already lives in your primary database, its unique index provides the same mutual exclusion for free and inside the same transaction as the effect.
- What happens if the process holding an in-progress key crashes?Without recovery the key stays in progress forever and every retry gets 409 — the operation can never complete. A lease with an expiry lets a later request take the key over once the lease lapses, using a conditional update so only one taker wins. The lease must exceed the maximum plausible execution time, or a takeover can run concurrently with a still-live original.
- Is returning 409 to the concurrent duplicate a good client experience?It is acceptable because the caller is already a retrying client and will get the real response on its next attempt. The alternative — waiting briefly and then replaying the completed response — gives a better first-call experience but holds server threads during exactly the traffic spikes where duplicates are most common, so many APIs choose 409 for its availability properties.
Two people reach for the last seat: you don't resolve it by both looking and both sitting — you resolve it by whoever gets the ticket stub first, and the ticket machine issues only one.
saying these in an interview costs you the question
- Checking with a SELECT and then executing — both concurrent requests proceed.
- Assuming duplicates are rare enough that the race can be ignored.
- Blocking indefinitely while waiting for the in-flight request, with no deadline.
- Leaving in-progress rows with no lease, permanently poisoning the key after a crash.
- Marking the key completed outside the transaction that commits the effect.