skip to content

You are implementing support for the Idempotency-Key request header on a POST endpoint. What does the server store, and what is the request-handling flow for a first request versus a replay?

level: middleimportance: must knowfreq 55%

answer

  1. Row: (scope,key) + fingerprint + state + stored response + expires_at
  2. INSERT claims the key; never SELECT-then-act
  3. Same transaction as the effect where possible
  4. Replay returns the original status/body verbatim, flagged replayed
  5. Don't freeze 5xx — the key would be bricked

basics

~20 s

Store a row keyed by (scope, key) holding a fingerprint of the request, a state (in-progress / completed), and the saved status, headers and body. First request: insert the row, execute, save the response. Repeat: find the completed row and return the stored response without re-executing.

solid answer

~50 s

**The record**: `(scope, idempotency_key)` as the primary key, plus a request fingerprint (hash of method, path and body), a state (`in_progress` / `completed`), the stored response status/headers/body, `created_at` and `expires_at`. **First request**: insert the row as `in_progress` under a unique constraint — the insert, not a prior SELECT, is what claims the key. Execute the business operation, then write the response into the row and mark it `completed`, ideally in the same transaction as the effect. Respond. **Replay**: the lookup finds a `completed` row. Verify the fingerprint matches, then return the stored status, headers and body verbatim, usually flagged with a response header such as `Idempotent-Replayed: true`. Do **not** re-execute. **Row already `in_progress`**: a concurrent retry is running — respond 409 Conflict (or 425 Too Early) rather than executing. Store only outcomes you want frozen: a 5xx normally should not be locked in, so the client's retry can genuinely re-attempt.

code

json · 11 lines
json
{
  "scope": "acct_123",
  "idempotency_key": "8f14e45f-ceea-467a-9f5a-3c3f2f0d1b77",
  "request_fingerprint": "sha256:1c9f...",
  "state": "completed",
  "response_status": 201,
  "response_headers": {"Location": "/v1/payments/pay_9K2"},
  "response_body": {"id": "pay_9K2", "amount": 2000, "status": "succeeded"},
  "created_at": "2026-08-12T09:14:02Z",
  "expires_at": "2026-08-13T09:14:02Z"
}

go deeper

for a junior

Describe the two paths — first request executes and saves the response, repeat returns the saved response — and name the fields stored.

for a middle

Add the state column, insert-first-under-a-unique-constraint instead of select-then-act, and which response classes should and should not be frozen.

for a senior

Lead with atomicity between the effect and the dedupe record, discuss store placement (primary DB vs Redis) and its failure modes, and cover lease/recovery for crashed in-progress rows.

for a principal

Treat it as a platform capability: uniform semantics and response replay across services, storage growth and retention cost, passing keys through to downstream providers, and how it composes with outbox/event publication.

## The data model A workable dedupe record looks like this: - **scope** — who the key belongs to: API key, tenant, or account id. Never a global namespace (covered below). - **idempotency_key** — the client-supplied value. Together with scope this is the primary key or unique index. - **request_fingerprint** — a hash over the method, request target and body, used to detect a key reused with a different payload. - **state** — `in_progress` or `completed`. This is what makes concurrent retries safe. - **response_status, response_headers, response_body** — the frozen outcome to replay. - **resource_id** (optional) — a direct pointer to the entity created, useful for reconciliation. - **created_at, expires_at** — retention control. - **locked_at / lease_expires_at** (optional) — so a crashed in-progress row does not poison the key forever. ## The flow **Step 1 — claim the key by writing, not by reading.** The instinct is `SELECT ... ; if absent then execute`. That is a check-then-act race: two concurrent retries both see nothing and both execute. Instead, attempt `INSERT` of an `in_progress` row and let the unique index arbitrate. Exactly one insert succeeds; the other gets a unique-violation and knows a sibling request owns the key. **Step 2 — validate the fingerprint.** If a row exists and the fingerprint differs, the client has reused a key for a different payload. Reject (commonly 422; some APIs use 400 or 409) and execute nothing. **Step 3 — execute the operation.** **Step 4 — persist the response and mark completed.** If the dedupe store lives in the same database as the business data, do steps 3 and 4 in one transaction; then either both happened or neither did, and a retry after a crash is genuinely a first attempt. This is the single most valuable implementation decision in the whole pattern, and it is the reason people put the dedupe table in the primary database rather than in a separate cache. **Step 5 — on replay, return the stored response verbatim.** Same status, same body, same `Location`. Clients parse these; changing 201 to 200 on replay breaks callers that branch on the status. Signal the replay in a separate response header — `Idempotent-Replayed: true` is the common convention — so callers and your own logs can distinguish it. ## Which responses to freeze - **2xx** — freeze. That is the point. - **4xx caused by the request** (validation errors) — usually freeze, since the same request will fail identically anyway and freezing prevents a bad retry loop from doing work. Some APIs choose not to, so that a client fixing its payload… but note that a fixed payload changes the fingerprint and would be rejected as a mismatch anyway; the client should use a new key. - **5xx and timeouts** — normally do **not** freeze. If you store "500" against the key, the client's retry gets 500 forever and the operation can never succeed. Instead leave the key in a retryable state (or delete the row) so a retry re-executes. This is why the state column matters: a crashed request must leave behind either nothing or an expired lease, not a permanent failure. ## Response-size and storage realities Storing full response bodies is fine for small JSON but not for large payloads. Common mitigations: cap the stored body size, store only what is needed to reconstruct (status + resource id, re-rendering the representation on replay), or compress. Whatever you choose, replays must remain faithful — a replay that returns a *re-read* of the resource can differ from the original response if the resource changed afterwards, which is a subtle and real behavioural difference to decide deliberately. ## Where to put the store **Same relational database as the business data** — strongly preferred when the effect is local, because you get atomicity between the effect and the dedupe record for free, plus durability and a unique index for the race. Cost: extra write load on the primary and a table that needs pruning. **Redis or another cache** — fast and easy TTLs, but it is a second system with no transactional relationship to your effect. A crash between "charge succeeded" and "key marked completed" is unrecoverable; an eviction or failover losing keys silently permits duplicate execution. Acceptable when the effect is cheap to repeat, dangerous for money. ## Idempotency for effects you don't own If the operation calls a downstream provider (a PSP), the cleanest design is to **pass your key through** as the downstream idempotency key. Then even if your own record is lost, the provider deduplicates, and the two systems agree on what "the same operation" means. ## Pitfalls worth naming in an interview - SELECT-then-INSERT instead of INSERT-and-catch-violation (the race). - Marking `completed` before the effect commits (a crash then replays a success that never happened). - Returning a different status on replay than on the original. - Freezing 5xx responses and permanently bricking the key. - No TTL, so the table grows without bound. - Storing the key without a scope, so two tenants can collide.

  • Should the server store a 500 response against the idempotency key?
    Normally no. Freezing a 500 means every retry with that key replays the failure and the operation can never succeed, so the client would have to invent a new key — defeating the point. Leave the key retryable (delete the row or expire its lease) so a retry genuinely re-attempts.
  • Why insert the dedupe row before executing rather than checking whether the key exists first?
    A SELECT-then-execute sequence is check-then-act: two concurrent retries can both find nothing and both execute. Inserting first under a unique constraint makes the database arbitrate — exactly one insert wins, and the loser learns from the constraint violation that a sibling request owns the key.
  • What is the advantage of putting the dedupe table in the same database as the business data?
    You can commit the business effect and the dedupe record in one transaction, so a crash leaves either both or neither. With a separate store such as Redis you can charge a card and then fail to record the key, and the next retry charges again — the exact failure the pattern exists to prevent.

saying these in an interview costs you the question

  • Checking for the key with a SELECT and then executing (a check-then-act race).
  • Marking the key completed before the business effect has committed.
  • Returning a different status code on replay than the original request produced.
  • Storing 5xx responses against the key, permanently blocking that operation.
  • Keeping keys forever with no TTL, or storing them without a tenant/account scope.

context