skip to content

When implementing the Idempotency-Key HTTP header, what namespace should a key be scoped to, how long should stored keys be retained, and what happens to a client that retries after that retention window expires?

level: seniorimportance: should knowfreq 38%

answer

  1. Scope = (principal, key); global namespace leaks across tenants
  2. Endpoint in scope vs in fingerprint = a deliberate choice
  3. ~24h TTL: > client retry window, < storage pain
  4. Prune by partition drop or batched delete on expires_at
  5. After expiry a retry re-executes — reconcile, don't retry

basics

~20 s

Scope keys per authenticated principal (API key, account or tenant) — never globally, or one tenant's key can collide with another's and replay their response. Retain typically 24 hours. After expiry the key is unknown, so a retry re-executes and duplicates the effect; publish the window and require clients to stop retrying before it.

solid answer

~60 s

**Scope.** The lookup key is `(principal, idempotency_key)`, where the principal is the API key, account or tenant. A global namespace is a correctness *and* security defect: two tenants choosing the same value would collide, and the second would be served the first's stored response — a cross-tenant data leak. Whether the endpoint is also part of the scope is a design choice: including it isolates operations, excluding it means the same key on a different endpoint is a fingerprint mismatch. **TTL.** 24 hours is the common convention. It must comfortably exceed the client's total retry window, and it bounds a table that would otherwise grow with every write request. Enforce it with an `expires_at` column plus a pruning job or partition drop — a `DELETE ... WHERE expires_at < now()` over a hot table needs batching and an index. **After expiry** the key looks brand new, so a late retry **re-executes** — a duplicate charge. That is why the retention window is a published contract: clients must exhaust retries well inside it, and beyond it must reconcile by querying rather than retrying.

go deeper

for a junior

Say keys are scoped per account or API key, kept for about a day, and that a retry after expiry would run the operation again.

for a middle

Explain why a global namespace collides and leaks, how the TTL relates to the client's retry window, and that expires_at needs an actual pruning job.

for a senior

Cover the endpoint-in-scope-versus-in-fingerprint choice, partition-drop pruning versus batched deletes, response-body storage cost, and how clients should reconcile rather than retry past the window.

for a principal

Position the header as a TTL-bounded transport dedupe and argue for a permanent domain-level unique constraint alongside it, plus the published retention contract clients build their retry and recovery policy against.

## Scope: what namespace does a key live in? The stored record must be looked up by more than the key string. **Always include the authenticated principal** — API key, account id, or tenant id. Two reasons: 1. *Correctness.* Keys are client-chosen. Nothing stops two customers picking the same value, especially if either uses something weaker than a UUID (a sequence number, an order id, a timestamp). 2. *Security.* If scopes collide, tenant B sends its key and receives **tenant A's stored response body** — someone else's payment id, amounts, and metadata. That is a cross-tenant data leak triggered by an ordinary request, and it is the strongest single argument for scoping. It also means a key presented by a different principal is **not** a payload mismatch. It is a different namespace, and the server must handle it as an unknown key without hinting that the value exists elsewhere. **Whether to include the endpoint** in the scope is a genuine choice: - *Include it* (`(principal, endpoint, key)`): the same key used on `POST /payments` and `POST /refunds` names two independent operations. Convenient for clients that derive keys from a business id, but it removes a safety net — a client that meant to retry a payment and accidentally hit the refund endpoint gets an execution rather than a rejection. - *Exclude it* (`(principal, key)`) and put the method and path into the **fingerprint**: a key reused across endpoints trips the mismatch check and is rejected. Stricter, and usually what you want, because a key names one intent and one intent lives at one endpoint. The IETF draft leaves scope server-defined, so whichever you pick must be documented explicitly. ## TTL: how long to keep keys **24 hours is the de-facto standard** (Stripe's published retention). The number is not sacred; the constraints are. *Lower bound.* The window must exceed the client's entire retry campaign — including a client that queues an operation, crashes, restarts, and drains its queue later. If clients retry with exponential backoff for an hour, a 15-minute TTL guarantees the last retries land in an expired namespace and double-execute. Some APIs offer 7 days or more precisely because their clients are batch systems with long recovery cycles. *Upper bound.* Storage and write amplification. Every mutating request creates a row containing a response body. At high volume this table can rival the business tables in size, and it is pure overhead. Longer retention also means longer-lived keys that could collide within a scope. *Mechanics.* Store `expires_at` and prune. Options: a batched delete job (`DELETE ... WHERE expires_at < now() LIMIT n` in a loop, with an index on `expires_at`), time-based **partitioning** with a partition drop (cheapest at scale — no row-by-row deletes, no vacuum pressure), or a native TTL if the store provides one (Redis, DynamoDB). Naive unbatched deletes on a hot table are a classic self-inflicted incident. *Body size.* Cap what you store. A response body of a few kilobytes times millions of requests times 24 hours is real money. Truncating or storing only status plus resource id is a legitimate optimization, but then a replay may not be byte-identical to the original — decide that deliberately and document it. ## What happens after expiry The uncomfortable answer: **the key is simply unknown, so the operation runs again**. Idempotency is not a permanent property of the key; it holds only within the retention window. A client that retries a 26-hour-old payment against a 24-hour TTL creates a second charge, and the server has no way to know. The mitigations are contractual and operational, not technical: - **Publish the window** in the API documentation, in the same breath as the header itself. - **Tell clients to bound their retries** well inside it. - **Give clients a reconciliation path**: a way to query by their own reference (`GET /payments?client_reference=…`) so that after the window, the correct behavior is *look it up*, not *retry it*. This is the durable answer — a business-level natural key that outlives the dedupe record. - **Consider a longer retention for high-value operations** even if the general TTL is shorter, or persist a permanent business-level unique constraint (e.g. a unique index on the client reference for payments) that survives key expiry. That last point is worth saying out loud in an interview: the idempotency key is a *transport-level* dedupe with a TTL; a unique constraint on a business identifier is a *domain-level* guarantee with no expiry. The strongest systems have both, and the second is what actually saves you at hour 26. ## Summary Scope by principal, always; by endpoint, deliberately. Retain around 24 hours — long enough to cover every retry your clients perform, short enough to keep the table prunable, enforced by partition drops rather than ad-hoc deletes. And be explicit that the guarantee ends when the record does, which is why long-lived correctness belongs to a business-level unique constraint rather than to the header.

  • What concretely goes wrong if keys share one global namespace across all customers?
    Two customers can pick the same key value, and the second one receives the first one's stored response — including another tenant's resource ids and amounts. It is simultaneously a correctness bug (their operation silently never runs) and a cross-tenant data leak triggered by an ordinary request.
  • How would you prune the dedupe table at high write volume?
    Prefer time-based partitioning and drop whole partitions once they age out — no row deletes, no index churn, no vacuum pressure. If partitioning is not available, run a batched delete loop against an index on expires_at with a bounded LIMIT per statement, so you never take a long-running delete on a table that is also the hot write path.
  • How can a client behave correctly for an operation whose idempotency key has already expired?
    It should stop retrying and reconcile instead: query the API by its own business reference to see whether the operation exists. That is why the server should expose a lookup on a client-supplied reference and, for high-value operations, back it with a permanent unique constraint — a domain-level guarantee that outlives the TTL-bounded dedupe record.

saying these in an interview costs you the question

  • Storing keys in a single global namespace with no tenant scope.
  • Setting a TTL shorter than the client's own retry window.
  • Claiming an idempotency key protects against duplicates forever.
  • Keeping records indefinitely with no pruning strategy.
  • Treating a key seen under a different account as a payload mismatch instead of an unknown key.

context