Your POST endpoint honours the Idempotency-Key header, but the operation it performs is a call to an external payment provider you do not control. How do you keep the stored idempotency record and the external side effect consistent across crashes?
answer
- No 2PC with a third party — converge, don't atomically commit
- Commit intent (in_progress + lease) BEFORE calling out
- Pass your key through as the provider's idempotency key
- Never mark completed before the effect is confirmed
- Outbox for events; terminal state + reconciliation queue for stuck rows
basics
~20 sYou cannot make a local commit and a remote call atomic, so you record intent first, then call, then record the outcome — and you pass your key through as the provider's idempotency key. A crash mid-flight is then recoverable: re-calling the provider with the same key deduplicates at their end.
solid answer
~60 sThere is no two-phase commit with a third-party API, so the design is **intent-first plus downstream idempotency**. 1. Commit an `in_progress` record for `(scope, key)` with a lease, before calling out. This durably records that the operation was attempted. 2. Call the provider, passing **your** idempotency key as **their** idempotency key. 3. Commit the result into the record and mark it `completed`. A crash between 2 and 3 leaves an in-progress record with a lease. When it expires, a retry re-calls the provider with the same key; the provider deduplicates and returns the *original* charge, so recovery converges rather than double-charging. If the provider supports lookup by key, prefer querying before re-calling. What you must never do is mark the record completed before the effect is confirmed (replays would claim a success that never happened) or rely on a non-durable store such as an unreplicated cache for the record — a lost key means a second real charge. For emitted events, use an outbox written in the same transaction as the completion.
code
http · 7 linesPOST /v1/charges HTTP/1.1
Host: psp.example.com
Idempotency-Key: 8f14e45f-ceea-467a-9f5a-3c3f2f0d1b77
Authorization: Bearer <psp-token>
Content-Type: application/json
{"amount":2000,"currency":"usd","source":"card_1"}go deeper
Recognize that a local database commit and a remote API call cannot be made atomic, and that recording the attempt before calling out is what makes recovery possible.
Describe the intent-first ordering and passing your key through as the provider's idempotency key so a retry cannot double-charge.
Add lease-based takeover, preferring a provider lookup over a blind re-call, why the dedupe record must be durable and transactional rather than cached, and the outbox for emitted events.
Frame the achievable guarantee — at-most-once effect via convergence, not atomicity — make the provider's semantics and retention an explicit dependency of your correctness argument, and design the terminal states, reconciliation path and alerting for what automation cannot close.
## The core impossibility An idempotency-key implementation whose effect is local is easy: commit the business row and the completed dedupe record in one database transaction, and every crash leaves you in a consistent state. The hard version is when the effect lives in someone else's system — a payment provider, a shipping carrier, an email sender. You cannot enlist their API in your transaction, and they will not participate in two-phase commit. So there is always a window where your local record and their state can disagree. The design goal shifts from "never disagree" to **"always be able to converge"**. ## Intent-first ordering The ordering that survives crashes is: 1. **Commit intent.** Write and commit the `(scope, key)` record as `in_progress`, with a request fingerprint and a lease expiry, *before* any outbound call. This is durable evidence that an attempt is under way, which is what recovery will key off. 2. **Perform the effect**, passing your key downstream. 3. **Commit the outcome.** Write the provider's result into the record and mark `completed`. Compare the alternatives and why they fail: - *Call first, record after.* A crash after the charge leaves no local trace at all. The retry sees an unknown key and charges again — unless the provider deduplicates, which is exactly why step 2's key pass-through carries so much weight. - *Mark completed before confirming.* A crash or a provider failure after the mark means every replay confidently reports success for a payment that never happened. This is the worst failure mode because it is invisible: no error, no alert, just a customer who was never charged and an order that shipped. ## Key pass-through: making recovery converge The single most important technique is to send **your** idempotency key as the provider's `Idempotency-Key`. It has the effect of extending your dedupe boundary into their system: any number of retries or takeovers of the same logical operation collapse to one charge at their end, because both systems now agree on the operation's identity. Derive their key deterministically from yours (the key itself, or a hash of key + operation name if you make several distinct downstream calls in one request — each distinct call needs a stable, distinct key). Never generate a fresh downstream key per attempt; that is the same bug as regenerating your own key per retry, just one layer down. ## Recovery from an abandoned in-progress record When a later retry finds an in-progress record with an expired lease, it may take over. Take over **atomically** — a conditional update that re-stamps the lease so only one taker wins — and set the lease longer than the maximum plausible execution time, because taking over an operation that is still running is how you reintroduce the duplicate. Then prefer **query over re-execute**: if the provider exposes lookup by idempotency key or by your client reference, read their state and reconcile into your record. Re-calling with the same key is the fallback and is safe if their dedupe window covers your recovery window — note that their TTL is now part of your correctness argument, so it belongs in your design docs, not just theirs. Some operations simply cannot be resolved automatically (provider has no lookup, key window expired). For those, the record should transition to a state that a reconciliation job or a human queue picks up. Designing that terminal state explicitly, rather than leaving rows stuck in-progress forever, is what separates a system that degrades from one that silently loses money. ## Storage placement follows from all this The record must be **durable and transactional with your local state**. That argues for the primary relational database rather than a cache: an eviction, a failover that loses recent writes, or a non-replicated node means a lost key, and a lost key for a payment means a second real charge. Caches are fine when the effect is cheap to repeat; they are not fine here. ## Events and downstream fan-out If completing the operation also publishes an event (`PaymentCaptured`), do not publish inside the request path — a crash between commit and publish loses the event, and publishing before commit can announce work that rolls back. Write the event to an **outbox table in the same transaction as the completed record**, and let a relay publish it at-least-once. Consumers then need their own idempotency, keyed on the event id — the same pattern recursing one level out. ## What a principal-level answer emphasises - The guarantee you can actually offer is **at-most-once effect with at-least-once attempt and a convergence path**, not distributed atomicity. - The provider's idempotency semantics and retention window become part of *your* correctness argument. - Every state must be recoverable: no row may sit in-progress forever, and there must be a defined path — automatic reconciliation or an operational queue — for the cases automation cannot close. - Observability is part of the design: alert on in-progress records older than the lease, on takeover rates, and on reconciliation mismatches, because these are the signals that the invisible failure mode is happening.
- Why is marking the idempotency record completed before confirming the external call so dangerous?Because every subsequent replay reports a success that never happened, with no error anywhere in the system. The customer is not charged, downstream processes proceed as if they were, and nothing alerts — it is a silent, self-concealing failure, which is far worse than a duplicate you can detect and refund.
- What can you do when the provider offers no idempotency key and no lookup by your reference?You lose automatic convergence and must fall back to detection: record intent before calling, and on recovery leave the operation in an explicit unresolved state that a reconciliation process or a human queue handles, typically against the provider's settlement file or dashboard. You should also weigh not auto-retrying at all, since a blind retry against a non-deduplicating provider is a guaranteed duplicate whenever the ambiguity was case two.
- How does the provider's own key retention window affect your design?Their TTL becomes part of your correctness argument: a takeover or recovery that re-calls them after their window has closed will create a second effect. Your lease and reconciliation timings must therefore fit inside their retention, and that dependency should be documented and monitored rather than assumed.
- Where does the outbox pattern fit into this?When completing the operation must also publish an event, writing the event to an outbox table inside the same transaction as the completed record makes publication as durable as the completion. A relay then delivers at-least-once, and consumers deduplicate on the event id — the same idempotency problem recursing one level outward.
You cannot make posting a letter and writing your diary entry a single act. So you write 'about to post' first, put your reference number on the letter, and if you wake up unsure, you ask the post office about that reference rather than sending a second letter.
saying these in an interview costs you the question
- Claiming a distributed transaction or two-phase commit across your database and a third-party API.
- Calling the provider first and recording the key afterwards.
- Marking the record completed before the external effect is confirmed.
- Keeping the dedupe record only in a non-durable cache for a money-moving operation.
- Generating a fresh downstream idempotency key on each retry attempt.
- Leaving in-progress records with no terminal state or reconciliation path.