skip to content

A payment consumer stores every message ID it has successfully processed in a 'processed_messages' database table, and checks that table before applying a payment. What race condition can still cause a duplicate payment if two copies of the same message arrive at nearly the same time, and how would you close it?

level: middleimportance: must knowfreq 80%

answer

  1. check-then-act race
  2. unique constraint on message_id
  3. INSERT ... ON CONFLICT DO NOTHING
  4. claim before vs after side effect
  5. partial-claim / crash between insert and payment

basics

~20 s

A table that lists 'already handled' message IDs can still let two duplicates through if both are checked at the same instant before either writes its ID down, like two people checking an empty sign-up sheet at once and both writing their name first. Fix it by making the 'write' step atomic, for example a database unique constraint.

solid answer

~50 s

Checking a processed_messages table and then inserting into it as two separate steps is a classic check-then-act race: if two copies of the same message are processed concurrently (two consumer threads, two instances, or a redelivery racing the original), both can read 'not present' before either writes, and both proceed to charge the customer. The fix is to make the check-and-mark atomic with the side effect: put a unique constraint on the message ID column and perform the payment and the insert in the same database transaction, letting the second insert fail with a constraint violation that the consumer catches and treats as 'already handled, skip.' An INSERT ... ON CONFLICT DO NOTHING (or equivalent), combined with checking rows-affected, is the idiomatic pattern; it turns the race into a database-enforced serialization point instead of an application-level race.

go deeper

for a junior

Should recognize that 'check first, then act' can race if two things happen close together, even without naming the formal term.

for a middle

Should name the check-then-act race explicitly and propose a database uniqueness constraint as the fix.

for a senior

Should design the atomic claim-and-mark pattern (unique constraint + ON CONFLICT or equivalent) and discuss ordering it relative to the side effect, including the partial-claim failure mode.

for a principal

Should discuss this in the context of distributed transactions/outbox patterns when the side effect isn't in the same database as the dedup table, and weigh claim-before vs claim-after trade-offs system-wide.

## The sequence that looks correct on paper The naive implementation of a processed-message table looks correct on paper: `SELECT 1 FROM processed_messages WHERE message_id = ?`; if not found, do the payment, then `INSERT INTO processed_messages (message_id) VALUES (?)`. Read individually each statement is fine, but the sequence is a textbook **check-then-act race condition**. If two copies of the same logical message are being handled at overlapping times — say - the broker redelivered a message because an ack was lost in transit, and the original processing hadn't finished yet, or - two consumer instances in the same consumer group both somehow received a copy due to a rebalance glitch — both executions can run the SELECT before either has run the INSERT. Both see 'not found,' both conclude they're the first to handle it, both charge the card, and only afterward do both INSERTs race to write the same `message_id`, with one succeeding and the other, if unprotected, either erroring out too late to prevent the double charge or silently overwriting. ## Why the naive version survives testing **Why the naive version exists:** it's the intuitive translation of 'remember what you did' into code, and it works fine in testing where messages are processed serially by a single thread with no true concurrency. It only breaks under real concurrent or overlapping redelivery, which is exactly the condition integration tests rarely reproduce; this is a common reason dedup bugs escape to production despite passing all unit and integration tests. ## The fix: one atomic claim The fix is to collapse the check and the mark into a single atomic operation, ideally in the same transaction as the side effect itself. Concretely: 1. Put a `UNIQUE` (or `PRIMARY KEY`) constraint on the `message_id` column in `processed_messages`. 2. Then in one database transaction attempt to insert the `message_id` and perform the payment write together. 3. If the INSERT violates the unique constraint, the database itself tells you someone already claimed this `message_id`, and you catch that specific error and treat it as a no-op skip-and-ack, never reaching the payment step at all. This converts a race that used to depend on timing into a **serialization point enforced by the database's own concurrency control** (row-level locking / unique index), which is far more reliable than an application-level if-check because the database guarantees only one of two concurrent inserts for the same key can succeed. An equivalent idiom in Postgres is `INSERT INTO processed_messages (message_id) VALUES ($1) ON CONFLICT (message_id) DO NOTHING`, then checking the row count returned: | Row count | What it tells you | |---|---| | zero rows affected | already processed, skip | | one row affected | you own this message, proceed | ## Claim before, or claim after **Trade-offs:** - **Claim first.** Doing the insert-and-check first, before the business side effect, means you claim the message before you know the business logic will succeed; if the payment step then fails for an unrelated reason (card declined, downstream timeout) and the message needs to be retried, you must either delete your claim row on failure or design retries to bypass the dedup check for legitimate re-attempts, which reintroduces complexity. - **The alternative**, do the payment first, then the atomic insert-and-check as the final step gated by a transaction that wraps both, avoids the false-claim problem but requires the payment write and the `processed_messages` write to be transactionally atomic together (same database, or a two-phase/outbox pattern if they're different systems), which is not always possible if the side effect is a call to an external payment gateway rather than a local database write. ## What goes wrong under load **Failure modes in production:** - **Teams that skip the unique constraint** and rely purely on select-then-insert in application code see duplicate charges appear specifically under load spikes or after incidents that trigger mass redelivery, exactly the conditions where true concurrency is highest and the race window is most likely to be hit. - **Another failure mode is partial claim:** the INSERT succeeds but the process crashes before the payment call completes, leaving the `message_id` marked as processed forever even though the payment never happened; this needs either transactional atomicity between the two writes or a status column (`PENDING/COMPLETED`) with a reconciliation job that revisits stuck rows. ## The same shape behind an idempotency key **A concrete real-world pattern:** this is essentially what a database-backed idempotency-key implementation, like Stripe's, does internally. The idempotency key has a unique constraint, the first request to insert it wins and proceeds to do the real work, and any concurrent or later request with the same key either blocks briefly on the row lock or gets redirected to return the already-computed result once the first request commits.

  • Would putting a mutex/lock inside the application process fix this race?
    Only if there's exactly one process and one thread ever handling messages, which defeats the purpose of horizontally scaling consumers. As soon as you run more than one consumer instance, an in-process lock can't see across instances, so the race reappears at the multi-instance level. The fix has to be enforced by something shared across all instances, the database's unique constraint, not process-local locking.
  • What should the consumer do if the INSERT fails due to the unique constraint?
    Treat it as this message was already claimed or processed and simply ack the message without redoing the side effect; it's the expected, successful outcome of a duplicate delivery, not an error to alert on. It should still be logged or counted as a metric so the team can see how often redelivery happens, but it shouldn't page anyone or retry the payment.
  • How would you handle the case where the payment succeeds but the process crashes before the INSERT into processed_messages commits?
    If the two writes aren't in the same transaction, that gap is unavoidable with this pattern alone, so you need either a transactional outbox that writes both in one local transaction, or a downstream idempotency key passed to the payment gateway itself, so even a full reprocessing from scratch calls the gateway with the same key and gets the original result back instead of a second charge.

Like two people simultaneously checking a paper sign-up sheet that has no seat assigned yet; both see it's empty and both write their name down, because looking at the sheet and writing on it weren't one atomic action. A numbered ticket dispenser fixes it by making 'take a number' a single indivisible act enforced by the machine.

saying these in an interview costs you the question

  • Doesn't recognize check-then-act as a race condition
  • Proposes only an application-level if-statement with no database constraint
  • Assumes single-threaded consumers eliminate the race entirely
  • No plan for what happens if the process crashes between the claim and the side effect
  • Treats a unique-constraint violation as an error condition to alert/retry rather than an expected skip

context