Code running at the SERIALIZABLE isolation level must expect transactions the database aborts with a serialization failure (SQLSTATE 40001). Describe the retry pattern you would implement around such transactions, and the mistakes that make a retry loop wrong.
answer
- 40001 is normal, not a bug
- retry the whole transaction, reads included
- capped attempts + exponential backoff + jitter
- no emails/payments/publishes inside the retried block
- retry at the outermost boundary; metric the abort rate
basics
~20 sWrap the whole transaction in a loop: on a serialization failure, roll back and re-run everything including the reads, with a capped number of attempts and jittered backoff. Keep external side effects out of the retried block, and retry only serialization and deadlock errors.
solid answer
~60 sAt SERIALIZABLE the engine enforces correctness by aborting transactions it cannot fit into a serial order, so **a serialization failure is a normal outcome, not an error condition**. Every write path needs a retry wrapper. The pattern: a function that takes the transaction body as a closure, opens a transaction, runs it, commits; on SQLSTATE 40001 (and deadlock errors) it rolls back and re-runs the whole body, re-executing the reads because values from the aborted attempt are unusable. Cap attempts — typically 3 to 5 — with exponential backoff plus jitter, then surface a clean try-again error to the caller. The classic mistakes: retrying a single statement instead of the transaction; reusing values read in the failed attempt; retrying non-retryable errors such as constraint violations; unbounded retries that amplify contention into a storm; and performing non-idempotent side effects — sending mail, charging a card, publishing to a broker — inside the retried block. Side effects belong after commit, or in an outbox row written transactionally. Track the abort rate as a metric; a rising rate is a contention signal.
code
text · 12 linesfor attempt in 1..MAX:
try:
begin(isolation = SERIALIZABLE)
result = body() # reads AND writes live here
commit()
return result # side effects happen AFTER this line
catch e where isSerializationFailure(e) or isDeadlock(e):
rollback()
sleep(backoff(attempt) + jitter())
catch e:
rollback(); throw e # constraint/syntax/auth: never retry
throw RetriesExhaustedgo deeper
Know that this level can abort your transaction and that the fix is to run the whole transaction again.
Describe the wrapper concretely: capped attempts, backoff with jitter, reads inside the retried block, retryable-error classification.
Add operational depth — outbox for side effects, retry at the outermost boundary, connection hygiene after rollback, abort-rate metrics, and index and transaction-length work to cut conflicts.
Set the platform policy: where retry lives in the stack, what abort rate is acceptable, and when to convert hot paths to explicit locking or constraints instead.
## Why the failure is expected SERIALIZABLE guarantees that the committed outcome matches some serial execution. An engine can only deliver that by refusing schedules that admit no such order — either by making transactions wait (locking implementations, where the residual failure is deadlock) or by aborting them when tracked read-write dependencies look dangerous (optimistic implementations, where the failure arrives around commit). Both hand the application the same contract: **your transaction may be told that did not work, run it again.** A transaction can be aborted through no fault of its own. In optimistic implementations even a read-only transaction can be the victim, and detection is conservative: the engine aborts on patterns that *might* be non-serializable, so some aborts are false positives. None of that is negotiable from the application side; the only correct response is to re-run. ## The pattern The transaction body must be a **re-runnable unit**. In practice that means one function, executed inside a wrapper: 1. Open the transaction at SERIALIZABLE. 2. Execute the body: reads, computation, writes. 3. COMMIT. On success, return. 4. On a retryable error — serialization failure (SQLSTATE 40001) or deadlock — ROLLBACK, sleep for a backoff interval with jitter, and repeat from step 1. 5. After N attempts, give up and return a distinct, retriable-looking error to the caller. Key properties: - **The reads are inside the retried block.** This is the whole point. The previous attempt's snapshot has been declared unusable; recomputing from those values would reproduce exactly the non-serializable outcome the engine rejected. - **Bounded attempts.** Unbounded loops turn contention into a self-sustaining storm: every retry consumes CPU, locks and connections, which raises the conflict rate, which causes more retries. - **Backoff with jitter.** Without jitter, conflicting transactions retry in lockstep and collide again. - **Precise error classification.** Retry class-40 errors. Do not retry unique-constraint violations, check-constraint failures, syntax errors, permission errors or application exceptions — those will fail identically forever. - **Connection state.** After the abort, the connection must be rolled back before reuse; a pooled connection returned in an aborted state poisons the next borrower. ## Side effects are the real trap Anything the retry cannot undo must not live inside the retried block: - Sending an email or SMS, publishing to a message broker, calling a payment API: on retry it happens twice. - Mutating in-memory caches or counters: the retried attempt sees corrupted state. - Consuming a non-repeatable resource, such as popping from an in-process queue. The standard remedies are to perform external effects only after a successful commit, or to write an **outbox** row inside the transaction and have a separate process deliver it — the outbox row is rolled back with the aborted attempt, so no duplicate is ever emitted. ## Where the retry belongs in the stack At the **outermost** transaction boundary. Nested retries — an inner service retrying inside an outer transaction that then also retries — multiply attempts geometrically and cannot work anyway, because an aborted transaction cannot host further statements. In frameworks with declarative transactions, the retry advice must sit outside the transaction advice, and any HTTP or RPC layer above it should not add its own blind retry on top. ## Making retries rarer Retry logic is the safety net, not the design: - **Keep transactions short.** Conflict probability grows with the window; move computation, remote calls and user think-time out of the transaction. - **Narrow the read set.** Optimistic implementations track what you read; a full scan reads everything and conflicts with everyone, while an index lookup reads little. Missing indexes are a common cause of high abort rates. - **Order access consistently** to cut deadlocks in locking implementations. - **Push invariants into constraints** where possible — a unique index enforces uniqueness with no serializable schedule required. - **Consider explicit locking** for a few notoriously hot paths, converting an abort-and-redo into a short, predictable wait. ## Operability Emit metrics: attempts per transaction, retry rate by transaction type, and exhausted-retry count. A retry rate creeping from 0.1% to 5% is an early warning of a hot key, a lengthened transaction or a lost index — usually visible well before latency alarms fire.
- Why can't you just re-issue the failing statement instead of the whole transaction?Once a serialization failure is raised the transaction is aborted; no further statement can execute in it until you roll back. Beyond the mechanics, the engine's verdict is that the set of things this transaction read and wrote has no valid serial position, so only redoing the reads can produce a correct result. Retrying a single statement would either fail again or commit a result the engine already rejected.
- How do you keep a retried transaction from sending duplicate emails or duplicate payments?Keep external effects out of the transaction body. Either perform them only after a successful commit, or write an outbox row inside the transaction and let a separate dispatcher send it; an aborted attempt rolls the outbox row back, so nothing is emitted. If an external call truly must happen inside, it needs its own idempotency key so repeats are collapsed downstream.
saying these in an interview costs you the question
- Treating SQLSTATE 40001 as an unexpected error to alert on rather than a normal, retryable outcome.
- Retrying the statement rather than the transaction, or reusing values read in the aborted attempt.
- Unbounded or un-jittered retry loops that amplify contention.
- Retrying non-retryable errors such as unique-constraint violations, which can never succeed.
- Leaving emails, broker publishes or payment calls inside the retried transaction body.