A service intermittently fails with the database's deadlock error, which rolls the transaction back entirely. Design the application-side handling for that error: where the retry lives, how many attempts, what must not be inside the retried block, and how you distinguish it from errors that must not be retried.
answer
- retry at the transaction boundary, re-read
- only deadlock/serialization codes are retryable
- 3-5 attempts, exponential + jitter
- no emails/payments/cache writes inside
- outbox for messages; metrics on rate
basics
~20 sWrap the whole transaction in a bounded retry — about three to five attempts with exponential backoff and jitter — that re-reads and re-decides each time. Catch only the deadlock and serialization error classes; never retry constraint or business errors. Keep external side effects outside the retried block, and emit metrics.
solid answer
~60 sPut the retry **at the transaction boundary**, wrapping the entire unit of work, so each attempt opens a new transaction and re-reads from current data. Retrying just the failed statement is wrong: the rollback discarded everything, and values read earlier are stale. Rules that matter: - **Classify errors precisely.** Retry only the engine's deadlock and serialization-failure codes. Unique-constraint violations, foreign-key errors, permission errors, and domain failures are deterministic — retrying them just multiplies load and hides bugs. - **Bound it.** Three to five attempts, then fail the request with a clear error. Unbounded retries under contention turn a deadlock into livelock and a thread-pool outage. - **Backoff with jitter.** Without randomisation the same pair of transactions re-collides on the same schedule. - **Keep side effects out.** Emails, payment calls, message publishes, and in-memory mutations must live after commit — the block can run several times. Use an outbox for messaging. - **Make the work re-runnable.** No dependence on state accumulated by the failed attempt. - **Instrument.** Count deadlocks and attempts per operation; a rising rate is a design defect, not a retry-budget problem.
code
text · 17 linesattempt = 0
loop:
begin transaction
try:
re-read inputs <-- must be inside the loop
recompute decision
write
commit
break
catch e:
rollback
if not isDeadlockOrSerializationFailure(e): throw e
attempt += 1
if attempt >= MAX (3-5): throw e
sleep(base * 2^attempt * random(0.5, 1.5))
// emails, payments, message publishes: AFTER the loop, or via outboxgo deeper
Know that a deadlock error means the transaction was rolled back and should be re-run, and that the retry wraps the whole transaction, not one statement.
Specify the mechanics: precise error classification, a bounded attempt count, exponential backoff with jitter, and re-reading inside each attempt.
Add the correctness and blast-radius reasoning — side effects outside the block or via an outbox, latency budget versus caller timeouts, connection-pool pressure, coordination with upstream retries, and metrics that distinguish tolerance from a fix.
Make it a platform concern: one shared retry abstraction with consistent semantics, idempotency conventions for transaction boundaries, deadlock-rate SLOs with alerting, and a policy for when contention gets redesigned rather than retried.
## Why retry is the correct response — and only half the answer A deadlock error carries a specific contract: your transaction was chosen to break a wait cycle, it was rolled back **completely**, and the conditions that caused it were timing-dependent, so running again is likely to succeed. That makes retry correct. But retry is a *tolerance* mechanism, not a fix; a service whose deadlock rate is climbing needs its access ordering or transaction scope repaired, and the retry layer exists to keep the occasional collision from reaching a user. ## Where the retry lives At the **transaction boundary**, wrapping the whole unit of work: begin, read, decide, write, commit. This is non-negotiable, and it is where most implementations go wrong. Retrying an individual statement inside a transaction that no longer exists is meaningless. Retrying at the boundary but reusing values loaded before the failure is worse — the surviving transaction has since changed exactly the rows you were fighting over, so replaying a decision computed from stale reads can produce a wrong result rather than an error. Each attempt must re-read. Structurally, this argues for one shared helper — a `withTransactionRetry { … }` wrapper, an interceptor around service methods, or the framework's declarative equivalent — rather than per-call-site try/catch. A single implementation means the classification, the bounds, the backoff, and the metrics are consistent, and new code inherits them. ## Classifying the error Retry must be driven by the engine's **specific error class**, not by string matching on a message and not by catching a broad exception type. Retryable: deadlock detected, serialization failure, and — with more care — lock-wait timeout, which is retryable but often signals a slow holder rather than a cycle. Not retryable: unique or foreign-key constraint violations, check-constraint failures, syntax or type errors, permission denials, and any domain-level rejection. These are deterministic; a second attempt produces the same failure while doubling the load and burying the real bug. A subtlety worth raising: some frameworks wrap driver exceptions, and a careless mapping can collapse a deadlock and a constraint violation into the same generic type. Verify what your data-access layer actually throws, and assert it in a test. ## Bounding and spacing the attempts **Bound the attempts** — three to five is typical. Beyond that, either the contention is structural or the workload is saturated, and continuing to retry converts a localised problem into a system-wide one: connections stay checked out, thread pools fill, upstream timeouts fire, and load amplification pushes the deadlock rate higher still. When the cap is reached, fail the request with a clear error the caller can act on. **Back off exponentially with jitter.** Two transactions that deadlocked will, if retried immediately and in lockstep, collide again in exactly the same way — the classic path from deadlock to livelock, where work is repeatedly done and thrown away and nothing progresses. Randomised delay decorrelates the retries so one attempt wins. ## What must not be inside the retried block Anything the database rollback cannot undo: - **External calls** — payment authorisations, emails, SMS, webhooks, third-party APIs. If the transaction is retried, these run again; a duplicate charge is far worse than a failed request. - **Message publishes** — use the transactional outbox pattern: write the intent to a table inside the transaction, and let a separate relay publish it after commit. - **In-memory or cache mutations** — updating a shared cache, a counter, or an object's state inside the block leaves stale mutations behind when the attempt is rolled back. - **Non-idempotent business logic** — anything that consumes a nonce, generates a sequence value with a visible side effect, or writes files. The general rule: the block must be safe to execute *n* times with only the last one committing. If some step cannot be made re-runnable, move it after commit, or make it idempotent with a natural key or an idempotency token. ## Interaction with the rest of the stack Retries consume the request's latency budget, so the total worst case — attempts times per-attempt latency plus backoff — must fit inside the caller's timeout, or the caller gives up while the database keeps working. Retries also hold a connection for the whole sequence, so a burst of contention can exhaust the pool; some designs return the connection between attempts precisely to avoid that. And retries interact with upstream retries: a client that retries on a 500 while the service retries internally multiplies attempts, so the layers must be coordinated rather than each defending itself. ## Observability Emit a counter per operation for deadlock errors and a histogram for attempts-until-success, and alert on rate rather than on individual events. Log the engine's deadlock report or its identifiers, because that is what names the conflicting statements and their access order. Two patterns should trigger investigation rather than tuning: a specific pair of operations that always appear together in reports (an ordering bug), and one operation that is always the victim (a starvation pattern, usually a small transaction colliding with a large batch). ## Then fix the cause With retries in place, use the reports to remove the deadlock: order the access to shared rows consistently across all code paths, sort keys before batch updates, shorten transactions so locks are held briefly, avoid user think-time inside a transaction, and take the strongest lock you will need up front instead of upgrading a shared lock to exclusive mid-transaction.
- Why must the reads be inside the retry loop rather than fetched once before it?Because the rollback discarded the transaction and the surviving transaction has since modified exactly the rows that caused the conflict. Replaying a decision computed from pre-failure reads writes a result derived from stale data, which can be silently wrong rather than merely failing. Re-reading inside each attempt is what makes the retry semantically equivalent to running the operation later.
- How do you keep a message publish inside the same logical unit of work if it cannot go inside the retried transaction?Use a transactional outbox: insert the message into an outbox table as part of the same transaction, so it is rolled back with everything else if the attempt fails, and have a separate relay process read committed outbox rows and publish them after the fact. Consumers should be idempotent because the relay guarantees at-least-once delivery.
- What goes wrong with an unbounded retry loop under heavy contention?It converts a transient deadlock into livelock and load amplification. Each attempt holds a connection, does work, and throws it away, so the connection pool drains, thread pools fill, latency climbs past upstream timeouts, and the extra concurrency raises the deadlock rate further. Bounding attempts and adding jittered backoff caps the blast radius and surfaces the underlying design problem instead of hiding it.
- Should a lock-wait timeout error be retried the same way as a deadlock?It can be retried, but it means something different and deserves separate treatment. A deadlock is a cycle that the engine resolved immediately; a lock-wait timeout means one holder kept a lock longer than the threshold, so retrying will likely hit the same slow holder. Track the two separately, retry timeouts more conservatively, and treat a rising timeout rate as a signal to shorten the holding transaction rather than to raise the limit.
saying these in an interview costs you the question
- Catching a broad exception type and retrying everything, including constraint violations
- Retrying only the failed statement, or reusing data read before the rollback
- Unbounded retries, or retries with no backoff or no jitter
- Leaving emails, payment calls, or cache writes inside the retried block
- Treating a rising deadlock rate as something to absorb with a bigger retry budget instead of fixing the access order