skip to content

In a data-access layer, where must a retry of a failed unit sit, and what must each attempt rebuild?

level: seniorimportance: must knowfreq 58%

answer

  1. retry the boundary, not the statement
  2. one attempt, one transaction
  3. reload inside every attempt
  4. re-derive the decision, not just the write
  5. ambiguous outcomes need idempotency

basics

~20 s

Outside the transaction boundary. Each attempt starts a new unit with an empty tracked set and reloads the data it works on, because objects from the failed attempt belong to a rolled-back transaction. Bound the attempts and keep the operation idempotent.

solid answer

~50 s

A retry placed inside the unit is useless: the transaction is already condemned, so the second attempt writes into something that will be rolled back regardless. The retry therefore wraps the **whole boundary**, and one attempt means a new transaction, a fresh tracked set, and a reload of every object the work reads - the copies from the previous attempt hold versions read in a transaction that no longer exists. The decision the work makes must be re-derived inside the attempt too; re-applying a decision computed from stale data is how a retry corrupts state rather than repairing it. Retry only the transient kinds, cap the attempts and the total time, and treat a failure around commit as ambiguous - the work may already be durable, so retrying it safely requires the operation to be idempotent.

code

pseudocode · 16 lines
pseudocode
attempt = 0
while true:
    attempt = attempt + 1
    unit = beginUnit()                      # new transaction, empty tracked set
    try:
        order = load(Order, orderId)        # read inside the attempt
        if order.balanceDue > 0:            # decision re-derived, not carried in
            order.status = "PAID"
            flush(unit)
        commit(unit)
        return DONE
    catch failure:
        rollback(unit)                      # tracked set discarded with the unit
        if not isTransient(failure) or attempt >= maxAttempts:
            raise failure
        pauseBeforeNextAttempt(attempt)

go deeper

for a junior

Remember that a retry means running the whole piece of work again in a new transaction, not calling the failed line again. The objects from the previous attempt cannot be reused.

for a middle

Explain why a retry inside the boundary writes into a condemned unit, and list what one attempt must contain: a new transaction, an empty tracked set, the reads, and the writes.

for a senior

Demonstrate the classification - transient, permanent, ambiguous - a bounded loop, and a clean separation between the retried body and side effects that cannot be rolled back.

for a principal

Set the policy: which operations are allowed automatic retry at all, what idempotency they must provide, what the caps are, and what a climbing retry rate obliges the team to investigate.

## Why the retry cannot live inside the unit When a statement is rejected, the transaction is normally marked so that it can only be rolled back, and the tracked set holds objects that no longer describe stored rows. A loop placed around the failing call therefore re-runs work into a container that is already destined for the bin. It looks like a retry and behaves like an expensive way of failing. The boundary is the unit of retry because the boundary is the unit of atomicity. Anything smaller cannot be re-run in isolation; anything larger drags unrelated work into repeated execution. ## What one attempt must contain 1. **A new transaction.** Started by the attempt, ended by the attempt, whichever way it goes. 2. **A fresh tracked set.** Empty at the start; discarded with the transaction at the end. 3. **The reads.** Every object the work needs is loaded again inside the attempt. Objects carried in from the previous attempt hold values and version numbers read in a transaction that was rolled back. 4. **The decision.** Whatever the work computes from what it read - can this order ship, is the balance sufficient - is computed again. Carrying the previous attempt's conclusion into a new attempt applies a judgement about state that has since changed, which is worse than the original failure. 5. **The writes.** Recorded and flushed inside the attempt, committed by it. The rule that follows from this list: **the retried block starts before the first read, not before the first write.** ## Which failures are worth retrying - **Retry**: kinds that are definitive and transient - a unit aborted to break a cycle of lock waits, a conflict that the engine could not order, an optimistic version check that failed because someone else got there first. Nothing was applied; another attempt against current data is likely to succeed. - **Do not retry**: constraint violations, conversion errors, anything caused by the input. The identical work fails identically, and the attempts only add load while the caller waits. - **Retry only if idempotent**: anything ambiguous, above all a connection lost around commit. The engine may have committed before the link dropped, so an attempt could apply the work twice. ## Ambiguity and idempotency Ambiguous failures are the reason retry is a design decision rather than a wrapper. If the operation can be identified - a client-supplied request key, a natural key, a state transition that is only legal once - then a second attempt either finds the work already done or performs it, and both outcomes are correct. If it cannot, a retry may double a payment or duplicate a row, and the honest answer is to surface the ambiguity rather than guess. ## Bounding the loop - Cap the **number of attempts**, small and explicit. - Cap the **total elapsed time**, because attempts under contention are not cheap and a caller is waiting. - Space attempts out rather than re-running immediately, so a contended row is not hammered by the same client. How to space them - the shape of the delay, and what several clients retrying together do to a dependency - is a general resilience concern rather than a data-access one. - Record every retry. A retry rate that climbs is a design signal, not noise: it usually means two code paths touch the same rows in a different order, or a unit is holding contended rows for far too long. ## What must stay out of the retried block Anything that cannot be undone by a rollback, because the block runs more than once: - Calls to external systems, notifications, and messages published to other services. Work that must happen exactly once after the data is durable belongs in the mechanism for running work after a successful commit, not in the retried body. - Consuming a stream or an input that cannot be re-read. - Mutating shared in-memory state - a cache entry, a counter - which the rollback will not undo, leaving it inconsistent with the rows after a failed attempt. ## Where layers differ Some data-access layers offer a built-in wrapper that re-runs a declared boundary on selected failure kinds; others leave it entirely to the caller. The built-in form is convenient precisely because it enforces the rule above - it re-enters the boundary from outside, so a fresh unit is automatic. Where it does not exist, the same shape has to be written by hand, and the mistake to watch for in review is a retry loop that sits one level too far in, inside the boundary it is supposed to re-enter.

  • Why must the reads be inside the retried block and not just the writes?
    Because the work's decision depends on them. Objects read in the failed attempt come from a rolled-back transaction and their versions are stale, so re-applying a conclusion drawn from them writes a judgement about state that has since changed. Reloading inside the attempt is what makes the second run a genuine re-execution rather than a replay.
  • How do you stop a retry from sending the same notification twice?
    Keep the side effect out of the retried body. Record the intent as part of the transactional work and let the actual send happen once the unit has committed, through the mechanism for work that runs after a successful commit. Anything performed inside the block runs once per attempt, and a rollback cannot take it back.
  • What does a rising retry rate usually tell you?
    That contention is structural rather than incidental - typically two paths touching the same rows in different orders, or a unit holding hot rows across slow work. Retries then paper over a design problem while adding load. Treat the rate as a metric with a threshold, and investigate the access pattern rather than raising the attempt cap.

saying these in an interview costs you the question

  • Putting the retry loop inside the transaction it is meant to re-run
  • Reusing objects loaded during the failed attempt in the next one
  • Retrying a constraint violation caused by the submitted input
  • Retrying an ambiguous commit failure on a non-idempotent operation
  • Leaving external calls or notifications inside the retried block
  • Running an unbounded retry loop with no cap on attempts or elapsed time