A pessimistic lock request through a data-access layer can block indefinitely; what bounds that wait, and how must callers handle the outcome?
answer
- the default is an unbounded queue
- a parked thread also holds a connection
- bound it below your own deadline
- failure at the boundary, not a slow success
- nothing from the failed attempt is reusable
basics
~20 sBy default the request queues for as long as the holder keeps the row. A wait timeout or an immediate-failure request bounds it, turning the block into an error that rolls the unit of work back.
solid answer
~50 sAn unbounded lock request is a hidden availability risk: the waiting caller parks a worker and a pooled connection for as long as the current holder's transaction runs, so a single slow holder turns into pool exhaustion and failures on endpoints that never touch that row. Two knobs bound it — a wait timeout that gives up after a set period, and a no-wait request that fails immediately if the row is held — and a layer usually maps the resulting database error onto a portable acquisition-failure exception. Both outcomes are failures, not slow successes: the transaction is rolled back and the tracked set from the failed attempt must not be reused. So the caller catches it outside the boundary, decides whether the use case can honestly be attempted again, and reports a clear conflict if not. Set the bound below the request's own deadline so the timeout you see is yours.
go deeper
Know that a lock request can wait, and that the wait can be bounded so the call fails instead of hanging forever.
Explain the three behaviours — wait, wait with a timeout, fail immediately — and that a bounded request surfaces as an error that rolls the unit of work back.
Show the operational chain: parked worker plus held connection, pool exhaustion, unrelated endpoints failing; then bound the wait under your deadline and export the wait metrics.
Argue the ordering of fixes — shorten the holder's boundary first, tune bounds second — and set which classes of path may block on a lock at all.
The interesting part of a pessimistic lock request is not the case where the lock is free. It is the case where somebody else already holds the row. ## The default is "wait as long as it takes" Unless you say otherwise, a lock request queues behind the current holder and stays queued until that holder's transaction commits or rolls back. There is no ceiling implied by the request itself; the ceiling is whatever the holder's boundary happens to be, plus everyone else already in the queue. In a request-serving service that is more dangerous than it sounds, because the wait is not free while it lasts: - A worker thread is parked in a blocking database call. - A pooled connection is held for the whole wait, on top of the connection the holder is using. - The caller upstream — a browser, a gateway, another service — is running its own timeout, which will usually fire first and abandon the request while your side keeps waiting. The failure mode is therefore not "this one request is slow". It is **pool exhaustion**: enough waiters accumulate on one hot row that unrelated endpoints cannot get a connection, and a localised contention problem is reported as a total outage. ## The two bounds | Option | Behaviour when the row is held | Fits | |---|---|---| | Wait (default) | Queue until the holder's transaction ends | Short, predictable holders; background work | | Wait with a timeout | Queue, then fail after a set period | Interactive paths that must answer within a deadline | | No wait | Fail immediately, without queueing | Paths where a held row means "someone else is already doing this" | A no-wait request is not merely an impatient timeout of zero. It changes the semantics of the call from "get me this row eventually" to "tell me whether the row is free", which is a genuinely different question and often the one the use case is really asking. ## How the failure reaches your code Whichever bound fires, the database raises an error rather than returning rows, and the layer normally translates it into a portable acquisition-failure exception instead of leaking a vendor code. Two consequences matter more than the exception's name: 1. **The transaction is finished.** The unit of work is rolled back at the boundary, and any objects tracked in the failed attempt are unusable — their state reflects work that never committed. Nothing from the failed attempt may be carried into a second one. 2. **It is a conflict, not a bug.** It says another transaction held the row, which is a normal condition on a contended row, and it should not be logged and alerted like a defect. So the handling lives **outside** the boundary, where a fresh unit of work can be started, and it starts by deciding a question the layer cannot answer: is repeating this use case honest? Repeating a pure read-modify-write usually is, since everything is read again from the database. Repeating anything that already had an effect outside the transaction is not. ## Choosing the bound Some practical rules that survive contact with production: - Bound the wait **below the request's own deadline**, so the error you handle is your acquisition failure rather than an upstream cancellation you cannot see or explain. - Keep the bound well **above the normal holder's duration**, or you convert ordinary contention into a stream of failures. If you cannot pick a value that satisfies both, the holder's boundary is too long and that is the thing to fix. - Prefer **no wait** where a held row means the work is already being done by someone else and duplicating it is pointless. - Give background and batch paths a longer bound than interactive ones; they have no user waiting and their retries are cheap. - Make acquisition failures and lock wait time **visible as metrics**, not just log lines. The rate of failures and the distribution of wait times are what tell you whether the bound is set sensibly, and they are the early warning that a holder's boundary has grown. ## The design point underneath Every knob here manages a symptom of the same cause: someone holds the row for a long time. Shortening the boundary that holds the lock — no remote calls, no user interaction, no unrelated work between claim and commit — reduces waits, acquisition failures and pool pressure at once, and it is the fix that keeps working as load grows. The timeout is what makes the failure survivable and diagnosable in the meantime.
- Why is an unbounded wait an availability problem rather than just a latency problem?Each waiter holds a worker and a pooled connection for the whole wait. Enough of them on one hot row exhausts the pool, so endpoints with nothing to do with that row start failing too. The blast radius of one slow holder becomes the whole service.
- What must not be carried from a failed acquisition into a second attempt?Anything from the rolled-back unit of work: the tracked objects and any values computed from them reflect a transaction that never committed. A second attempt starts a new boundary and re-reads everything it needs, so the retry decision belongs outside the boundary, not inside it.
- When does a no-wait request beat a short timeout?When a held row means the work is already in progress elsewhere and duplicating it is pointless. Then the useful answer is an immediate 'someone else has it', not a delayed one, and failing at once frees the worker and the connection instead of parking them for the timeout.
saying these in an interview costs you the question
- Leaves the wait unbounded because contention is assumed to be rare
- Treats an acquisition failure as a bug to alert on rather than a conflict
- Retries inside the same rolled-back unit of work
- Sets the lock timeout above the request's own deadline
- Thinks a waiting request costs nothing while it waits
- Tunes the timeout instead of shortening the boundary that holds the lock