Why declare a timeout on a transaction boundary, and what happens to the work when that timeout expires?
answer
- a deadline set when the unit begins
- bounds lock and connection hold time
- checked between statements, pushed down where possible
- expiry means full rollback, never partial
- shorter than the caller's own deadline
basics
~20 sA timeout caps how long one unit may hold locks and a connection, so one stuck statement cannot pile callers up behind it. When the deadline passes the unit is failed and rolled back in full - nothing partial survives.
solid answer
~50 sA deadline is recorded when the boundary begins. Its job is blast-radius control: without one, a blocked statement can hold row locks indefinitely while every other caller queues behind it, and a slow path becomes an outage. Enforcement is layered and varies. A layer typically checks the remaining budget before issuing each statement and refuses to start one past the deadline; where the driver allows it, the remaining time is passed down as a per-statement limit so a statement already in flight can be aborted; the database may enforce limits of its own. That means the deadline is not a hard wall on wall-clock time - a single long statement, or slow application code between statements, can overrun it. When it does fire the unit rolls back in full. Size it below the caller's own deadline, so the work gives up before the client stops waiting on locks it still holds.
go deeper
Know that a boundary can carry a deadline, and that when it passes the whole unit is rolled back rather than partly saved.
Explain where the deadline is checked - between statements, and pushed down to the statement where the driver allows it - and why that makes it a soft rather than an exact wall.
Reason about blast radius: a unit with no deadline turns one stuck statement into a queue of blocked callers, and a timeout shorter than the caller's deadline is what keeps retries from stacking.
Treat deadlines as a system-wide budget: each layer's limit strictly inside the one above it, values chosen per use-case shape, and timeout rates tracked as a contention signal rather than as noise.
## What a declared timeout actually is A timeout attached to a boundary is a **deadline recorded when the unit begins**: start time plus the declared budget. It is an attribute of the unit, not of any one statement, and it exists for one reason - to bound how long this unit may hold resources that other callers need. The resources in question are the locks its writes have taken and the connection it occupies. A unit with no deadline holds both for as long as it takes, which in the bad case is "until something else times out". ## Why it matters more than it looks Contention is non-linear. One transaction stuck for thirty seconds on a hot row does not cost thirty seconds; it costs thirty seconds multiplied by everything that queues behind it, and those waiters hold their own locks while they wait. A deadline is the cheapest available circuit breaker: it converts an unbounded stall into a bounded failure that the caller can report or retry. ## How the deadline is enforced Enforcement is layered, and layers genuinely differ in how much of it they do: - **Before each statement.** The layer compares the remaining budget against the deadline and refuses to start a new statement once it has passed. This is the most portable check and the most common one. - **Per statement, pushed down.** Where the driver supports a statement-level limit, the layer passes the remaining time down, so a statement that blocks can be aborted by the database or the driver rather than waited out. - **By the engine.** Databases commonly offer their own statement and lock-wait limits, configured independently of the application. - **Not at all, between statements.** Time spent in application code inside the boundary is generally not interrupted; the check happens the next time the layer is asked to do something. The honest summary: the deadline is a promise about *when the unit stops making progress*, not a guarantee that it dies at that instant. ## What happens when it fires 1. The unit is failed. There is no partial commit and no "commit what we have" - the whole boundary rolls back. 2. Whatever the layer accumulated for this unit is discarded along with it. 3. The caller receives a timeout-flavoured error, distinguishable in most layers from a constraint failure or a conflict. 4. Locks are released as the transaction ends, which is the entire point of the exercise. Retry is a decision, never automatic. It is safe only when the use case is idempotent or has produced no effect outside the database, and a blind retry of a timing-out unit is the classic way to turn a slow system into a dead one. ## Timeout versus cancellation These are different mechanisms and mixing them up causes real incidents. | | Transaction timeout | Caller cancellation | |---|---|---| | Owner | the unit of work | the request or job | | Fires on | its own deadline | the client hanging up or a shutdown | | Effect by itself | rolls the unit back | nothing, unless propagated | A cancelled request does not end a transaction. The thread carries on holding it until the current statement returns and the boundary is closed. To make cancellation mean anything, it has to be turned into a statement cancellation plus a rollback - otherwise the client disconnects, retries, and now two transactions contend for the same rows. ## Choosing a value - Start from the **caller's** deadline and set the transaction's below it, leaving slack to roll back and report. A transaction timeout longer than the client's patience guarantees abandoned work still holding locks. - Size it against the unit's **honest worst case** under load, not its median. A deadline that trips routinely is noise; one that never trips is decoration. - Keep write units short enough that the deadline is generous. A timeout is a safety net for the tail, not a substitute for a boundary that is too wide. - Watch out for a **statement** limit set larger than the unit's deadline: the layer's check between statements will fire eventually, but a single query can still overrun the unit by the difference. ## Traps worth naming - Assuming the deadline preempts application code - a slow loop inside the boundary runs to completion. - Setting one global value across use cases with wildly different shapes. - Treating a timeout error as a bug in the database rather than as evidence about lock waits and unit duration. - Retrying on timeout without idempotency, which multiplies the load that caused the timeout.
- Is a timeout error safe to retry?Only when the use case is idempotent or has had no effect outside the database. The unit itself rolled back, so the data is clean, but the cause was usually contention - and retrying immediately adds load to exactly the thing that was slow. Retry with a bounded count and a growing delay, or surface the failure.
- What does a rising rate of transaction timeouts usually tell you?That units are holding locks longer than the workload tolerates. Look first at boundary width - computation or waiting inside the boundary - then at the rows being contended and how long they stay locked. Raising the timeout treats the symptom and moves the queue somewhere less visible.
saying these in an interview costs you the question
- Expects the deadline to interrupt application code running inside the boundary
- Thinks a timeout commits the work completed so far before failing
- Assumes cancelling the caller's request also ends its open transaction
- Retries a timed-out unit blindly without checking it is safe to repeat
- Sets the transaction timeout longer than the client's own deadline