skip to content

A unit of work already holds a loaded copy of a row; why must a later pessimistic lock request force a re-read?

level: middleimportance: must knowfreq 55%

answer

  1. the in-memory copy came from an unlocked read
  2. identity map answers the second load
  3. a window between read and claim
  4. lock protects the row, not your old values
  5. re-read under the lock or lock on load

basics

~20 s

The copy in the tracked set was read without a lock, so another transaction may have committed since. A lock taken without re-reading protects a row whose current values you have never seen. Your write then loses that change.

solid answer

~50 s

A layer that tracks loaded objects keeps an identity map: the second request for the same row normally returns the in-memory copy without touching the database. Those values were read unlocked, so between that read and your lock request another transaction can commit a change to the row. If the lock request only takes the lock, you end up in the worst position available: you hold the row exclusively, you believe your fields are current, and they are not — you then compute from the old values and write them back, losing the other transaction's change exactly as if you had never locked at all. So the lock request must re-read the row under the lock and refresh the tracked copy, and anything already computed from the pre-lock values must be recomputed. The cheaper fix is to ask for the lock on the original load.

go deeper

for a junior

Hold on to one sentence: a lock protects the row from the moment it is taken, not the values you read before taking it.

for a middle

Explain the identity map's role — the second load is answered from memory — and why that turns a late lock request into a lock over stale fields.

for a senior

Talk about how the bug hides: single-threaded tests pass, the locking statement is in the log, and the loss only shows up at production concurrency.

for a principal

Push the decision upstream: make write paths declare their intent at load time so late locking is the rare, reviewed exception rather than the default habit.

This is the trap that makes pessimistic locking look like it is working when it is not, and it exists only because the layer keeps loaded objects in memory. ## Why the copy is there at all A unit of work maintains an **identity map**: one in-memory object per row identity for the life of the unit. It is what makes repeated lookups cheap and what guarantees that two code paths handling "order 42" hold the same object rather than two divergent copies. The consequence is that a second load of an already-tracked row is usually answered from memory, with **no statement sent to the database at all**. That is fine while you are only reading. It becomes dangerous the moment you decide, part-way through the unit of work, that you now need to lock the row. ## The window nobody sees Consider the ordinary sequence: 1. The unit of work loads the row with a plain read. No lock is taken, and under a version-based engine the read does not even block anyone. 2. Business logic runs. Meanwhile another transaction updates that row and commits. 3. Your code decides the update needs protecting and issues a pessimistic lock request for the object it already has. 4. The lock is granted — the other transaction has already committed, so nothing is in the way. You now hold an exclusive claim on the row and an object whose fields are from **step 1**. Every guarantee people associate with locking is absent: the values you are about to compute from are stale, and when the unit of work flushes, your write puts the pre-step-2 state back. The other transaction's change is gone. The lock did not fail; it protected the wrong instant in time. ## What a correct lock request does A lock request against an already-tracked object cannot be answered from the identity map. It must go back to the database, take the lock, and **read the row's current state under that lock**, refreshing the tracked copy with what comes back. Layers differ in how much of this they do for you: some treat a lock request as implying a refresh, others take the lock and leave your fields untouched unless you explicitly re-read, and some carry a version column and will fail the lock request outright when they see the row has moved — a loud failure that is far better than the silent one above. Because that behaviour is not uniform, the defensive habit is to state it explicitly rather than assume it. | Situation | Safe? | Why | |---|---|---| | Lock mode requested on the original load | Yes | Claim and values arrive in the same statement; no window exists | | Lock request on a tracked object, with a re-read under the lock | Yes | The refreshed values are the ones the lock protects | | Lock request on a tracked object, no re-read | **No** | The lock protects a row whose current state you have not seen | | Plain re-read first, lock afterwards | **No** | Reintroduces the same window between reading and claiming | ## The rule that falls out of it **Anything read before the lock is untrusted after it.** In practice: - Ask for the lock on the load that first brings the row in, whenever the code already knows it is going to write. This removes the problem instead of managing it. - When the decision to lock genuinely comes later, treat the lock request as a re-read: discard derived values, recompute totals and checks from the refreshed fields, and re-evaluate the business precondition that made you decide to write. - Never validate an invariant before acquiring the lock and then write after it. The validation has to happen on post-lock values, or it is decorative. - Be equally careful with a copy that has left the unit of work entirely. Values held across a boundary are older still, and a lock taken in a new transaction says nothing about what happened while they were away. ## Why it is worth naming in an interview The failure is invisible in testing. Single-threaded tests pass, the lock is genuinely acquired, the statement log shows a locking read, and the lost update only appears under real concurrency at the rate at which step 2 lands in the window. Candidates who have debugged this reach for it immediately; candidates who have only read about locking tend to believe that acquiring the lock is the whole of the job. The distinguishing sentence is short: **the lock protects the row from the moment it is taken, not the values you read before you took it.**

  • How do you avoid the problem instead of handling it?
    Decide early. If the use case is going to write the row, request the lock mode on the load that first brings it in, so the claim and the values arrive in one statement and no window exists. Late locking is only for paths where the decision genuinely depends on what was read, and those paths must recompute from the refreshed values.
  • What if the object was loaded in an earlier transaction and carried across the boundary?
    It is older again, and nothing about the previous boundary constrains what happened afterwards. A lock taken in the new transaction protects the row from that moment on, so the carried values must be treated as input to be re-checked, never as the current state to write back.
  • Does a version column make the late lock safe?
    It makes the failure loud rather than silent: a layer that notices the stored version has moved can refuse the lock request or fail the write instead of quietly overwriting. That is a much better outcome, but it is still a failed attempt to be handled, not a reason to skip the re-read.

saying these in an interview costs you the question

  • Thinks acquiring the lock retroactively validates values read earlier
  • Assumes a second load of a tracked row always hits the database
  • Checks the business precondition before locking and writes after it
  • Believes a lock request cannot be answered from memory, so no re-read is needed
  • Says the problem cannot happen because the tests pass