skip to content

Compare basic two-phase locking with strict 2PL and with rigorous (strong strict) 2PL: in each variant, which locks are held until the transaction commits, and what does a database engine gain by choosing a stricter variant?

level: middleimportance: should knowfreq 40%

answer

  1. Basic: release early → dirty reads, cascades
  2. Strict: X locks to commit → cascadeless, safe undo
  3. Rigorous: S and X to commit → order = commit order
  4. Strictness costs lock hold time
  5. Isolation levels = how much read-side 2PL you keep

basics

~20 s

Basic 2PL may release any lock once it stops acquiring. Strict 2PL holds all exclusive (write) locks to commit or abort. Rigorous 2PL holds all locks, shared and exclusive, to commit. Stricter buys recoverability, no cascading aborts, and commit-order serialization — at the cost of concurrency.

solid answer

~60 s

All three obey the growing/shrinking rule; they differ in how late the shrinking phase happens. - **Basic 2PL** — release locks as soon as no more will be acquired. Highest concurrency, but another transaction can read data you wrote and have not committed, which permits dirty reads, **cascading aborts** and schedules that cannot be recovered after a crash. - **Strict 2PL** — hold every **exclusive** lock until commit or abort; shared locks may go earlier. Nobody can read or overwrite uncommitted data, so schedules are cascadeless and rollback by restoring before-images is always safe. - **Rigorous / strong strict 2PL (SS2PL)** — hold **all** locks, shared and exclusive, until commit. Now the serialization order equals the **commit order**, which simplifies recovery, replication and distributed commit; it is also what SERIALIZABLE means in a lock-based engine. The cost is monotonic: each step lengthens lock hold time, so contention, lock-wait time and deadlock rate all rise. Most commercial lock-based engines implement rigorous 2PL and let you *opt down* through isolation levels.

go deeper

for a junior

Know the ordering — basic, then strict (write locks to commit), then rigorous (all locks to commit) — and that stricter means safer but slower.

for a middle

Attach each variant to its concrete guarantee: cascadeless and recoverable for strict, commit-order serialization for rigorous.

for a senior

Explain why write-lock strictness is non-negotiable for undo correctness while read locks are the isolation dial, and quantify the cost as lock hold time.

for a principal

Discuss the composition properties — SS2PL under two-phase commit giving global serializability — and when you would move a workload off locking entirely rather than pay that hold time.

## The one rule they share Every variant here is two-phase: a **growing phase** in which the transaction only acquires locks, then a **shrinking phase** in which it only releases them. The variants differ solely in *when the shrinking phase is allowed to start*, and that single knob controls a large amount of practical database behaviour. ## Basic 2PL A transaction releases each lock as soon as it knows it will acquire nothing more. This yields conflict-serializable schedules (that guarantee comes from the two-phase rule alone) and the most concurrency of the three, because locks are held for the shortest time. But it releases write locks *before commit*. So T2 can acquire a lock on a row T1 has already modified and read T1's uncommitted value. Three bad things follow: 1. **Dirty reads** — T2 sees data that may never exist in a committed state. 2. **Cascading aborts** — if T1 rolls back, T2 must roll back too, and anything that read T2's writes must roll back as well; a single failure can unwind an unbounded fan-out of transactions. 3. **Non-recoverable schedules** — if T2 *commits* before T1 aborts, the system is stuck: it cannot un-commit T2, yet T2's result was derived from work that never happened. Rollback also gets structurally hard, because restoring T1's before-image of the row would silently wipe T2's later write. Because of this, basic 2PL is a teaching device rather than an implementation choice. ## Strict 2PL Strict 2PL adds one requirement: **all exclusive (write) locks are held until the transaction commits or aborts.** Shared locks may still be released earlier. This produces a **strict schedule** — no transaction reads or overwrites an item written by an uncommitted transaction. The payoffs are exactly the three problems above, reversed: - No dirty reads of a write-locked item and therefore **no cascading aborts** ("cascadeless", also called avoiding-cascading-aborts, ACA). - Schedules are **recoverable**: a transaction can only have read committed data, so it never commits on top of work that later disappears. - **Rollback is trivially correct**: since nobody else has touched the item since your write, undo by restoring the before-image cannot destroy another transaction's update. This is why physical/physiological undo logging pairs so naturally with strict 2PL. What strict 2PL does *not* give you is protection on the read side. A shared lock released early means the transaction is no longer two-phase for reads, so repeated reads of the same row can return different values once another transaction's write lock is available. That is precisely how the weaker isolation levels are built. ## Rigorous (strong strict) 2PL Rigorous 2PL holds **every** lock — shared and exclusive — until commit or abort. The shrinking phase collapses into a single instant at commit. The extra property this buys is ordering: because a transaction's lock point now coincides with its commit, and precedence edges always run from an earlier lock point to a later one, **the serialization order equals the commit order**. That is enormously convenient: - The recovery log, ordered by commit, replays to the same state. - Replicas can apply the commit stream serially and stay consistent. - In a distributed setting it composes cleanly with two-phase commit — each participant holds locks until the coordinator's decision, so the global order is the global commit order. (Distributed SS2PL is the standard way to get global serializability without a central scheduler.) Rigorous 2PL is what a lock-based engine's SERIALIZABLE isolation level means in practice. ## The cost curve, and how isolation levels ride it The three variants form a monotonic trade: each step holds locks longer, so lock-wait time, blocking probability, deadlock rate and lock-manager memory all increase, while achievable concurrency falls. Rather than exposing "pick your 2PL variant", engines expose **isolation levels**, which are effectively the same dial: - **READ UNCOMMITTED** — do not take read locks at all. - **READ COMMITTED** — take a shared lock for the instant of the read and drop it immediately (not two-phase for reads); write locks still held to commit, i.e. strict on the write side. Allows non-repeatable reads. - **REPEATABLE READ** — hold shared locks on rows you read until commit, making reads two-phase as well; classically still allows phantoms unless range/predicate locks are added. - **SERIALIZABLE** — rigorous 2PL plus range locking to close phantoms. Notice that even the weakest useful level keeps the *strict* part (write locks to commit). That is not a tuning choice: giving it up would break rollback and recovery, not merely relax isolation. The read side, by contrast, is genuinely negotiable, and that is where the isolation-level dial lives. ## What to say when asked A crisp answer names the three variants by which locks they hold to commit, ties strict to *cascading aborts and safe rollback*, ties rigorous to *serialization order equals commit order*, and closes with the trade: strictness is paid for in lock hold time, which is the dominant driver of contention.

  • Why do engines keep write locks until commit even at READ COMMITTED, where isolation is deliberately weak?
    Because holding exclusive locks to commit is about recovery, not isolation. If another transaction could overwrite or read your uncommitted row, rolling you back by restoring the before-image would destroy their update or force them to abort as well. Strictness on the write side is what makes undo logging and cascade-free rollback correct, so it is not part of the isolation dial.
  • Which variant do distributed transactions coordinated by two-phase commit rely on, and why?
    Rigorous (strong strict) 2PL. Each participant holds all its locks until the coordinator's commit/abort decision arrives, so every site's local serialization order agrees with the global commit order and the composed schedule is globally serializable. It also means locks are held across the coordination round trip, which is why 2PC latency directly inflates contention.

saying these in an interview costs you the question

  • Using "strict 2PL" and "2PL" interchangeably
  • Claiming strict 2PL prevents deadlocks
  • Saying strict 2PL holds shared locks to commit (that is rigorous)
  • Thinking basic 2PL is what production engines implement because it is the most concurrent
  • Believing strictness is only about isolation and has nothing to do with rollback correctness

context