skip to content

A saga has already completed step 1 (reserve inventory) and step 2 (charge payment) while step 3 (arrange shipping) is still pending. Another transaction reads the order in this intermediate state. What consistency problem does this create, and how does the 'semantic lock' technique help address it?

level: principalimportance: nice to knowfreq 30%

answer

  1. application-level flag, not a DB lock
  2. PENDING/status field + owning saga ID
  3. one of the classic saga countermeasures
  4. fixes dirty reads/lost updates from lack of isolation
  5. orphaned-lock risk needs timeout/cleanup

basics

~20 s

A semantic lock is a flag added to a record (like 'pending') while a saga is still in progress, so other parts of the system know not to fully trust or freely modify that record until the saga finishes.

solid answer

~60 s

Since sagas don't give isolation, another transaction can read a record mid-saga and see a partially-applied state - inventory decremented but payment not yet confirmed, for example. A semantic lock is an application-level marker (a status field like PENDING, IN_SAGA, or CONFIRMING, plus an owning-saga ID) set on the record when a saga step touches it, which other transactions and business logic are expected to check and respect: they might block, queue, reject, or route around a locked record rather than treating it as a normal committed value. It's called 'semantic' because it's enforced by application logic and convention, not the database engine - there's no real lock being held, just a business-rule signal that consuming code must honor for the technique to work. It addresses the classic saga isolation anomalies: dirty reads (seeing pending, not-yet-final data), lost updates (a second saga overwriting a record another saga is mid-way through), and non-repeatable reads - at the cost of extra state to manage, and the risk of an orphaned lock if the owning saga crashes without releasing it, which needs its own timeout/cleanup mechanism.

go deeper

for a junior

Not expected to know this technique by name; a basic 'saga data can be seen in a not-fully-done state' intuition is enough.

for a middle

Should recognize the general problem (in-progress saga state is visible) even without knowing the specific 'semantic lock' term.

for a senior

Should know the term, describe the mechanism (status field, application-enforced), and give a concrete example.

for a principal

Should discuss the orphaned-lock failure mode, know it's one of several named saga countermeasures, and be able to compare it against alternatives like commutative updates or pessimistic view for a given scenario.

## One of a family of countermeasures **Semantic locking** is one of a family of countermeasures originally described for handling the isolation anomalies that sagas inherently create, and it's the most commonly used one in practice. Alongside it: - commutative updates, - pessimistic view, - reread value, - version files, - by value. ## Why the technique is needed To see why it's needed, recall that a saga is a sequence of independently-committed local transactions, with no cross-service lock held between steps - that's exactly what gives sagas their availability advantage over two-phase commit. The consequence is that between step 1 committing and the saga's final step committing, the record(s) touched by step 1 are in a normal, fully-committed, publicly-visible state as far as any other transaction in the system is concerned - there's nothing marking them as 'not really done yet.' If an order record has its inventory decremented by step 1 but the saga's payment step hasn't run yet, a completely unrelated transaction - a warehouse report, a different saga, a customer's own order-status page - can read that record and see 'inventory reserved,' with no signal that the reservation might still be undone by a compensation a moment later. The anomalies this creates: - **dirty read** - this is a dirty read in the classical database sense; - **lost updates** - sagas are also vulnerable to two sagas concurrently modifying overlapping data, where the second saga's write overwrites the first's without knowing about it; - **non-repeatable reads** - reading the same record twice within one logical operation and getting different answers because a saga committed a step in between. ## How the marker works A semantic lock addresses this by having the application, not the database engine, add an explicit marker to the record while it's mid-saga. Typically this is a status field on the record - `PENDING`, `RESERVING`, `IN_SAGA` - sometimes combined with the ID of the saga instance that owns the lock, set as part of the same local transaction that makes the change. Any other code path that reads this record is expected, by convention and by explicit checks in the business logic, to treat a locked record differently. It might: - refuse to let another saga touch the same record until the lock clears; - queue its own operation until the status changes; - show the customer 'processing' instead of a final order state; - or in some designs actively trigger compensation of an earlier conflicting saga to release the lock faster. ## Why the word is 'semantic' The word 'semantic' signals precisely that this is not a real database lock - no row-level lock is held, no other transaction is physically blocked from reading or even writing the row if it doesn't check the status field. The entire mechanism depends on every piece of code that touches this data agreeing to check and honor the flag; it's a convention enforced by discipline and code review, not by the database, which is both its main strength (cheap, doesn't hold real locks, doesn't block availability) and its main weakness (a single forgotten check anywhere in the codebase silently defeats it). ## The cost side The cost side of this trade-off is real. - You now have extra state to design, migrate, and reason about - every record touched by a saga needs a status field and the business logic to interpret it correctly in every context that reads it, which is easy to get partially right and miss an edge case (a new report added six months later that queries the table directly and doesn't know to filter out `PENDING` rows, for instance). - There's also the **orphaned-lock** problem: if the saga that set the lock crashes or gets stuck, the record can be left semantically locked forever unless there's a separate timeout or reconciliation job that detects stale locks and either resumes or force-clears them - which itself needs careful design so it doesn't clear a lock for a saga that's actually still legitimately in progress, just slow. ## A concrete example A concrete example: in an inventory-reservation saga, the moment stock is decremented, the product record is marked with a reservation status and the owning order ID; the product listing page and other checkout flows are written to treat that status as 'unavailable' rather than silently showing (or worse, allowing purchase of) inventory that's provisionally spoken for but not yet guaranteed. If the payment step later fails and the saga compensates by releasing the stock, the status flips back to available. This is a very common real pattern in e-commerce and ticketing systems (seat-hold timers on ticketing platforms are functionally the same idea - a semantic lock with a built-in expiry), and it's worth knowing by name specifically because interviewers use it to check whether a candidate understands that saga isolation problems are real and need an explicit, named countermeasure, rather than being hand-waved away as 'eventually consistent, so it's fine.'

  • What happens if a saga crashes while holding a semantic lock and never releases it?
    The record stays marked as locked/pending indefinitely unless something else intervenes, which effectively makes that data permanently unavailable to the rest of the system - a real production bug. The standard fix is a reconciliation or timeout job that periodically checks for locks older than some threshold, checks the owning saga's actual status, and either resumes/completes the saga or force-releases the lock as part of a compensation.
  • Are there alternatives to semantic locking for handling saga isolation problems?
    Yes - commutative updates (design operations so order doesn't matter, avoiding the conflict entirely), pessimistic view (sequence saga steps so the riskiest one runs first, minimizing the anomaly window), reread value (re-check a value hasn't changed before finalizing), and versioning (keep old and new values so concurrent readers can choose which to see). Semantic locking is popular because it's the most straightforward to implement and reason about, but it's not the only tool, and some alternatives avoid the orphaned-lock risk entirely.
  • Why doesn't the database just do this locking automatically?
    Because the database has no concept of 'this row is part of a multi-step, cross-service business process that hasn't finished yet' - that's a business-level, cross-service fact that only the application knows, since the saga's other steps live in entirely different databases the current database can't see. A real database lock also can't span services, and even within one service it would reintroduce the blocking/availability problem sagas are specifically designed to avoid.

It's like a 'reserved' sign a restaurant puts on a table before the party has actually arrived and paid - the table isn't physically locked, any staff member could still seat someone there, but everyone is expected to see the sign and treat the table as unavailable until it's cleared.

saying these in an interview costs you the question

  • Confuses semantic locks with actual database row locks
  • Doesn't know semantic locks can be silently ignored by code that doesn't check the status field
  • Has no answer for what happens to an orphaned semantic lock
  • Thinks sagas provide isolation by default and semantic locks are unnecessary
  • Can't name a concrete anomaly (dirty read/lost update) that semantic locks address

context