skip to content

Because a saga is a sequence of independently committing local transactions rather than one ACID transaction, other operations can observe and act on data mid-saga. What isolation anomalies does this create, and what techniques mitigate them without resorting to a distributed lock?

level: principalimportance: nice to knowfreq 30%

answer

  1. saga gives up 'I' of ACID between step commits
  2. lost update / dirty read / fuzzy read named anomalies
  3. semantic lock = status flag, not a DB lock
  4. commutative updates avoid lost updates by construction
  5. protect only high-value fields, not everything

basics

~20 s

While a saga is still in progress, other users or processes can see and act on the half-finished data, which can cause weird bugs like double-spending a discount or acting on data that later gets reversed. Teams fix this with tricks like marking records 'pending' so others know to be careful with them, instead of using a database-wide lock.

solid answer

~60 s

Sagas give up the 'I' in ACID: between step commits, the system is in a business-visible intermediate state that other concurrent transactions — including other sagas — can read and act on, producing classic anomalies. Lost update: two transactions both read a stale pre-write value and both write, and one overwrite is lost. Dirty read: a transaction reads data written by a saga step that later gets compensated away, acting on a value that's semantically 'never really happened.' Non-repeatable/fuzzy read: a saga reads the same aggregate twice across two steps and gets different values because another transaction updated it in between. Mitigations avoid distributed locks (which would reintroduce 2PC-like blocking) and instead use semantic locks — a status flag marking a record PENDING so other transactions treat it specially; commutative updates (design operations like 'add/subtract N' so that order of application doesn't matter, avoiding lost updates); pessimistic view (route reads of in-flight aggregates through a check that returns a 'not yet available' response); and re-ordering steps so the riskiest read happens as late as possible, after critical state has stabilized.

go deeper

for a junior

Not expected to know the named anomalies; should just recognize that a half-finished saga can be seen by other parts of the system, so things can look 'weird' mid-flight.

for a middle

Should be able to describe one concrete anomaly with a simple example, like two things trying to use the same reserved item.

for a senior

Should name at least lost update and dirty read specifically, and describe semantic locking as a status-flag pattern applied to a specific field, not a database-level fix.

for a principal

Should discuss all three anomaly types with concrete scenarios, contrast semantic lock / commutative updates / pessimistic view as distinct mitigations, and make the cost-benefit call on which fields warrant protection versus which can tolerate the anomaly.

## Why a saga gives up isolation A single ACID transaction gives you Isolation as one of its four guarantees: concurrent transactions can't observe each other's uncommitted intermediate state, and the database enforces this with locks or MVCC snapshots under the hood. A saga deliberately gives this up — it exists specifically because holding those locks across service boundaries (which is what 2PC would require) is unacceptable for availability. The consequence is that between the commit of saga step 2 and the commit of saga step 3, the effects of step 2 are fully durable and visible to every other transaction in the system, including transactions that have nothing to do with this saga, for as long as the saga takes to either finish or compensate. ## The classic anomalies This produces classic anomalies from the original sagas literature, applied here in a microservices context. - **A lost update** happens when two concurrent operations both read a value before either writes, and one write clobbers the other; in a saga context this shows up when, say, an inventory saga step and an unrelated direct-purchase transaction both read `available_quantity = 10` before either commits its change, and whichever commits second silently discards the other's decrement, leaving the counter wrong by the lost amount. - **A dirty read** happens when a transaction reads data that a saga step wrote but that step is later compensated away — for example, a loyalty-points balance that a saga step incremented gets read and spent by the customer in a completely different transaction, and moments later the saga fails and compensates by decrementing the points back, leaving the balance negative or inconsistent relative to what was actually spent. - **A non-repeatable (fuzzy) read** is when a saga itself reads the same aggregate at two different steps and gets two different answers because an unrelated transaction modified it in between — a discount-eligibility check at step one says "eligible," but by a later step a concurrent transaction has already consumed the customer's one-time discount, and the saga proceeds on stale information. ## Why the naive fix is worse The naive fix — take out a database lock and hold it across the whole saga — reintroduces exactly the coupling and blocking problem sagas were adopted to avoid; it's effectively hand-rolled 2PC and defeats the purpose. The established mitigations instead work at the application/business level. ## The mitigations 1. **Semantic lock**: mark the record with an explicit status flag (e.g., a `PENDING` or `RESERVED` state on the row) the moment a saga step touches it. Other transactions are written to check this flag and respond accordingly — refuse the operation, queue it, or route it through a compensating-aware path — rather than the database silently allowing an unaware read or write to interleave. This is "semantic" because it's enforced by application logic reading and respecting a business field, not by the database's native locking machinery, so it doesn't hold a real lock or block other unrelated work. 2. **Commutative updates**: design the operations themselves so that the order they're applied in doesn't matter to the final result — "increment balance by N" and "decrement balance by N" commute regardless of interleaving, unlike "set balance = X" which requires knowing the current value and can silently lose an update if two set operations race. Where you can express saga effects as commutative deltas rather than absolute overwrites, lost updates stop being possible by construction. 3. **Pessimistic view**: for reads that must not observe an in-flight saga's intermediate state, route them through a check that recognizes a semantic-locked record and returns a controlled "not currently available" response (or blocks/queues briefly) rather than returning a value that might be compensated away moments later. This trades a slightly worse UX (a temporary "please retry" for a record mid-saga) for avoiding a dirty read of state that's not yet final. 4. **Reordering and step design**: place operations vulnerable to fuzzy reads (like discount-eligibility checks) as late in the saga as practically possible, ideally re-validated right before the point of no return, rather than trusting a check performed several steps and network round-trips earlier. ## Why these are only partial These are genuinely partial mitigations, not a full restoration of isolation — they reduce the frequency and blast radius of anomalies for the specific fields you've thought to protect, but a saga fundamentally cannot offer the blanket guarantee a single ACID transaction does, and teams need to explicitly decide, per field, whether the business impact of a possible anomaly (an over-sold item, a double-spent discount) is tolerable or needs one of these mitigations. In practice, most teams apply semantic locks only to a small number of high-value fields — inventory counts, account balances, one-time discount flags — rather than trying to protect every field a saga touches, because the engineering cost of full protection everywhere would erode most of the availability benefit the saga was adopted for in the first place.

  • Why not just hold a database lock on the affected rows for the duration of the whole saga?
    That reintroduces the exact blocking and cross-service coupling problem 2PC has and sagas were adopted to avoid — a lock held across multiple network round-trips and service calls ties every other transaction's availability to this saga's completion time, which can be seconds to minutes. Semantic locks give a similar protective effect without ever taking a real, held database lock.
  • Are these anomalies specific to choreography, or do they also affect orchestration-based sagas?
    They affect both — the anomalies come from the lack of a spanning ACID transaction across the steps' local commits, which is a property of the saga pattern itself, not of how the next step gets triggered. Orchestration makes the saga's own progress easier to track, but doesn't by itself prevent an unrelated concurrent transaction from reading or writing the same data mid-saga.
  • How would you decide which fields need a semantic lock versus which can tolerate the anomaly?
    Weigh the business cost of the anomaly against the cost of protecting it: a one-time discount code or an account balance where double-spending has real financial or fraud impact typically warrants a semantic lock, while a low-stakes display field (like a 'last viewed' timestamp) that might briefly show stale or dirty data usually doesn't justify the added complexity.

It's like a house under renovation that stays occupied and open to visitors between construction phases instead of being sealed off until the whole job is done: a visitor mid-renovation might see a wall half-torn-down and later put back (a dirty read), or two contractors might each grab the 'last bag of cement' without knowing the other already took it (a lost update) — the fix isn't locking the house for the whole renovation, it's putting up a clear sign on the specific room that's mid-work.

saying these in an interview costs you the question

  • Claims sagas provide the same isolation as a single ACID transaction
  • Proposes a cross-service database lock held for the saga's full duration as the fix
  • Can't name a concrete anomaly (lost update, dirty read, fuzzy read) or give an example
  • Thinks semantic locks are literal database locks rather than an application-level status-flag convention
  • Assumes orchestration alone solves the isolation problem

context