skip to content

Serializable Snapshot Isolation keeps snapshot isolation's non-blocking reads but adds runtime detection of read-write anti-dependencies and aborts transactions that could produce a non-serializable result. Explain how that detection works, why it aborts transactions that would in fact have been fine, and how you would decide whether to run a high-throughput OLTP system on it.

level: principalimportance: should knowfreq 30%

answer

  1. anti-dependency = read then someone wrote
  2. pivot: in-edge + out-edge = dangerous structure
  3. SIREAD markers block nobody
  4. necessary not sufficient -> false aborts
  5. long transactions degrade tracking precision

basics

~20 s

SSI tracks which transactions read data that another later wrote (anti-dependencies) and aborts when a transaction has both an incoming and an outgoing such edge — the dangerous structure. Detection is conservative, so some aborts are false positives. Adopting it means budgeting for aborts, short transactions, and a universal retry path.

solid answer

~60 s

SSI records reads — at row and index-range granularity — so it can detect a **read-write anti-dependency**: transaction A read a version that transaction B then replaced. The theory says every non-serializable snapshot-isolation history contains a cycle with two *consecutive* anti-dependency edges, so SSI does not build the full dependency graph. It just watches for a **pivot**: a transaction with both an incoming and an outgoing anti-dependency edge. When one appears and the relevant transactions commit, it aborts one of them with a serialization failure. The dangerous-structure test is necessary but not sufficient for a cycle, so SSI produces **false positives** — it aborts some serializable executions. Read tracking is also approximate: predicate reads are tracked by index ranges or pages, and under memory pressure the information is summarised more coarsely, adding more false positives. Adoption criteria: an abort rate you can absorb, uniformly short transactions, a central bounded retry with backoff, no non-idempotent side effects inside transactions, and read-only work marked as such so it can be exempted.

code

text · 7 lines
text
T1 --rw--> T2 --rw--> T3        (T3 may equal T1)
              ^
              pivot: has an incoming AND an outgoing
                     read-write anti-dependency edge

rw edge A --rw--> B  ==  A read a version that B later overwrote,
                         so any serial order must put A before B

go deeper

for a junior

Know that SSI means the database detects the conflicts snapshot isolation misses and aborts one transaction, and that your code must retry.

for a middle

Explain read-write anti-dependencies and that detection watches for a transaction with both an incoming and an outgoing edge, plus the retry obligation this creates.

for a senior

Add the false-positive story — necessary-not-sufficient test, range-granularity read tracking, degradation under memory pressure — and the operational rules: short transactions, side effects outside, read-only work marked read-only.

for a principal

Frame it as a system-wide tax versus targeted hardening: quantify the abort rate at production concurrency, weigh the retry-tail latency and CPU cost against the cost of auditing every new query for write skew, and decide where invariants are enforced across the whole platform.

## The idea Strict two-phase locking achieves serializability pessimistically: readers block writers, writers block readers, and correctness costs latency. Snapshot isolation achieves scalability by never blocking reads, at the cost of correctness — it validates only write sets. Serializable Snapshot Isolation is the optimistic middle: keep MVCC snapshot reads exactly as they are, *observe* the read-write conflicts the snapshot rule ignores, and abort transactions when the observed pattern could not have come from any serial order. ## The theorem it rests on Dependencies between concurrent transactions come in three forms. Write-read: B reads what A wrote. Write-write: B overwrites what A wrote. Read-write, called an **anti-dependency**: A read an item, and B then wrote a newer version — so any equivalent serial order must place A before B. Serializability means this dependency graph is acyclic. Snapshot isolation structurally rules out cycles that would need a write-read or write-write edge between concurrent transactions: a transaction never reads a concurrent write, and concurrent writes to one item cannot both commit. So any cycle that survives must be built with anti-dependency edges — and Fekete et al. proved something sharper: **every non-serializable snapshot-isolation history contains a cycle with two consecutive anti-dependency edges**. There is a transaction in the middle — the **pivot** — that has an incoming anti-dependency (someone read what it wrote) and an outgoing one (it read something someone else wrote). This is what makes runtime enforcement cheap. SSI never materialises the whole graph or searches it for cycles. It maintains two bits per transaction — "has an inbound conflict", "has an outbound conflict" — and fires when both are set on the same transaction and the commit ordering makes the situation unsafe. ## Tracking reads To notice an anti-dependency, the engine must remember what each transaction read. It does that with non-blocking read markers — often called SIREAD locks, though they are not locks in the blocking sense: they conflict with nothing and never make anyone wait; they are bookkeeping. They are recorded at the row level for point reads, and at index-range or page level for scans, so that a *predicate* read ("all reservations for room 5 today") can be matched against a later insert into that range. Markers must persist past the reading transaction's own commit, because the write that conflicts with a read may arrive later and the pivot may only be revealed then. ## Why it aborts innocent transactions Three independent sources of false positives, and being able to name them is the mark of real understanding: 1. **The dangerous structure is necessary, not sufficient.** A pivot with both edges *can* occur in perfectly serializable executions. SSI aborts on the pattern rather than proving a cycle exists, because proving it would require the full graph and, worse, information about transactions that have not finished. 2. **Read tracking is coarse.** A range read is registered against index pages or ranges, so a write anywhere in that range counts as a conflict even if it touches no row the reader would have returned. Fine granularity costs memory; coarse granularity costs precision. 3. **Memory pressure degrades precision further.** The tracking structures are bounded; when they fill, the engine summarises per-transaction read sets into coarser aggregates and, in the extreme, treats old committed transactions conservatively. Long-running transactions are the usual cause, because they keep read markers and old snapshots alive. The consequence is that abort rate is not purely a function of true conflicts; it is a function of true conflicts *plus* transaction duration *plus* read-set breadth. ## Operating it **Retry is not optional.** SSI converts anomalies into transient serialization failures, and a system without a bounded retry with exponential backoff and jitter converts them into user-visible errors. The retry must wrap the entire unit of work so it re-reads under a fresh snapshot, must be capped, and must be able to distinguish serialization failures from constraint or domain errors, which must never be retried. **Side effects must be outside.** Anything not rolled back by the database — emails, payment calls, message publishes — cannot live inside a transaction that may be aborted and replayed. Move them after commit, ideally via an outbox. **Transactions must be short and narrow.** Every extra millisecond and every extra row read widens the window and the read set. Analytical scans inside serializable transactions are the fastest way to make abort rates unmanageable; route them to a read-only path or a replica. **Mark read-only work as read-only.** Engines can exempt read-only transactions from much of the tracking, and a deferrable read-only mode can wait for a snapshot guaranteed safe, avoiding aborts entirely for reporting queries. **Measure before and after.** The decision hinges on the actual abort rate at production concurrency, the added latency of the retry tail, and the CPU cost of tracking. A conflict rate of a fraction of a percent is comfortable. Double-digit abort rates mean retries dominate throughput, and the answer is to reduce contention structurally — partition hot keys, shorten transactions, move counters out of the critical path — not to widen the retry budget. ## When not to use it If only a handful of code paths carry cross-row invariants, it can be cheaper and more predictable to leave the system on snapshot isolation and harden those paths individually with constraints, materialized conflicts, or explicit locking. SSI is the right default when invariants are numerous, ad hoc, or written by many teams — when you would rather pay a uniform tax than audit every new query for write skew.

  • Why does SSI abort transactions that were actually serializable?
    Its test is the dangerous structure — a transaction with both an incoming and an outgoing anti-dependency edge — which is necessary for a non-serializable history but not sufficient. Proving an actual cycle would require the full dependency graph plus knowledge of transactions still in flight. On top of that, read tracking is recorded at index-range or page granularity and is summarised more coarsely under memory pressure, so some recorded conflicts do not correspond to real ones.
  • How do long-running transactions specifically hurt an SSI deployment?
    They hold their snapshot and their read markers alive, so the tracking structures grow and old row versions cannot be cleaned up. When the tracking memory fills, the engine falls back to coarser summaries, which increases false-positive aborts across the whole system — not just for the long transaction. A single reporting query inside a serializable transaction can therefore raise abort rates for unrelated OLTP traffic.
  • How would you decide between running everything under SSI versus hardening a few code paths under snapshot isolation?
    It depends on how many invariants span multiple rows and who writes them. If the risky paths are few and stable, constraints, materialized conflicts, or explicit locking on those paths give predictable latency with no global abort tax. If invariants are numerous, ad hoc, or authored by many teams, a uniform serializable default is cheaper than auditing every new query for write skew — provided you measure the abort rate and can absorb the retry tail.

Instead of stopping every car at an intersection, you record which cars passed which junction and pull over anyone whose path forms the one traffic pattern known to be present in every collision — safe, cheap, and it occasionally stops a driver who would not have crashed.

saying these in an interview costs you the question

  • Saying SSI builds and searches the full dependency graph for cycles
  • Describing SIREAD read markers as locks that block writers
  • Claiming SSI never aborts a serializable transaction
  • Deploying serializable isolation without a bounded retry with backoff
  • Treating abort rate as fixed rather than driven by transaction length and read-set breadth
  • Leaving non-idempotent side effects such as emails or payment calls inside a retryable transaction

context