skip to content

How does linearizability differ from serializability, and how can a database be serializable but not linearizable, or linearizable but not serializable?

level: middleimportance: must knowfreq 85%

answer

  1. serializability: transactions, any equivalent serial order
  2. linearizability: one object, real-time order
  3. orthogonal properties, not a spectrum
  4. strict serializability = both combined
  5. Spanner's TrueTime buys the real-time part

basics

~20 s

Serializability makes whole transactions touching many pieces of data behave as if run one at a time, but ignores real-world timing. Linearizability keeps one piece of data instantly up to date in real time, but says nothing about multi-item transactions.

solid answer

~50 s

Serializability is a transaction-isolation guarantee: the outcome of executing concurrent, possibly multi-object transactions must equal some serial, one-at-a-time execution of those transactions, but that equivalent order need not match real/wall-clock time. Linearizability is a real-time, single-object guarantee: operations on one object must appear to happen atomically at a point between invocation and completion, respecting the actual real-time order in which they occurred. The two are orthogonal. A system can be serializable yet not linearizable -- for example, a snapshot-isolation-based database can pick a valid serial order for transactions that doesn't match wall-clock order, so a transaction that reads after another committed in real time might still miss its effect. Conversely, a system can be linearizable per key, like a plain distributed key-value store, without offering any multi-key transactional isolation at all. Combining both properties yields 'strict serializability': transactions ordered as if run serially and consistent with real time.

go deeper

for a junior

Should know serializability is about multi-item transactions and linearizability is about one item, without needing the formal real-time definition.

for a middle

Should be able to state that serializability doesn't fix real-time order while linearizability does, and give a rough example of each existing without the other.

for a senior

Should name strict serializability as the combination, explain why it costs more (real-time coordination), and connect it to a concrete mechanism like a bounded clock or single commit authority.

for a principal

Should be able to advise on when a system genuinely needs strict serializability versus when plain serializability (or even weaker isolation) is sufficient, weighing the latency and availability cost against the correctness the extra guarantee buys, and reference a real system's approach (e.g. TrueTime, consensus-based commit ordering).

## What each one orders Serializability and linearizability answer two different questions, and it's the difference in what they order that trips people up. - **Serializability asks:** given a set of concurrent transactions, each potentially reading and writing many objects, is there some serial order — run transaction 1 completely, then transaction 2, then transaction 3 — that would have produced the same final state and the same values each transaction observed? If yes, the execution is serializable, however the operations were actually interleaved in real time. Crucially, that equivalent serial order is chosen for logical consistency, not for timing: two transactions `T1` and `T2` might genuinely execute with `T1` finishing in wall-clock time before `T2` even starts, yet a serializable system is free to treat them "as if" `T2` ran first, provided that ordering is internally consistent with what each transaction actually read and wrote. - **Linearizability, in contrast,** is defined per single object and explicitly pinned to real time: it requires that the apparent order of operations on that object match the order they actually occurred in wall-clock time, with no freedom to reshuffle. | Guarantee | Serializability | Linearizability | |---|---|---| | What it orders | transactions across many objects | operations on one object | | Order it must match | any equivalent serial order | real wall-clock time | | Where it comes from | the transaction-processing world | the distributed-systems and concurrent-programming world | ## Why the two exist This difference exists because the two guarantees were built to solve different problems. - **Serializability** comes from the transaction-processing world, where the goal is protecting invariants across multiple pieces of data touched together (e.g., debit one account and credit another must never be observed half-done). Its classical implementations, **two-phase locking** or **serializable snapshot isolation**, only need to reconstruct a logically consistent story after the fact; they were never designed with real-time obligations in mind, because a single-node database has no meaningful cross-node clock-skew problem to solve. - **Linearizability** comes from the distributed-systems and concurrent-programming world, where the goal is giving a replicated object the illusion of being a single, un-replicated copy that any client can trust the instant an operation completes, no matter which physical replica or datacenter answered. ## The trade-off **The trade-off:** achieving serializability alone is comparatively cheap. A single-node database, or even a replicated one using asynchronous or snapshot-based techniques, can serialize transactions using local logical clocks, locks, or MVCC snapshot ordering without any need to reason about physical wall-clock time or cross-node coordination latency. Adding the linearizability requirement on top — producing **strict serializability** — is what forces real coordination cost: the system must additionally guarantee that its chosen serial order tracks real time, which typically means: 1. synchronizing commit timestamps against a tightly bounded physical clock; 2. or funneling commits through a single ordering authority. Google Spanner is the textbook example: it achieves strict serializability globally by using `TrueTime`, a clock API with a bounded uncertainty interval, and a "commit wait" step that delays making a transaction visible until the uncertainty window has definitely passed, purely to buy the real-time guarantee that plain serializability doesn't provide. ## The failure mode The failure mode from confusing the two is a class of production bugs sometimes called **"stale read after commit"** anomalies: an engineer assumes that because the database is "serializable" (a term often used loosely as a synonym for "safe"), a transaction that reads after another one committed in real time will always see its effect. That's only guaranteed under strict serializability, not plain serializability. A concrete scenario: 1. a payment service commits a balance transfer on a primary; 2. an operator immediately calls a support agent who queries a read replica or a differently-ordered secondary transaction; 3. because the isolation level is "merely" serializable rather than strictly serializable, the query is permitted to reflect a serial order where its own transaction is treated as happening before the transfer, even though the transfer genuinely completed first in wall-clock time. The system hasn't violated its documented guarantee, but it has violated the engineer's mental model. ## Where each one shows up In practice, most operational databases only promise serializability — PostgreSQL's `SERIALIZABLE` isolation, for instance, is a single-node guarantee with no real-time cross-transaction promise beyond what a single clock already gives for free. The distributed systems that explicitly market **"external consistency"** or **"strict serializability"** — Spanner, CockroachDB, FoundationDB — are the ones that have paid extra engineering and latency cost specifically to close the real-time gap that plain serializability leaves open.

  • If a system is linearizable on every individual key, is it automatically serializable for transactions spanning multiple keys?
    No. Per-key linearizability says nothing about atomicity across keys, so a transaction touching two linearizable keys can still suffer classic multi-object anomalies like lost updates or write skew, because nothing prevents another transaction from interleaving between the two individual key operations.
  • Why don't more databases just implement strict serializability by default if it's the strongest guarantee?
    Because the real-time coordination it requires costs latency and availability -- every commit needs a globally trustworthy ordering mechanism such as a bounded clock or a consensus round-trip, which is unnecessary overhead for workloads that don't actually need cross-node real-time guarantees.
  • What isolation level do most production relational databases default to, and how does that relate to serializability?
    Most default to Read Committed or Repeatable Read / Snapshot Isolation, which are weaker than full serializability and can permit anomalies like write skew; true SERIALIZABLE isolation is usually opt-in because it carries extra locking or validation overhead.

Serializability is like a movie editor who can freely reorder scenes in the final cut as long as the plot still makes sense; linearizability is like a live news broadcast that must show events in the exact order they actually happened. Strict serializability is a movie that both makes narrative sense and is cut in the exact chronological order the scenes were filmed.

saying these in an interview costs you the question

  • Says linearizability and serializability are just two names for the same thing
  • Claims serializability guarantees real-time (wall-clock) ordering
  • Thinks per-key linearizability automatically gives multi-key transactional atomicity
  • Cannot name what strict serializability adds on top of plain serializability
  • Assumes any 'serializable' database is safe to reason about with real-time intuition

context