skip to content

Why is causal consistency often cited as the strongest consistency model that a distributed data store can provide while still remaining available under network partitions, and what does a team give up if they choose a stronger model instead?

level: principalimportance: must knowfreq 35%

answer

  1. CAC result: Mahajan/Alvisi/Dahlin, formalized by Attiya/Bailis/Fekete
  2. dependency check needs only already-observed info
  3. stronger models need real-time agreement on unknown/concurrent ops
  4. causal-safe invariants = mergeable; total-order invariants (inventory, unique username) need more
  5. Bolt-on Causal Consistency shim over Cassandra

basics

~20 s

Because checking 'did this depend on something else' can be answered using only information the client and the local replica already have, without waiting on other machines. Anything stronger — like a single global order for all writes — needs machines to talk to each other and agree in real time, which becomes impossible if a network split cuts them off from each other.

solid answer

~50 s

This claim traces to results (notably by Mahajan, Alvisi, and Dahlin, later formalized by Attiya, Bailis, and Fekete) showing causal consistency is the strongest consistency model that can be implemented while always being available, low-latency, and partition-tolerant — because determining whether an operation's causal dependencies are satisfied only requires information already reachable via that operation's own causal history, which a replica can check locally. Any model stronger than causal — sequential consistency, linearizability, or anything requiring a real-time total order — needs replicas to coordinate about operations they haven't causally observed, and coordination is exactly what a partition can block. The trade-off a team accepts by choosing something stronger is CP-style behavior in CAP terms: during a partition, the system must block or error on some writes or reads rather than serve them from a possibly-stale replica, sacrificing availability for a global ordering causal consistency deliberately does not promise.

go deeper

for a junior

Not expected to know the theorem; can be asked, at a high level, why 'strongest possible while still always answering requests' sounds like a useful property for an always-on app.

for a middle

Should be able to connect causal consistency's availability to CAP at a basic level: it keeps working during partitions because it doesn't need cross-replica coordination for concurrent operations.

for a senior

Should be able to explain mechanically why dependency-checking only needs already-observed information, and identify at least one class of invariant (uniqueness, inventory) that causal consistency cannot enforce alone.

for a principal

Should be able to cite the shape of the theoretical result (even if not exact authors/dates), reason about which specific invariants in a real system need to be carved out for stronger coordination, and discuss engineering gaps between the theoretical ceiling and a real deployment (e.g., causal-stalling).

## The result the claim rests on The claim rests on a specific theoretical result, developed initially by **Mahajan, Alvisi, and Dahlin** and later given a rigorous formal treatment by **Attiya, Bailis, and Fekete** (often abbreviated as the **CAC** — causal consistency + availability — line of work): among the standard hierarchy of consistency models, causal consistency is the strongest one that can be implemented by a system that is - **always available** — every non-faulty replica answers every request; - **low-latency** — answers don't wait on remote coordination; - **tolerant of network partitions**. The mechanical reason is almost definitional once you look at what causal consistency actually requires a replica to check: to safely apply and expose operation B, the replica only needs to confirm that B's own causal dependencies — the writes the issuing client actually read or wrote before performing B — are already present. Crucially, those dependencies are, by the definition of causality, things that already reached the requesting client through some prior interaction, and therefore have already had a chance to propagate through normal replication to any replica that will now serve B. No replica ever needs to synchronously ask a possibly-unreachable peer 'has anything else happened that I don't know about yet?' — it only needs to know about the specific, already-observed things B depends on. ## Why anything stronger cannot stay available Contrast that with any model strictly stronger than causal — sequential consistency or linearizability, for instance — which require establishing an agreed order among operations that may be entirely concurrent and mutually unknown to each other, including operations neither replica has causally observed yet. Determining that order requires the replicas involved to communicate about the fact of each other's existence in real time, precisely the kind of cross-replica coordination a network partition can block. In CAP terms, causal consistency sits at the boundary: - it's the strongest model on the **AP** side of the trade-off; - while anything stronger forces a system into **CP** behavior — during a partition, a CP system must refuse to serve some subset of requests (or serve them only from whichever side has quorum) rather than risk two sides of the partition disagreeing about a total order they can't currently coordinate on. ## What a team gives up by staying at causal The cost a team accepts by staying at causal consistency is that it provides no help at all for invariants that require a true, system-wide total order across operations that are not causally related. The textbook examples: - selling the last unit of inventory; - enforcing a globally unique username; - preventing a double-spend on an account balance. Two clients performing these operations concurrently, from different regions, are — by definition — concurrent, not causally related, so a causally consistent system is free to let both locally 'succeed' independently, leaving the application to reconcile an outcome that may be fundamentally irreconcilable (you can't un-sell a physical item already promised to two different people). For invariants like these, a team genuinely needs stronger coordination — consensus, a single-writer partition, or a synchronous quorum check — and has to accept the corresponding latency and availability cost, at least on that specific code path, even if the rest of the system happily runs on causal consistency. ## The failure modes A common production failure mode is a team mismatching the guarantee to the invariant in either direction: - assuming causal consistency silently prevents a race condition it was never designed to prevent, leading to double-allocation bugs under concurrent load; - or, more conservatively, defaulting an entire system to strong/linearizable consistency out of caution and paying a global coordination tax on every operation, including the vast majority that never actually needed a total order. There's also a subtler practical gap between the theorem and a given implementation: the CAC result establishes a ceiling on what's theoretically achievable, but a specific causally consistent system can still erode its own practical availability if it isn't engineered carefully — for example, if a replica blocks indefinitely waiting for a slow dependency to arrive (the **causal-stalling** failure mode) rather than bounding that wait, the deployed system behaves worse than the theorem's ideal. ## Where it shows up A concrete real-world illustration of teams actively exploiting this result is Bailis et al.'s **Bolt-on Causal Consistency** work, which retrofits causal consistency as a shim layer on top of existing, already-deployed eventually-consistent stores like **Cassandra** — explicitly motivated by wanting causal ordering's stronger guarantees without sacrificing the underlying store's availability, precisely because the theory says that trade isn't necessary.

  • If causal consistency is provably the strongest available-and-partition-tolerant model, why do systems like Spanner or CockroachDB offer strong/linearizable consistency at all?
    Because 'available under partition' isn't always the top priority — some invariants (unique constraints, financial transfers, inventory decrements) are only safe under real coordination, so those systems deliberately accept CP-style behavior (reduced availability or added latency during partitions or contention) in exchange for a total order strong enough to enforce those invariants correctly. The CAC result tells you the theoretical ceiling for staying available; it doesn't say every application should target that ceiling.
  • Can a causally consistent system still experience a request that hangs or times out during normal operation, not just during a partition?
    Yes — if a replica is waiting on a dependency that's slow to arrive (a straggler replica, a congested link, a large fan-out of dependencies) it can stall a read or write even with no partition present. This causal-stalling behavior is an engineering property of a specific implementation, not something the availability theorem rules out; the theorem is about what's achievable in principle, not a guarantee every implementation hits it.

It's like verifying a rumor by only checking sources you personally already heard it from — you never need to phone a stranger on the other side of town to confirm it, so a cut phone line to that stranger doesn't stop you. Confirming a strict global 'who said what first, across everyone' order, though, does require calling around, and a cut line stops that cold.

saying these in an interview costs you the question

  • Claims causal consistency is 'as strong as' linearizability, missing that it's strictly weaker by design.
  • Cannot explain mechanically why checking causal dependencies avoids needing to contact unreachable replicas.
  • Thinks causal consistency is sufficient for enforcing global uniqueness or inventory-limit invariants.
  • Treats the CAC theoretical result as a guarantee that any specific causally-consistent implementation is fully available in practice.

context