skip to content

In a distributed key-value store, what does 'causal consistency' guarantee about the order in which different replicas observe writes, and how does that differ from plain eventual consistency?

level: juniorimportance: must knowfreq 65%

answer

  1. happens-before relation
  2. cause visible before effect, always
  3. concurrent writes = no order guarantee
  4. session consistency in Cosmos DB
  5. dependency tracking cost

basics

~20 s

Causal consistency guarantees that writes which are causally related (one happened because of, or after seeing, another) are seen by every replica in that same order. Unrelated writes may appear in different orders on different replicas. Plain eventual consistency promises no ordering at all, just eventual agreement.

solid answer

~40 s

Causal consistency is a model between eventual and strong (linearizable) consistency. If write A 'happens-before' write B — because the same client wrote A then B, or a client read A then wrote B, or transitively through a chain of such relationships — every replica must apply A before B. Writes with no such relationship (concurrent writes) can be observed in different orders on different replicas, converging eventually via a conflict-resolution rule. Plain eventual consistency only guarantees replicas converge to the same final state if writes stop; it says nothing about the order a client sees writes arrive in, so a reply could be visible before the comment it replies to. Causal consistency rules that out cheaply, without the global coordination strong consistency needs.

go deeper

for a junior

Should state the basic guarantee in plain language: causes are seen before effects, and give one concrete anomaly example (like a reply appearing before its post) that eventual consistency alone permits.

for a middle

Should articulate the happens-before relation precisely (program order, read-from, transitivity) and be able to say what happens to concurrent writes under causal consistency.

for a senior

Should discuss the cost side: dependency tracking overhead, causal stalling, and be able to compare causal consistency's guarantees against a concrete system (e.g., session consistency in Cosmos DB) and explain when it's insufficient.

for a principal

Should place causal consistency in the broader consistency spectrum and articulate the deeper reason it's a sweet spot (availability trade-offs), and know how to choose it vs. stronger/weaker models at a system-design level for a given product requirement.

## What the model guarantees **Causal consistency** is a consistency model for replicated data stores that sits strictly between **eventual consistency** (the weakest useful guarantee) and **strong/linearizable consistency** (the strongest, and most expensive). Its core idea is the `happens-before` relation, borrowed from Lamport's work on distributed event ordering. An operation A happens-before an operation B if: - **(1) Program order** — the same client performed A and then B in program order. - **(2) Read-from** — B is a read that observed the value written by A (a read-from relationship). - **(3) Transitive** — the relation is transitive: A happens-before C if A happens-before B and B happens-before C. Two operations related by happens-before are **causally related**; two operations related in neither direction are **concurrent**. Causal consistency's guarantee: every replica must apply causally related writes in the order happens-before dictates — no replica ever observes an effect before its cause. Concurrent writes carry no such obligation — different replicas can apply them in different orders, as long as they eventually converge to the same value via a deterministic conflict-resolution rule (e.g., `last-writer-wins` or an application merge function). ## Why the model exists This model exists because plain eventual consistency is too weak to build usable applications on. Its only promise is: if writes stop arriving, every replica eventually holds the same data. It says nothing about the order a client sees writes arrive in along the way. That opens the door to real anomalies: - a user could load a comment thread and see a reply ('Great point!') rendered above the comment it replies to, because the reply happened to propagate to that replica first; - a chat client could show 'yes' before the question it answers. These are exactly the class of bug unstructured eventually-consistent systems produce when they replicate independently per key. Causal consistency closes this gap cheaply: it only tracks and respects dependencies that are actually causally linked, not a single global order on all operations, avoiding the expensive coordination (consensus, locking, a single serialization point) strong consistency needs. ## The trade-off The trade-off is real on both sides. - In exchange for a far more intuitive **programming model** — 'you never see an effect before its cause' — the system gives up any agreed ordering for concurrent writes. - Two users editing the same document at once from different regions can each see their own edit applied first, then reconciled later; an application that cannot tolerate that (e.g., a bank ledger where every debit must be globally ordered against every credit) cannot rely on causal consistency alone. - On the systems side it isn't free either: preserving happens-before requires tracking, for every write, which prior writes it causally depends on, and withholding a write from visibility on a replica until all its dependencies are already visible there — that bookkeeping and enforcement both cost metadata and sometimes latency. ## The failure mode in production The most common production failure mode is a symptom of that bookkeeping: a system that cannot yet prove a write's dependencies have arrived at a given replica must either delay exposing the write there (increasing tail latency, sometimes called **causal stalling**) or risk violating the guarantee. Under network partitions or straggler replicas, dependency chains can pile up and cause visible writes to lag behind what a naive eventually-consistent system would have shown immediately — engineers sometimes mistake this for a bug rather than the cost of the guarantee. ## Where it shows up - **Azure Cosmos DB's 'Session' consistency level** is a well-known concrete example, a practical, client-scoped variant of causal consistency: within a single client session, reads are guaranteed to reflect that session's own prior writes and move monotonically forward, at far lower latency and availability cost than the database's 'Strong' consistency level. - Academic systems like **COPS** demonstrated a full, cluster-wide causal consistency model could be built with modest overhead while remaining available during partitions, which is why causal consistency is often cited as the strongest model achievable without sacrificing availability.

  • If causal consistency is stronger than eventual consistency, why doesn't every distributed database just use it by default?
    Because tracking causal dependencies and enforcing 'dependency before visibility' costs metadata and can add latency when a replica must wait for a write's dependencies to arrive before exposing it. For workloads that genuinely don't care about ordering (e.g., independent counters or unrelated keys), plain eventual consistency is cheaper and sufficient, so many systems default to it and offer causal consistency as an opt-in stronger mode.
  • How does a client typically observe a causal-consistency violation in practice?
    A classic symptom is a 'reply before post' or 'answer before question' anomaly — a client reads a value derived from an earlier write without ever seeing that earlier write, because the two writes propagated to that replica out of order. Another is a client re-reading its own prior write and getting a stale value, which violates the read-your-writes session guarantee causal consistency is meant to provide.

Like a group text thread where replies always show up after the message they're replying to, but two people's unrelated jokes posted at the same moment might land in a different order for different readers — nobody is confused about who's replying to what, but simultaneous unrelated chatter can interleave differently.

saying these in an interview costs you the question

  • Says causal consistency means all replicas apply every write in the same global order (that's strong/sequential consistency, not causal).
  • Thinks causal consistency guarantees linearizable reads.
  • Believes eventual consistency already implies causal ordering.
  • Cannot explain what 'happens-before' means or give an example of two concurrent operations.

context