skip to content

In the COPS system for geo-replicated causally consistent storage, what does the '+' in 'causal+ consistency' add on top of basic causal consistency, and why did COPS need a separate mechanism (COPS-GT) for read-only transactions across multiple keys?

level: seniorimportance: should knowfreq 30%

answer

  1. causal+ = causal ordering + convergent conflict resolution
  2. concurrent writes to same key need deterministic merge (e.g. LWW)
  3. COPS-GT: bounded 2-round read-only transactions across keys
  4. motivating example: privacy setting vs. old post visible
  5. Eiger extends to read-write transactions

basics

~20 s

The '+' adds a rule for what happens when two writes conflict (weren't ordered by causality) — everyone must resolve them the same way, so replicas converge. COPS-GT is a special read path that grabs several keys' values together, so a client reading multiple related items never sees a causally inconsistent mix, like an old version of one item next to a newer, dependent version of another.

solid answer

~50 s

Basic causal consistency only orders causally related operations and leaves concurrent (conflicting) writes to the same key unordered, risking replicas permanently disagreeing on that key's value. Causal+ adds convergent conflict handling: a deterministic, commutative merge or arbitration rule (e.g., last-writer-wins by version) so all replicas resolve the same conflict to the same final value. COPS enforces this at the single-key put/get level cheaply using per-key dependency metadata. But a client reading several different keys in sequence with plain single-key gets could still observe a causally inconsistent snapshot across keys — e.g., see a friend's updated privacy setting but an old post that setting should have hidden. COPS-GT (get transactions) fixes this by fetching a client's requested keys together and checking, in roughly two rounds of communication, that the returned values are mutually causally consistent, retrying only the stale ones — giving read-only transaction guarantees without full multi-key locking.

go deeper

for a junior

Not generally expected to know COPS by name; can be prompted with the scenario and asked what could go wrong reading two related pieces of data separately.

for a middle

Should understand that concurrent writes to the same key need some agreed resolution rule, even if they can't name COPS specifically.

for a senior

Should know COPS by name, explain causal+'s convergent conflict handling, and explain the cross-key read consistency problem and why single-key causal consistency doesn't solve it.

for a principal

Should be able to explain the COPS-GT bounded-rounds design and its cost/benefit versus alternatives (full transactions, locking, application-level retries), and place it in context against successor work like Eiger.

## What the '+' adds Basic causal consistency, on its own, has a gap: it only imposes ordering on operations that are causally related. Two writes to the same key that are concurrent — neither happened-before the other — have no ordering rule at all, meaning different replicas could apply them in different orders and, if nothing else intervenes, permanently disagree about that key's final value. **Causal+ consistency**, the model COPS (Clusters of Order-Preserving Servers, Lloyd et al., 2011) implements, closes this gap by adding **convergent conflict handling**: whenever two concurrent writes touch the same key, every replica must resolve the conflict using the same deterministic rule — commonly `last-writer-wins` by version number, though the system can plug in an application-supplied merge function — so every replica converges to the identical final value, even though they may have observed the conflicting writes in different orders along the way. This is what makes it 'causal, plus something': - **causal ordering** for related operations; - **plus** a guaranteed, agreed-upon resolution for unrelated ones. ## How COPS implements it per key COPS implements this at the granularity of individual key-value puts and gets. Every stored value carries a small amount of dependency metadata identifying its **nearest** causal dependencies — the same nearest-dependency bookkeeping used generally to keep causal consistency's overhead bounded. On a `put`, the server: 1. checks the new write's dependencies are already satisfied locally before applying it (the **causal-cut** rule); 2. if it conflicts with another concurrent put to the same key, invokes the conflict handler to decide the surviving value deterministically. Because every replica runs the identical handler on the identical set of conflicting writes, they converge to the same result without synchronous coordination at write time — exactly what preserves COPS's core goal of remaining available and low-latency across geographically distant data centers. ## Why single-key operations are not the whole story Single-key operations, however, are not the whole story for real applications, which routinely need to read several related keys together. This is where the second half of the question comes in: even with every individual `get` returning a causally consistent value for its own key, a client issuing several independent gets in sequence or in parallel can still end up with a causally inconsistent set across keys. The paper's motivating example is a social network: a user tightens their privacy setting (one key) and then deletes an old post that setting was meant to hide (a causally dependent write on another key); a viewer who reads the post-list key slightly before the deletion propagates, but the ACL key slightly after the tightened setting propagates, ends up seeing exactly the content the user tried to hide — each individual key was internally fine, but the pair, read together, wasn't causally consistent. ## What COPS-GT adds, in two rounds **COPS-GT** (get transactions) exists specifically to close this cross-key gap for read-only access, without resorting to full distributed transactions or locking. At a high level it works in a bounded two rounds of communication regardless of how many keys are requested or how deep their dependency chains run: 1. the first round fetches all requested keys, in parallel, from the local cluster, along with each returned value's dependency information; 2. the second round checks whether any returned value is missing a dependency that another returned value in the same set requires, and if so, re-fetches only the offending key at a version that satisfies it. This bounded-rounds design is the key engineering result — a naive approach walking the full dependency graph could need round trips proportional to the depth of the causal chain, unacceptable for interactive read paths. ## The trade-off The trade-off is real: - COPS-GT adds an extra round trip's worth of latency compared to a plain, non-transactional multi-get; - it only covers read-only access — it does not give general multi-key read-write transactions or full ACID semantics, and concurrent writes to different keys within the same logical operation still aren't atomically ordered; - it also inherits causal consistency's broader operational costs: dependency metadata still needs background garbage collection, and the conflict handler must genuinely be deterministic and commutative or replicas can still diverge. COPS's ideas were later extended by the **Eiger** system to cover read-write transactions as well, distinct follow-on work beyond COPS-GT itself.

  • Why can't COPS just use a single global lock or coordinator to solve the multi-key read consistency problem?
    A global coordinator or lock would reintroduce the synchronous cross-datacenter coordination that causal consistency (and causal+ specifically) is designed to avoid, undermining the low-latency, always-available property that's the whole point of the model. COPS-GT instead solves it with a bounded, mostly-local check-and-retry protocol that avoids a single point of coordination.
  • What happens if the conflict handler used for causal+ convergence isn't actually deterministic across replicas?
    Replicas can permanently diverge on that key's value, because each one may compute a different 'winner' from the same set of conflicting concurrent writes — defeating the purpose of causal+'s convergent conflict handling. The handler must be a pure, deterministic function of the conflicting writes (or their metadata, like version numbers), not something that depends on local state like arrival order or wall-clock time at the replica.

Basic causal consistency is like making sure everyone agrees a reply comes after its original message; causal+ is adding a house rule for what happens when two people talk over each other at the same moment (whoever has the later timestamp wins, and everyone applies that same rule); COPS-GT is like grabbing a matched pair of documents from a filing cabinet and double-checking neither one is an outdated version relative to the other before handing them both over.

saying these in an interview costs you the question

  • Says causal+ means all writes are globally ordered, contradicting the point of the '+'.
  • Cannot explain why single-key causal consistency isn't sufficient for multi-key reads.
  • Thinks COPS-GT provides full multi-key read-write transactions.
  • Believes COPS-GT requires a number of round trips proportional to dependency depth rather than a bounded number.

context