skip to content

Consistency Models

The spectrum of guarantees between strong and eventual: linearizability, causal and session guarantees, convergence, quorums, conflict resolution, CRDTs and reconciliation. Interviewers use this area to find out whether eventually consistent means anything precise to you.

part ofDistributed & scalable systemsoverview, primer and where to startread it →
on this pageshow

questions

page 2 of 2

What is a 'sloppy quorum' in a quorum-replicated store like Amazon Dynamo or Cassandra, how does it differ from a strict quorum, and what consistency guarantee do you give up by using one?

level: middleimportance: should knowfreq 55%

basics

~20 s

A sloppy quorum lets the system write to any healthy nodes it can reach instead of insisting on the exact nodes normally responsible for that piece of data. It keeps writes working during outages, but it means you can no longer be sure a later read will overlap with that write.

open as a page

In the COPS system for geo-replicated causally consistent storage, what does the '+' in 'causal+ consistency' add on top of basic causal consistency, and why did COPS need a separate mechanism (COPS-GT) for read-only transactions across multiple keys?

level: seniorimportance: should knowfreq 30%

basics

~20 s

The '+' adds a rule for what happens when two writes conflict (weren't ordered by causality) — everyone must resolve them the same way, so replicas converge. COPS-GT is a special read path that grabs several keys' values together, so a client reading multiple related items never sees a causally inconsistent mix, like an old version of one item next to a newer, dependent version of another.

open as a page

In a Dynamo-style system that combines hinted handoff with 'sloppy quorums' (accepting writes on any N reachable nodes rather than strictly the designated replica set), what durability and convergence risk does this combination introduce during an extended network partition?

level: seniorimportance: should knowfreq 35%

basics

~20 s

If the 'right' replicas are unreachable, the system just writes to other, reachable machines instead and promises to forward the data later. If the outage lasts too long, that promise can expire before it's kept, and the write can effectively vanish.

open as a page

When two replicas run anti-entropy repair using Merkle trees, why is comparing hash trees dramatically cheaper than comparing the raw datasets directly, and what has to happen when a mismatch is found?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Instead of comparing every single record, each replica builds a tree of checksums summarizing chunks of its data. Comparing just the top checksums tells you instantly if anything differs at all, and you only dig deeper into the branches that don't match.

open as a page

What is strict serializability, and why is it considered strictly stronger than either serializability or linearizability alone?

level: seniorimportance: should knowfreq 50%

basics

~10 s

Strict serializability means transactions behave both as if run one at a time AND in the real order they actually happened. It's serializability plus the real-time guarantee that linearizability adds.

open as a page

A social app is built as several microservices — posts, comments, notifications — each backed by its own eventually consistent store. A user reads a post, then writes a comment on it, expecting that anyone who can see their comment can also see the post it refers to. What has to happen at the infrastructure level to actually guarantee this writes-follow-reads property across service boundaries, and where does it typically break?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Every service in the chain needs to pass along a marker, like a version or vector clock entry, of what the user has already read, and the comment service must attach that marker to the write so downstream readers of the comment are forced to also be caught up on the post before they see it.

open as a page

Beyond Merkle-tree anti-entropy, some systems reconcile replicas using log-based mechanisms — for example shipping a replication log, or Kafka-style log compaction that retains only the latest value per key. Compare log-based reconciliation to Merkle-tree anti-entropy: what problem does each solve best, and what does each cost?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Log-based reconciliation replays a stream of changes in order to catch a replica up or keep only the newest value per key; Merkle trees compare two already-diverged snapshots to find differences after the fact. Logs are great for continuous sync, Merkle trees are the safety net for drift the log itself might have silently missed.

open as a page

In a leaderless replicated store using vector-clock-based conflict detection, deletes are often implemented as tombstone writes rather than physically removing data. Explain why plain deletes are dangerous under concurrent conflict resolution, and what problems tombstones introduce over time that operators must manage.

level: principalimportance: should knowfreq 35%

basics

~20 s

If you just erase a deleted record, a slightly-late update from before the delete can make it reappear, because there's nothing left to compare against. So systems keep a 'this was deleted' marker (a tombstone) around for a while instead of erasing right away — but keeping markers around forever wastes space, so they eventually get cleaned up too.

open as a page

You're designing a multi-region e-commerce platform and need to decide where different features - shopping cart, inventory count, order history - sit on the consistency spectrum from linearizable down to eventual. Walk through how you'd choose, and name a real system that lets you tune this per operation.

level: principalimportance: should knowfreq 60%

basics

~20 s

Different parts of an app need different guarantees - showing 'item in stock' can be a little stale, but charging a card exactly once needs strong guarantees. Good systems let you pick the guarantee per operation instead of one setting for everything.

open as a page

Under what precise conditions is convergence actually guaranteed in an eventually consistent system, and how do CRDTs provide a stronger guarantee ('strong eventual consistency') than a generic last-write-wins scheme?

level: principalimportance: should knowfreq 30%

basics

~20 s

Convergence isn't automatic just because updates eventually arrive everywhere — the merge rule also has to give the same result no matter what order updates show up in. CRDTs are specially designed data types built so that's always mathematically true; simpler rules like 'newest timestamp wins' can quietly lose data instead.

open as a page

When operating CRDTs at scale in production, what are the concrete costs (metadata growth, tombstone accumulation, causal stability tracking) and in what situations should you avoid choosing a CRDT for a piece of state at all?

level: principalimportance: should knowfreq 35%

basics

~20 s

CRDTs make merging automatic, but the bookkeeping (unique tags, deleted-item markers, per-replica counters) keeps growing and someone has to periodically clean it up safely, which is tricky. And CRDTs are the wrong tool whenever you need a strict rule across the whole system at once, like 'account balance can never go negative,' because no replica can enforce that alone without talking to the others first.

open as a page

During a network partition that leaves fewer than R (or fewer than W) replicas reachable from a client's coordinator node, what happens to quorum reads and writes in a strict-quorum system, and how do techniques like sloppy quorums, hinted handoff, and read repair change that behavior? Also: name a case where R + W > N is satisfied yet a client can still observe stale or conflicting data.

level: principalimportance: should knowfreq 45%

basics

~20 s

If not enough copies are reachable, a strict system just refuses the read or write rather than risk giving a wrong answer — it picks correctness over always answering. Looser variants keep answering by writing to backup nodes and fixing things up later, but that reopens the door to occasionally seeing old or conflicting data.

open as a page

Sequential consistency guarantees a single global order of operations that all processes agree on, but unlike linearizability it doesn't require that order to respect real-time (wall-clock) ordering across processes. Why would a system designer choose sequential consistency over linearizability, and what real-time anomaly does that trade-off allow?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Sequential consistency is theoretically cheaper than linearizability because it doesn't need to match real clock time, just one consistent story everyone agrees on. The catch: an operation that really finished earlier in time might still appear to be seen after a later one starts, as long as everyone sees the same story.

open as a page

How do production consensus-based stores like etcd or ZooKeeper provide linearizable reads without paying the full write-path (quorum round-trip) latency on every single read?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Instead of the slow write process for every read, the leader proves cheaply it's still the real leader -- one check-in with other nodes (read-index), or a time-limited promise nobody took over (lease) -- then answers locally.

open as a page

A globally distributed product wants to offer read-your-writes and monotonic reads to every user via sticky session routing to a 'home region' replica, plus nightly Merkle-tree anti-entropy across all regions. What does this design cost at scale, and under what circumstances would you deliberately relax or drop these guarantees for parts of the product?

level: principalimportance: nice to knowfreq 20%

basics

~20 s

Pinning every user to one region for consistency hurts latency for traveling users and makes failover harder, and comparing huge datasets nightly is expensive CPU and network work — so smart teams only pay for these guarantees on the few features where users would actually notice a violation, like their own posts, and skip them for things like view counters where nobody cares.

open as a page

showing 31–45 of 45