Consistency Models
The spectrum of guarantees between strong and eventual: linearizability, causal and session guarantees, convergence, quorums, conflict resolution, CRDTs and reconciliation. Interviewers use this area to find out whether eventually consistent means anything precise to you.
part ofDistributed & scalable systemsoverview, primer and where to startread it →on this pageshowhide
explore
- Read/Write Quorums5 questions
- Reconciliation & Session Guarantees6 questions
- Consistency Model Spectrum5 questions
- Convergence & Propagation6 questions
- Conflict Resolution6 questions
- CRDTs6 questions
- Linearizability vs Serializability6 questions
- Causal Consistency5 questions
- Backend Developerroleanchors this topic
- Full Stack Developerroleanchors this topic
- Java Backend Developerroleanchors this topic
- Kotlin Backend Developerroleanchors this topic
- System Designskillanchors this topic
- API Designskill
- Data Engineerrole
- DevOps / SRE Engineerrole
- Forward Deployed Engineerrole
- Game Developerrole
- PostgreSQL DBArole
- Server-Side Game Developerrole
- Software Architectrole
- Software Design & Architectureskill
questions
page 2 of 2What is a 'sloppy quorum' in a quorum-replicated store like Amazon Dynamo or Cassandra, how does it differ from a strict quorum, and what consistency guarantee do you give up by using one?
basics
~20 sA sloppy quorum lets the system write to any healthy nodes it can reach instead of insisting on the exact nodes normally responsible for that piece of data. It keeps writes working during outages, but it means you can no longer be sure a later read will overlap with that write.
In the COPS system for geo-replicated causally consistent storage, what does the '+' in 'causal+ consistency' add on top of basic causal consistency, and why did COPS need a separate mechanism (COPS-GT) for read-only transactions across multiple keys?
basics
~20 sThe '+' adds a rule for what happens when two writes conflict (weren't ordered by causality) — everyone must resolve them the same way, so replicas converge. COPS-GT is a special read path that grabs several keys' values together, so a client reading multiple related items never sees a causally inconsistent mix, like an old version of one item next to a newer, dependent version of another.
In a Dynamo-style system that combines hinted handoff with 'sloppy quorums' (accepting writes on any N reachable nodes rather than strictly the designated replica set), what durability and convergence risk does this combination introduce during an extended network partition?
basics
~20 sIf the 'right' replicas are unreachable, the system just writes to other, reachable machines instead and promises to forward the data later. If the outage lasts too long, that promise can expire before it's kept, and the write can effectively vanish.
When two replicas run anti-entropy repair using Merkle trees, why is comparing hash trees dramatically cheaper than comparing the raw datasets directly, and what has to happen when a mismatch is found?
basics
~20 sInstead of comparing every single record, each replica builds a tree of checksums summarizing chunks of its data. Comparing just the top checksums tells you instantly if anything differs at all, and you only dig deeper into the branches that don't match.
What is strict serializability, and why is it considered strictly stronger than either serializability or linearizability alone?
basics
~10 sStrict serializability means transactions behave both as if run one at a time AND in the real order they actually happened. It's serializability plus the real-time guarantee that linearizability adds.
A social app is built as several microservices — posts, comments, notifications — each backed by its own eventually consistent store. A user reads a post, then writes a comment on it, expecting that anyone who can see their comment can also see the post it refers to. What has to happen at the infrastructure level to actually guarantee this writes-follow-reads property across service boundaries, and where does it typically break?
basics
~20 sEvery service in the chain needs to pass along a marker, like a version or vector clock entry, of what the user has already read, and the comment service must attach that marker to the write so downstream readers of the comment are forced to also be caught up on the post before they see it.
Beyond Merkle-tree anti-entropy, some systems reconcile replicas using log-based mechanisms — for example shipping a replication log, or Kafka-style log compaction that retains only the latest value per key. Compare log-based reconciliation to Merkle-tree anti-entropy: what problem does each solve best, and what does each cost?
basics
~20 sLog-based reconciliation replays a stream of changes in order to catch a replica up or keep only the newest value per key; Merkle trees compare two already-diverged snapshots to find differences after the fact. Logs are great for continuous sync, Merkle trees are the safety net for drift the log itself might have silently missed.
In a leaderless replicated store using vector-clock-based conflict detection, deletes are often implemented as tombstone writes rather than physically removing data. Explain why plain deletes are dangerous under concurrent conflict resolution, and what problems tombstones introduce over time that operators must manage.
basics
~20 sIf you just erase a deleted record, a slightly-late update from before the delete can make it reappear, because there's nothing left to compare against. So systems keep a 'this was deleted' marker (a tombstone) around for a while instead of erasing right away — but keeping markers around forever wastes space, so they eventually get cleaned up too.
You're designing a multi-region e-commerce platform and need to decide where different features - shopping cart, inventory count, order history - sit on the consistency spectrum from linearizable down to eventual. Walk through how you'd choose, and name a real system that lets you tune this per operation.
basics
~20 sDifferent parts of an app need different guarantees - showing 'item in stock' can be a little stale, but charging a card exactly once needs strong guarantees. Good systems let you pick the guarantee per operation instead of one setting for everything.
Under what precise conditions is convergence actually guaranteed in an eventually consistent system, and how do CRDTs provide a stronger guarantee ('strong eventual consistency') than a generic last-write-wins scheme?
basics
~20 sConvergence isn't automatic just because updates eventually arrive everywhere — the merge rule also has to give the same result no matter what order updates show up in. CRDTs are specially designed data types built so that's always mathematically true; simpler rules like 'newest timestamp wins' can quietly lose data instead.
When operating CRDTs at scale in production, what are the concrete costs (metadata growth, tombstone accumulation, causal stability tracking) and in what situations should you avoid choosing a CRDT for a piece of state at all?
basics
~20 sCRDTs make merging automatic, but the bookkeeping (unique tags, deleted-item markers, per-replica counters) keeps growing and someone has to periodically clean it up safely, which is tricky. And CRDTs are the wrong tool whenever you need a strict rule across the whole system at once, like 'account balance can never go negative,' because no replica can enforce that alone without talking to the others first.
During a network partition that leaves fewer than R (or fewer than W) replicas reachable from a client's coordinator node, what happens to quorum reads and writes in a strict-quorum system, and how do techniques like sloppy quorums, hinted handoff, and read repair change that behavior? Also: name a case where R + W > N is satisfied yet a client can still observe stale or conflicting data.
basics
~20 sIf not enough copies are reachable, a strict system just refuses the read or write rather than risk giving a wrong answer — it picks correctness over always answering. Looser variants keep answering by writing to backup nodes and fixing things up later, but that reopens the door to occasionally seeing old or conflicting data.
Sequential consistency guarantees a single global order of operations that all processes agree on, but unlike linearizability it doesn't require that order to respect real-time (wall-clock) ordering across processes. Why would a system designer choose sequential consistency over linearizability, and what real-time anomaly does that trade-off allow?
basics
~20 sSequential consistency is theoretically cheaper than linearizability because it doesn't need to match real clock time, just one consistent story everyone agrees on. The catch: an operation that really finished earlier in time might still appear to be seen after a later one starts, as long as everyone sees the same story.
How do production consensus-based stores like etcd or ZooKeeper provide linearizable reads without paying the full write-path (quorum round-trip) latency on every single read?
basics
~20 sInstead of the slow write process for every read, the leader proves cheaply it's still the real leader -- one check-in with other nodes (read-index), or a time-limited promise nobody took over (lease) -- then answers locally.
A globally distributed product wants to offer read-your-writes and monotonic reads to every user via sticky session routing to a 'home region' replica, plus nightly Merkle-tree anti-entropy across all regions. What does this design cost at scale, and under what circumstances would you deliberately relax or drop these guarantees for parts of the product?
basics
~20 sPinning every user to one region for consistency hurts latency for traveling users and makes failover harder, and comparing huge datasets nightly is expensive CPU and network work — so smart teams only pay for these guarantees on the few features where users would actually notice a violation, like their own posts, and skip them for things like view counters where nobody cares.
showing 31–45 of 45