skip to content

What is a 'sloppy quorum' in a quorum-replicated store like Amazon Dynamo or Cassandra, how does it differ from a strict quorum, and what consistency guarantee do you give up by using one?

level: middleimportance: should knowfreq 55%

answer

  1. preference list vs any healthy node
  2. hint = intended owner metadata
  3. hinted handoff forwards later
  4. breaks R+W>N overlap
  5. Dynamo shopping cart example

basics

~20 s

A sloppy quorum lets the system write to any healthy nodes it can reach instead of insisting on the exact nodes normally responsible for that piece of data. It keeps writes working during outages, but it means you can no longer be sure a later read will overlap with that write.

solid answer

~50 s

A strict quorum requires W acknowledgments specifically from the N nodes designated as owners/replicas of a key — if enough of those specific nodes are unreachable, the write fails outright. A sloppy quorum relaxes this: if fewer than W of the designated replicas are reachable, the coordinator writes the data to other healthy nodes instead, temporarily 'borrowing' capacity outside the normal replica set, and marks those writes with a hint recording the intended owner. Once the real owner comes back, the hint is forwarded to it and the temporary copy is discarded — this is hinted handoff. The trade-off: because the write may have landed on nodes outside the designated set, a subsequent strict quorum read of the designated replicas is no longer guaranteed to overlap with it, so R+W>N stops being a hard guarantee — you gain availability during partitions/outages at the cost of that overlap guarantee, moving further toward eventual (rather than immediately-quorum-consistent) reads.

go deeper

for a junior

Should know that a sloppy quorum means writes can go to non-designated nodes to survive an outage, and that this is a deliberate trade for availability.

for a middle

Should be able to explain hinted handoff mechanically (hint metadata, forwarding on recovery) and state clearly that the R+W>N overlap guarantee no longer holds during the window.

for a senior

Should connect this to CAP/PACELC explicitly, discuss read repair and anti-entropy as the reconciliation mechanisms, and reason about when to enable/disable sloppy quorums for a given workload.

for a principal

Should discuss durability edge cases (hint-holder crash before handoff), the specific Dynamo design rationale (always-writable for shopping carts), and how to design monitoring/alerting to detect prolonged sloppy-quorum divergence in production.

## The strict quorum it relaxes To understand sloppy quorums you first need the strict quorum they relax. In a strict-quorum system, each key is deterministically mapped (often via consistent hashing) to a fixed set of N 'owner' nodes — the **preference list**. - A write must get W acknowledgments from nodes in that specific preference list. - A read must get R responses from that same list. That is exactly what makes the R+W>N overlap argument valid: both operations draw from the same fixed pool of N nodes. If, during a network partition or a node outage, fewer than W of the designated owners are reachable from the coordinator, a strict-quorum system simply fails the write — it refuses to substitute other nodes, because doing so would break the mathematical guarantee that makes quorums useful in the first place. ## What a sloppy quorum does instead A sloppy quorum takes the opposite stance on availability. Instead of failing the write when too few of the 'correct' owners are reachable: 1. The coordinator picks other healthy nodes in the cluster — any nodes it can reach — and writes the data there instead, counting those acknowledgments toward W. 2. Each such 'borrowed' write carries a **hint**: metadata recording which node was actually supposed to receive this data. This mechanism, **hinted handoff**, means the write still succeeds and the client still gets an ack, even though the data temporarily lives on the 'wrong' nodes. 3. Later, once the intended owner node recovers and rejoins the cluster, the nodes holding hints proactively push (hand off) the data to it, and then discard their temporary copy. The design goal is explicit: Amazon's Dynamo paper introduced this to guarantee that writes are always accepted (the shopping-cart use case prioritized 'always writable' over strict correctness — losing a cart update is worse than briefly serving a stale cart). ## What you give up The cost is precisely the overlap guarantee that R+W>N was built to provide. A strict quorum read still queries R nodes from the fixed preference list — but if a recent write went to sloppy/hinted nodes outside that list (because the real owners were briefly unreachable), the read's fixed preference-list nodes may not have that write yet at all, hint or no hint, until the handoff completes. So the read can return a version older than the latest acknowledged write, something the strict R+W>N argument explicitly promised couldn't happen. In other words, sloppy quorums intentionally trade the deterministic-overlap consistency property for availability during exactly the failure scenarios (partitions, node outages) where a strict system would instead refuse to accept the write at all. This is a direct, concrete instance of the availability side of the CAP trade-off, applied selectively (per-write) rather than as a blanket policy. ## Closing the divergence window Because sloppy quorums create a window of divergence, real systems pair them with repair mechanisms to eventually reconverge state: - **read repair** — when a read's R responses disagree, the coordinator pushes the newest version back to the stale replicas it just saw - **background anti-entropy processes** — e.g. Merkle-tree comparison between replicas run periodically to find and fix divergence that read repair never happened to trigger, since read repair only fixes keys that actually get read Without these, the temporary inconsistency introduced by sloppy quorums could persist indefinitely for cold, rarely-read keys. ## When to turn it on Operationally, the decision to enable sloppy quorums is a knob, not a fixed property of the algorithm — Cassandra exposes it as a per-operation config, sometimes as a 'write to any available replica' fallback behavior, while Dynamo enabled it as a core design default. | Workload | Usual stance | |---|---| | Teams that need strict linearizable-ish reads-after-writes (e.g. a financial ledger balance) | generally disable sloppy quorums and accept write failures during partitions | | Teams optimizing for 'never reject a write' (e.g. a shopping cart, a like counter, presence/session data) | accept the temporary staleness window in exchange for never bouncing a client's request during a transient outage |

  • Does a sloppy quorum ever get used for reads, or only for writes?
    The classic Dynamo mechanism is specifically about writes accepting substitute nodes; reads still query the designated preference list. Some implementations extend a similar 'accept any healthy replica' idea to reads under partition too, but the canonical hinted-handoff pattern is a write-side availability mechanism.
  • How does hinted handoff avoid data loss if the hint-holding node crashes before forwarding?
    It's not fully immune — if the node holding the hint dies before handoff completes and before the write also replicated normally elsewhere, that write copy can be lost. This is why sloppy quorums are explicitly a durability/availability trade-off, and why anti-entropy repair (comparing replicas independently of hints) exists as a second line of defense.
  • If you disabled sloppy quorums entirely, what would you get back, and what would you lose?
    You'd get back the deterministic R+W>N overlap guarantee — a strict-quorum read is guaranteed fresh relative to the last strict-quorum write. You'd lose availability during partitions: any write that can't reach W of its designated owners fails outright instead of degrading gracefully.

Normally your mail only goes to your own mailbox (the strict quorum). If your mailbox is jammed, a sloppy quorum is the mail carrier leaving your package with a neighbor instead, with a sticky note ('this is really for #42'), and later walking it over to your mailbox once it's fixed. The package didn't get lost, but if someone checks your mailbox in the meantime, they won't find it.

saying these in an interview costs you the question

  • Confuses sloppy quorum with just lowering R or W numerically
  • Thinks hinted handoff is a form of consensus or leader election
  • Believes sloppy quorums still guarantee R+W>N overlap
  • Can't say what problem (availability during partition) sloppy quorums solve
  • Assumes sloppy quorums are Cassandra/Dynamo-only trivia with no general applicability

context