An e-commerce team is building a 'Recently Viewed Items' feature and a 'Payment Balance' feature. One can tolerate showing slightly stale data for a few seconds; the other cannot ever show a wrong number. Which consistency model would you pick for each, and why does that trade-off exist at all?
answer
- multiple copies = coordination cost
- strong = leader/quorum before ack
- eventual = ack local, sync later
- CAP: partition forces A or C
- PACELC: even healthy, L vs C
basics
~20 sUse fast, "good enough" data for low-stakes info like recently viewed items, and exact, up-to-date data for money like a payment balance. Keeping everything instantly perfect everywhere is slow, so you only pay that cost where being wrong actually hurts.
solid answer
~30 s"Recently Viewed Items" can use eventual consistency: any replica can serve the read, writes propagate asynchronously, and a few seconds of staleness is invisible; this buys low latency and high availability. "Payment Balance" needs strong consistency: every read must reflect the latest committed write, so requests get routed to a leader or quorum enforcing ordering, accepting higher latency and reduced availability during a partition. The trade-off exists because synchronously agreeing across replicas costs time and, during a partition, availability (per CAP); you spend that cost only where incorrect answers cause real harm.
go deeper
Should articulate the basic trade-off (fast-but-maybe-stale vs. slow-but-exact) in plain language and correctly assign each feature to a model, without needing to name CAP or PACELC explicitly.
Should name CAP, explain that strong consistency typically requires a leader or quorum, and connect the trade-off to concrete latency/availability costs, not just 'strong is safer.'
Should reason about the trade-off per read/write path, mention PACELC's latency-vs-consistency axis even absent a partition, and give a production example of a real bug caused by a mismatched choice.
Should treat consistency-model selection as an architectural policy applied consistently across a system's feature catalog, discuss how to communicate these choices to a team, and connect it to SLA/SLO design and failure-mode runbooks.
## Why the data is copied at all Every piece of data in a distributed system is deliberately stored on more than one machine: - **Durability** — a disk dies, you don't lose data. - **Availability** — one node goes down, others still answer. - **Latency** — a reader in Tokyo shouldn't have to wait on a write that committed in Virginia. The moment you have multiple copies of the same fact, you must decide how tightly those copies have to agree before you'll let a client read from any of them. That decision is "the consistency model," and it is a knob you turn per feature or per read/write path — not a single global setting for the whole system. ## Strong versus eventual, mechanically | Compared on | Strong consistency | Eventual consistency | |---|---|---| | Usually implemented by | routing writes through a single leader replica, or requiring a quorum to agree before acknowledgment | typically acknowledged locally, then propagated asynchronously | | A later read | returns that write's value or something newer | can still return the old value | | Cost | latency, and availability during a network partition | correctness: a client can observe a value that's already stale | **Strong consistency** means that once a write is acknowledged, every subsequent read — from any replica, anywhere — returns that write's value or something newer; there is no window in which a client can observe stale data. Mechanically this is usually implemented by routing writes (and often reads) through a single leader replica, or by requiring a quorum of replicas to agree before a write or read is acknowledged, so ordering is enforced before the client ever sees a result. **Eventual consistency** means that after a write is acknowledged, replicas converge to the same value over time, but a read issued moments later against a different replica can still return the old value. Writes are typically acknowledged locally and propagated asynchronously — via replication logs, gossip, or background anti-entropy — so acknowledgment is fast and doesn't wait on far-away replicas to catch up. ## Where the cost lands The trade-off exists because coordination has a real cost, and that cost lands on one side or the other. - **Strong consistency costs latency**: every write, and often every read, waits on a round trip to a leader or on agreement from a majority of replicas, which is expensive across regions and adds up under load. - **It also costs availability during a network partition** — per the CAP theorem, if a replica can't reach the leader or reach quorum, it must refuse to serve rather than risk returning stale or conflicting data, so part of the system goes dark rather than goes wrong. - **Eventual consistency buys that latency and availability back** — reads are served from the nearest local replica, writes acknowledge instantly — but the cost lands on correctness: a client can observe a value that's already stale, two different clients can get two different answers to "what is X right now," and if two replicas accept conflicting writes independently, something or someone has to reconcile them later. **PACELC** sharpens this further: even with no partition at all (the "else, latency" branch), you still trade latency against consistency, because a synchronously replicated write is inherently slower than a fire-and-forget one — the cost isn't only a partition-time cost, it's a tax paid on every single operation. ## Two recognizable failure classes In production, picking the wrong model for a given path shows up as two distinct, recognizable failure classes. 1. **Choose eventual consistency where you actually needed strong**, and you get correctness bugs that are maddening to reproduce because they depend on replica lag and routing: a payment-balance read served from a lagging replica tells a customer they have funds they've already spent; a "your order is confirmed" screen reads its own write from a different replica right after checkout and shows "order not found," a classic read-your-own-writes violation. 2. **Choose strong consistency where eventual would have been fine**, and you get latency and availability bugs instead: a "recently viewed items" widget that blocks page render on a cross-region leader round trip for no real benefit, or — worse — an entire low-stakes feature going fully unavailable during a routine network blip in the leader's region, even though nobody actually needed that read to be perfectly fresh. ## A concrete split — DynamoDB exposes both A concrete, real split: Amazon's DynamoDB exposes both modes on the very same table. - Its **default read is eventually consistent** — cheaper, lower latency, and it may reflect a slightly older replication state. - A **strongly consistent read** can be requested explicitly, at higher latency and cost, and is guaranteed to reflect the result of the most recently successful write. A checkout flow decrementing inventory before confirming a purchase would request the strong read to avoid overselling the last unit of a product; a "customers who viewed this also viewed" panel would happily use the default eventual read, because a few seconds of staleness there costs the business nothing. ## The judgment The engineering judgment isn't "always pick strong" — that tanks latency and availability everywhere, including places where nobody cares — nor "always pick eventual," which quietly ships correctness bugs into money-moving paths. It's mapping each read and write path individually to how much staleness that specific feature can actually tolerate, and paying the coordination cost only where it's load-bearing.
- What does 'read-your-writes' consistency mean, and which of the two features in this scenario would break it under eventual consistency?Read-your-writes means a client is guaranteed to see the effects of its own prior write, even if other clients might not yet. Under eventual consistency, the payment-balance feature could break this guarantee if a customer's own follow-up read routes to a replica that hasn't received their write yet — the recently-viewed-items feature breaking this guarantee is far less costly since nobody depends on instantly seeing their own view history.
- How does causal consistency fit between eventual and strong consistency for a feature like a comment thread?Causal consistency guarantees that operations with a cause-effect relationship (a reply is seen after the comment it replies to) are observed in that order everywhere, while unrelated concurrent writes can still be reordered or seen at different times by different readers. It's a middle ground: cheaper than strong consistency because it doesn't need global ordering, but stronger than plain eventual consistency because it prevents nonsensical orderings like seeing a reply before the comment it replies to.
- If the payment-balance API is built on strong consistency via a single leader, what happens to write availability if that leader's region partitions off from the rest of the cluster?Writes to that leader either become unavailable for affected clients (if followers refuse to serve since they can't confirm leadership) or the cluster fails over to a new leader elected from the reachable majority, briefly pausing writes during election. Either way, you accept an availability hit specifically to avoid two leaders diverging and producing conflicting balances.
Like choosing between calling every sibling to privately confirm before updating the family group chat (accurate, slow) versus just posting and letting everyone catch up when they check the chat (fast, briefly out of sync) — you only make the slow call for news that must be right immediately, like 'Mom's flight got cancelled,' not for 'I had pizza for lunch.'
saying these in an interview costs you the question
- Says a system should always use strong consistency 'to be safe' without acknowledging the latency/availability cost
- Treats consistency as a single system-wide setting rather than a per-feature/per-path choice
- Confuses eventual consistency with 'the data is wrong' rather than 'the data is temporarily stale'
- Can't name a concrete mechanism (leader routing, quorum, async replication) for either model
- Doesn't recognize that a partition forces a real choice (CAP) rather than assuming both properties are simultaneously achievable