skip to content

You're picking a data store for a distributed lock service used for leader election versus a data store for shopping-cart sessions on a high-traffic e-commerce site. Walk through why you'd lean CP for one and AP for the other, and name the concrete cost each choice imposes on that specific system during a network partition.

level: seniorimportance: must knowfreq 70%

answer

  1. locks/leader-election = CP: wrong answer means split-brain
  2. shopping cart/session = AP: wrong answer is a mergeable annoyance
  3. rule of thumb: CP when wrong beats none, AP when none beats wrong
  4. minority side of a CP service refuses, doesn't lie
  5. AP service reconciles conflicts after the partition heals

basics

~20 s

For the lock service, two 'leaders' at once could break everything, so it's worth going offline rather than risk that (CP). For shopping carts, showing an old cart is annoying but rarely catastrophic, and going offline during checkout loses sales, so keep answering even with slightly stale data (AP).

solid answer

~40 s

A distributed lock or leader-election service must be CP: if a partition lets two nodes each believe they hold the lock, you get split-brain and potentially corrupted shared state, far worse than the service being briefly unavailable. The cost is that during a partition, the minority side of the lock service refuses to grant or renew locks, and any client relying on it stalls or errors until the partition heals. A shopping-cart store, by contrast, should be AP: showing a customer a slightly stale cart is a bad but recoverable experience, whereas refusing to let anyone add to their cart during a network blip directly costs revenue. The cost there is eventual, application-level reconciliation of conflicting cart edits once the partition heals.

go deeper

for a junior

Should be able to say, in plain terms, that a lock breaking is worse than a lock being briefly unavailable, and that a stale shopping cart is a minor annoyance.

for a middle

Should name split-brain as the specific danger for the lock service and describe, at a high level, what happens to requests on the losing side of the partition for each system.

for a senior

Should give the full cost-of-wrong-answer-vs-cost-of-no-answer reasoning for both systems, name a concrete mechanism, such as quorum, that enforces the CP choice, and describe a reconciliation strategy for the AP choice.

for a principal

Should generalize this into a reusable decision framework applicable across a whole architecture, recognize that the same physical system can host both CP and AP datasets side by side, and flag the failure mode of applying the wrong choice to a given dataset.

## The choice belongs to the dataset Choosing CP versus AP isn't a property of 'distributed systems' in the abstract — it's a decision made per dataset, based on what happens when that dataset gives a wrong answer versus what happens when it gives no answer. The clearest way to see this is to compare two systems with genuinely different failure economics: a distributed lock service used for leader election, and a shopping-cart store on an e-commerce site. ## Why the lock service leans CP A leader-election or distributed-lock service exists to answer one question correctly: exactly one process may hold the lock, or believe it is the leader, at any given time. The entire value of the service comes from that exclusivity guarantee — callers use it precisely because they need to coordinate access to something that breaks if two processes touch it concurrently, such as: - a shared file; - a job queue; - or a primary database role. If a network partition splits the cluster backing this service into two halves, and both halves independently decide to grant the lock, you get **split-brain**: two processes both believe they are the exclusive owner and act accordingly, potentially both writing to the same resource in conflicting ways. This is a correctness disaster, not a minor glitch. So the rational choice is **CP**: during a partition, the minority side of the cluster refuses to grant or renew any lock. Clients depending on that lock on the minority side stall, time out, or fail their operation. That is the accepted cost — a temporary outage for whatever depended on the lock — in exchange for the guarantee that nobody mistakenly believes they hold exclusive access when they don't. ## Why the shopping cart leans AP A shopping-cart store has the opposite failure economics. Its job is to remember what a customer put in their cart so they can check out later. If a partition happens and the cart service keeps answering using whichever local replica is reachable, the worst realistic outcome is that two conflicting edits happened on two sides of the partition and need reconciling once the partition heals, typically by a simple union-of-items merge or a last-write-wins rule with a re-add prompt if something's lost. That's annoying but recoverable, low-stakes. Compare that to making the cart service CP, so that during a partition it refuses requests on the minority side: on a high-traffic e-commerce site, that directly means customers can't shop, and every second of that outage is measurable lost revenue. So the rational choice here is **AP**: keep serving from local state, accept that carts can briefly diverge, and reconcile later. ## The rule of thumb This comparison generalizes into a rule of thumb. | The call | Where it lands, and why | |---|---| | **Pick CP when an incorrect answer is worse than no answer** | Coordination primitives like locks, leader election, unique-ID allocation, or anything enforcing an invariant such as 'balance never goes negative' fall here, because a wrong answer can cascade into corrupted downstream state. | | **Pick AP when no answer is worse than a slightly-wrong answer** | User-facing, individually-scoped, easily-mergeable data like shopping carts, likes, view counts, or session data falls here, because the business cost of an outage usually exceeds the cost of temporary staleness. | ## Getting the choice backwards The failure mode to watch for in production is applying the wrong choice to a dataset: - Teams sometimes make a lock service AP for uptime reasons, not realizing they've reintroduced the split-brain risk the lock was built to prevent — it now always answers, but no longer means anything. - The opposite mistake — making a high-traffic, low-stakes dataset CP — shows up as unnecessary, business-impacting outages during ordinary network blips a slightly-stale answer would have absorbed painlessly. The right call always comes back to naming, concretely, what a wrong answer costs versus what an outage costs for that specific piece of data.

  • What concrete mechanism stops both sides of a partitioned lock service from granting the same lock?
    A quorum requirement: the lock is only granted or renewed by a node that can confirm it's part of a majority of the cluster's total nodes. Since two disjoint majorities of the same node set can't both exist, at most one side of any partition can ever have a majority, so at most one side can legally grant the lock; the other side is mathematically unable to claim quorum and must refuse.
  • How would you reconcile two shopping carts that diverged during a partition once the network heals?
    The simplest common approach is a union merge: combine the items from both divergent cart versions rather than picking one as the winner, since silently dropping an added item is rarely what a customer wants. For fields where union doesn't make sense, like a chosen shipping address, a timestamp-based last-write-wins rule is common, sometimes with a prompt to the customer if the merge looks ambiguous.
  • Could the shopping-cart store still offer some consistency guarantee while remaining AP overall?
    Yes — AP doesn't mean 'no consistency guarantee,' it means the system won't sacrifice availability to get the strongest one. Many AP stores still guarantee, for example, that a client always sees their own writes reflected in their own subsequent reads, even while allowing different clients or replicas to briefly disagree. The exact strength of that weaker guarantee is a separate topic from the CP/AP choice itself.

A lock service is like a single conference room key — if two copies of the key both work at once, two meetings double-book the room and chaos follows, so better to lock everyone out until you're sure only one key is live. A shopping cart is like a sticky note on the fridge — if two family members each add an item to a duplicate copy during a power outage, you just merge the lists once the power's back; nobody's hurt by the temporary confusion.

saying these in an interview costs you the question

  • Picks CP or AP for 'distributed systems in general' rather than per dataset based on the cost of a wrong answer
  • Doesn't identify split-brain as the specific risk a CP lock service prevents
  • Assumes AP means no consistency guarantee at all
  • Can't name what the minority side of a CP system actually does versus an AP system
  • Treats the CP/AP decision as a one-time architectural label rather than something that can differ operation-by-operation within the same system

context