For a new write-heavy service, how would you choose between a leaderless wide-column store and one that serves each key range from a single server?
answer
- what hurts more, pause or staleness
- read-your-write assumptions
- counters and conditional updates
- where the writes originate
- who runs it
basics
~20 sChoose by which failure the product tolerates and which operations it needs. Leaderless stores favour write availability and multi-site writes but make conditional updates costly; range-served stores give strong single-row reads and cheap check-and-mutate at the price of brief per-range unavailability.
solid answer
~50 sI would start from the product's tolerances, not the store's features. **Is a few seconds of unavailability for some keys worse than a stale or conflicting read?** If writes must always be accepted, including from several sites at once, the leaderless model fits, and I would design around timestamp conflict resolution and keep conditional writes rare. If the code relies on **reading its own writes, counters, or check-and-set**, the range-owner model makes those cheap and correct by default, and I would plan for failover pauses and for hot rows landing on one server. Then the practical filters: key distribution (sorted keys need hotspot care), the operational model (a managed service versus a cluster the team runs, including repair and compaction), and how data reaches analytics. I would write the decision down as the specific operations and failure behaviours we are buying.
go deeper
Know that the two models behave differently when a server fails and when two writes collide.
Explain which operations are cheap in each model, conditional writes and counters, and which failure each shows, pause or staleness.
Map a real service's operations and failure tolerance onto the models and point out the design changes each would force.
Own the decision end to end, weigh product tolerance, operations and team capacity, write down the accepted costs, and name the triggers to revisit it.
## Frame the decision around behaviour Both models store data the same way — sorted keys, log-structured files, compaction. They differ in **who accepts a write and what happens when a server fails**. A good decision therefore starts from what the product can tolerate, not from benchmark numbers. ## Question 1: which failure hurts more? - **Leaderless replicas** keep accepting writes when nodes fail, as long as enough replicas answer. The failure mode is **staleness or conflict**: a read may miss a recent write, and concurrent writes are settled by timestamp. - **One server per key range** keeps a single writer per row. The failure mode is a **brief pause** for the failed server's ranges while they are reassigned and the log is replayed. A shopping cart or activity feed often prefers "always writable, occasionally stale". A balance-like counter or a uniqueness claim often prefers "briefly paused, never wrong". ## Question 2: which operations does the code need? | need | leaderless | range-served | |---|---|---| | read-your-own-write by default | needs appropriate request settings | yes, within a cluster | | atomic increments, check-and-mutate | consensus path, costly | local and cheap | | writes accepted in several sites at once | replicas in every site belong to one replica set | possible across replicated clusters, but asynchronous and last-write-wins, with single-row atomic operations confined to one cluster | | ordered scans across many keys | only within a partition | any key range | | hot single key | spread over replicas | one server absorbs it | If the design leans on the middle rows, the range-owner model removes whole classes of bugs. If it leans on multi-site writes and constant availability, the leaderless model is the natural fit. ## Question 3: the practical filters 1. **Key distribution.** Sorted-key stores need deliberate hotspot avoidance for sequential keys; hashed partitions spread by default but lose cross-key order. 2. **Operational model.** A managed service removes node operations but binds you to one provider; a self-run cluster needs people who handle compaction tuning, repair schedules and upgrades. 3. **Surrounding ecosystem.** Where does data go for analytics, how are backups taken, what change stream is available? 4. **Team experience.** A store the team can operate at 3 a.m. beats a theoretically better one. ## Question 4: what would make us revisit? Record the triggers in the decision: conditional writes becoming a hot path, a multi-region write requirement appearing, sustained hotspots, or operational load outgrowing the team. Naming them turns a one-way door into a reviewed one. ## Worked example A telemetry ingestion service writes millions of readings per minute from devices worldwide, reads by device and time, and tolerates slightly stale dashboards. Devices write from several regions and must never be refused. That points to the **leaderless** model, with keys bucketed by device and time and no conditional writes on the ingest path. A billing-usage service increments per-customer counters and must never double-count. That points to the **range-owner** model for its cheap atomic increments and read-your-writes, with a plan for failover pauses and for large customers' hot rows. ## Interview angle There is no single right answer. What the interviewer is testing is whether you choose from failure tolerance and required operations, name the costs you are accepting, and record what would make you change course.
- Could one system use both models?Yes, and many do. An ingestion path can land in a leaderless store while counters or uniqueness claims live in a range-served store or a relational database. The cost is two systems to operate and a boundary where data must be kept consistent.
- How would you test the choice before committing?Build the critical read and write paths against both, run a load test with realistic key distribution, and kill servers mid-test. Watch what the application sees, errors, pauses or stale reads, rather than only throughput.
saying these in an interview costs you the question
- Choosing by peak write benchmark alone without considering failure behaviour
- Assuming both models give the same single-row consistency by default
- Ignoring who will operate compaction, repair and upgrades
- Picking the leaderless model while building the hot path on conditional writes