skip to content

A candidate says the caught-up set is always a leader's list of current followers — on which cluster designs is that wrong?

level: seniorimportance: nice to knowfreq 36%

answer

  1. one phrase, several machineries
  2. standing membership or decided per write
  3. majority answering needs no list
  4. shared storage has no follower to be behind

basics

~20 s

Only leader-follower designs keep a standing membership list. Majority-write designs decide eligibility per write — whichever copies answer first — with no list and no drop-out event, and shared-storage designs have no per-node copies to be current at all.

solid answer

~40 s

The list is one shape of several. On **leader-follower** platforms the leader does maintain an explicit membership of copies it currently counts as current, with an allowance that decides who is in it and drop-out and re-admission as observable events. On **majority-write** platforms there is no standing list: a write is committed when a majority of copies has answered, so "caught up" is decided per write by who answered, and a slower copy just trails and catches up with no membership change. In designs where **shared durable storage replaces per-node copies**, brokers hold no per-record replicas, so the question becomes whether the underlying store accepted the bytes. And some brokers keep exactly one **mirrored copy** rather than a configurable set. The operator consequence is that "a copy is behind" costs something different in each.

go deeper

for a junior

Remember the phrase names a question — which copies count right now — and that different clusters answer it with different machinery, not one universal list.

for a middle

Describe at least two shapes and what distinguishes them: a maintained membership with drop-out and re-admission, versus eligibility decided per write by whoever answered.

for a senior

Show that you check the shape before reading any signal, because the same underlying trouble appears as a narrowing membership on one platform and as failing writes on another.

for a principal

Make the shape explicit in any estate-wide durability standard, so a tier written for one cluster design is not applied unchanged to a cluster whose success answer means something else.

## The question behind the vocabulary Every platform in this class must answer one question before it can say a write is safe: **which copies count right now?** The phrase "the caught-up set" names the answer, but the machinery behind it differs enough that an operator moving between platforms can carry the wrong mental model for years without noticing. ## Shape A — a leader with a membership of current followers This is the model most people mean. One copy leads; the others pull from it. The leader maintains an explicit membership of the copies it currently counts as current, judged by **the catch-up window** — how far behind a copy may fall in records or in elapsed time. Its distinguishing features: - membership is a **standing fact** you can ask about at any instant, independent of any particular write; - **dropping out** and **re-admission** are discrete events that happen to a named copy; - the durability promise is phrased relative to that membership, so the promise's strength moves with the membership's width. ## Shape B — a majority of copies answering Here a write is committed once **a majority of the copies has answered**. There is no maintained list of who is current, and no allowance to cross. - Eligibility is decided **per write**: the copies that answer in time for *this* write are the ones that made it count. - A slower copy is not ejected from anything. It simply was not among the fastest majority, and it continues to receive and apply records afterwards. - The failure mode is different in kind: when too few copies can answer, **the write itself stops succeeding**, loudly, rather than succeeding against a quietly narrowed membership. A candidate who insists on a list here will look for a drop-out event that does not exist, and will misread a slow copy as a durability incident when it is only a latency contributor. ## Shape C — durability from shared underlying storage Some designs put the data in a replicated store beneath the brokers rather than in N broker-held copies. The brokers become largely stateless with respect to record durability. - There is no follower pulling from a leader, so nothing is "behind" in this sense. - The equivalent question is whether the underlying store accepted and replicated the bytes, and its own redundancy — not a broker-level set — is the durability posture. - Broker-level copy counts and catch-up allowances may not exist as settings at all. ## Shape D — a single mirrored copy Some brokers, especially queue-shaped ones, keep exactly **one paired copy** rather than a configurable number. Currency still matters, but the vocabulary of a set of size N and an allowance for membership collapses: either the paired copy is up to date or it is not, and the operator's lever is whether to pair the queue at all. ## Side by side | | Standing membership? | "Behind" shows up as | What a slow copy costs | |---|---|---|---| | Leader with current followers | yes, maintained by the leader | a copy dropping out of the set | narrows what a durability promise is worth | | Majority write | no — decided per write | slower writes, then failures when too few answer | write latency, until too few copies remain | | Shared durable storage | not at the broker level | whatever the underlying store reports | depends on the store, not on broker copies | | Single mirrored copy | trivially, one paired copy | the pair being out of date | loss of the only redundancy that queue has | ## Why this matters operationally The practical trap is transfer of habit. An operator who learned Shape A expects a narrowing membership to be the early warning and treats successful writes as proof of nothing; an operator who learned Shape B expects the writes themselves to start failing and is not looking for a silent condition. Move the first to Shape B and they will hunt for a metric that does not exist; move the second to Shape A and they will trust successful answers that were satisfied by one machine. So the honest answer to "what is the caught-up set?" starts with a question of its own: **which of these does this cluster do?** That is not pedantry — it decides whether a durability problem here announces itself as a failed write or as nothing at all.

  • On a majority-write design, is there any sense in which a copy is "out of the set"?
    Only per write. For each write there is the set of copies that answered in time and made it count, and the rest catch up afterwards. That set is recomputed constantly and is not a property you can query about a copy the way a standing membership is, so there is no drop-out event and no re-admission.
  • Why does the shape change how a durability problem first shows itself?
    Because the two shapes fail in opposite directions. With a standing membership, the promise is relative to it, so a narrowed membership keeps returning success while protecting less. With a majority requirement, losing copies stops the write from committing at all. The first is a silent durability exposure, the second is a visible write failure.

saying these in an interview costs you the question

  • Assumes every platform maintains a standing membership of current copies
  • Expects a drop-out event on a design that decides eligibility per write
  • Thinks majority-write designs do not replicate each write at all
  • Believes broker-level copy counts exist on every platform in this class
  • Treats a slow copy as identically dangerous regardless of the cluster's design