skip to content

How does a cluster decide that a follower copy has fallen far enough behind its leader to drop out of the caught-up set?

level: middleimportance: must knowfreq 64%

answer

  1. close enough, not identical
  2. an allowance with a unit
  3. records short, or time since current
  4. a burst is not a sick disk
  5. too tight and copies flap

basics

~20 s

An allowance called the catch-up window sets how far behind a follower copy may be — a number of records, or elapsed time since it was last current. Past it the leader drops the copy; back inside it, the copy is re-admitted.

solid answer

~50 s

Platforms that maintain a caught-up set need a boundary, and that boundary is **the catch-up window**: how far a follower copy may trail its leader before it stops being counted. Two units are used. A **record-count** window asks how many records the copy is short; it is simple, but a traffic burst puts every healthy follower thousands of records behind at once, so the boundary fires on load rather than on sickness. An **elapsed-time** window asks how long it has been since the copy was last current, which a burst does not trip as long as the follower is still pulling. Crossing the boundary drops the copy from the set; closing the gap re-admits it. Sizing matters in both directions: too tight and healthy copies flap in and out under normal variation, too loose and a copy nobody would call current is still counted as one.

go deeper

for a junior

Remember that "caught up" is defined by an allowance, not by being identical — a copy is counted while it is within so many records, or so much elapsed time, of the leader.

for a middle

Be able to explain both units and the burst case that separates them, plus what happens at the boundary in each direction: drop out when the gap widens, re-admit when it closes.

for a senior

Show judgment about sizing: name flapping as the symptom of an allowance that is too tight for the hardware, and a stale copy still being counted as the symptom of one that is too loose.

for a principal

Treat the allowance as a threshold with a false-positive and a false-negative side, and decide what a house default should be across an estate whose clusters do not share hardware or traffic shape.

## Why a boundary is needed at all On a leader-follower design, every follower copy of a stream is always at least a little behind the leader — it has to read a record before it can hold it. So "caught up" cannot mean "identical". It has to mean "close enough", and something has to define close enough. That definition is the **catch-up window**: the allowance a copy has before the leader stops counting it as part of **the caught-up set**. The window is not a health check and not a failure detector. It is a currency threshold, evaluated against **the position a follower has replicated to** compared with the leader's newest record. ## The two units, and why the choice matters | Window unit | What it asks | Where it holds up | Where it misleads | |---|---|---|---| | A number of records | how many records short of the leader is this copy? | steady traffic, uniform record sizes | a burst pushes every healthy follower past it at once, so load looks like sickness | | Elapsed time | how long since this copy was last current with the leader? | bursts, uneven record sizes, mixed streams | a copy that keeps making requests but never closes the gap can look better than it is | The burst case is the one that decides the argument in practice. Suppose the write rate into a stream multiplies for thirty seconds. Every follower is now further behind in records than it was, not because anything is wrong with any of them but because there is simply more to pull. A record-count window cannot tell that apart from three sick disks; an elapsed-time window can, because each follower is still reaching the leader's current position moments after the leader reaches it. This is why platforms that use a time-based allowance describe it as *time since the copy was last caught up*, rather than *time since we last heard from it* — the second is liveness, which is a different subject, and a copy can be perfectly alive and hopelessly behind. ## Dropping out and being re-admitted Both transitions are routine, not exceptional: 1. A follower's disk starts retrying, or the link between its failure domain and the leader's saturates. Its gap widens. 2. The gap crosses the catch-up window. The leader removes it from the caught-up set. Nothing is deleted; the copy keeps every record it already pulled. 3. The copy carries on pulling. To rejoin, it must read faster than the leader is writing — it needs surplus rate, not merely a working disk. 4. Once it is inside the window again, the leader re-admits it and the set widens back. ``` leader writes ──────────────────────────────────────────▶ copy A close behind ............................. in the set copy B close behind ..╲ disk slows ╲............. drops out copy B ╲.. still pulling ..╱..... re-admitted ``` Step 3 is where a lot of intuition fails. A copy that dropped out because it cannot keep up with the incoming rate does not recover simply because time passes; it recovers when it has spare throughput to close a gap *and* keep pace with new arrivals. ## Sizing the window - **Too tight** and copies **flap**: normal variation — a compaction on the volume, a brief link contention, a garbage burst — pushes a healthy follower over the line and back repeatedly. Each transition re-narrows and re-widens the set, and the durability the stream offers oscillates for no useful reason. - **Too loose** and the set lies in the other direction: a copy well out of date is still counted, so writes are recorded as backed by copies that do not hold them yet. - The honest framing in an interview is that the window is a **false-positive / false-negative trade**, exactly like any threshold: tighten it and you eject healthy copies, loosen it and you keep stale ones. Platforms pick a default that tolerates ordinary variation; the point of knowing the setting exists is that a cluster whose hardware or traffic shape is unusual may need a different one. ## Where this does not apply A catch-up window belongs to designs that maintain a standing membership of current copies. Platforms that commit a write when a majority of copies has answered do not need one: a slow copy is simply not among the majority that answered *this* write, and it catches up afterwards with no drop-out event. In designs where durability comes from shared underlying storage rather than per-node copies, there is no follower to be behind in this sense at all. Before tuning any allowance, confirm your platform is the shape that has one.

  • Why do several platforms prefer an elapsed-time allowance over a record-count one?
    Because a record-count allowance cannot distinguish load from illness. When the write rate spikes, every healthy follower is suddenly thousands of records short and all of them cross the line at once, narrowing the set during exactly the burst you most wanted copies for. A time-based allowance asks when the copy was last current, which a healthy follower keeps satisfying through the burst.
  • A follower copy drops out and the operator sees it rejoin and drop again every few minutes. What is that telling you?
    It is flapping: the allowance is tight relative to that copy's normal variation, or the copy has only just enough throughput to hover at the boundary. Either the window is mis-sized for this hardware and traffic shape, or the copy's disk or link is marginal. The oscillation itself is the signal — a genuinely healthy copy sits well inside the boundary, not on it.
  • Does a copy have to close its gap completely before it counts again?
    Designs differ. Some re-admit a copy as soon as it is back inside the allowance, which by definition means a small residual gap is acceptable. Others require it to reach the leader's current position once before counting it. Either way the copy must out-pace the incoming write rate to get there, which is why a copy short on throughput can stay out indefinitely.

saying these in an interview costs you the question

  • Says a caught-up copy is byte-identical to its leader
  • Thinks the allowance measures whether the copy is alive rather than whether it is current
  • Assumes a record-count allowance and a time allowance behave the same under a traffic burst
  • Believes a copy rejoins automatically once its disk recovers, regardless of the incoming rate
  • Treats a tighter allowance as strictly safer, ignoring that healthy copies then flap in and out
  • Assumes every platform maintains such an allowance